Gemini 3.8 Flash Matches Opus 5 on Coding, Then Scores 19.1% as a General Agent
Google shipped Gemini 3.8 Flash on September 2 at the same $0.75 per million input tokens as 3.7 Flash. DeepSWE climbs from 65.3% to 73.7%, but general agent work lands at 19.1% against Claude Opus 5 at 51.8%.
- Gemini 3.8 Flash ships at the same $0.75 / $3.75 per million tokens as 3.7 Flash.
- DeepSWE climbs from 65.3% to 73.7%, level with Claude Opus 5 at 74.0%.
- General agent work stalls at 19.1% on Terminal-bench 4.0 against Opus 5 at 51.8%.
Google released Gemini 3.8 Flash on September 2, 2026, three weeks after 3.7 Flash shipped on August 13. By Google's own count it is the third Flash release in six weeks. Not a single line of the price sheet changed: $0.75 per million input tokens and $3.75 per million output, identical to 3.7 Flash, down to the same introductory-price expiry (December 31, 2026) and the same list price after it ($1.50 and $7.50).
When the price holds, one question is left. Can you move the work you run today on 3.7 Flash or Claude Sonnet 5 over to 3.8 Flash? The benchmark table Google published alongside the announcement answers in two directions. For bounded coding, document, and analysis tasks, yes. For operating a computer or running autonomously for long stretches, not yet.
What improved over 3.7 Flash
3.8 Flash beats 3.7 Flash on all 15 rows in the published table. The largest gain is coding. DeepSWE v1.1, a long-horizon software engineering benchmark, goes from 65.3% to 73.7%, up 8.4 points. In the same table Claude Opus 5 sits at 74.0% and GPT-5.6 Sol at 72.7%. A model priced at one-seventh of Opus 5's input rate has effectively pulled level.
The main rows are below. Google published the table only as an image, so these figures were read off the original and re-typeset here.
| Benchmark | 3.8 Flash | 3.7 Flash | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Input price (1M tokens) | $0.75 | $0.75 | $5.00 | $4.00 |
| Output price (1M tokens) | $3.75 | $3.75 | $25.00 | $20.00 |
| DeepSWE v1.1 (long-horizon coding) | 73.7% | 65.3% | 74.0% | 72.7% |
| Terminal-bench 2.1 (terminal coding) | 89.4% | 85.8% | 89.1% | 88.8% |
| Vals Finance Agent v2 | 61.4% | 59.0% | 58.6% | 53.8% |
| CharXiv (chart reading and reasoning) | 86.2% | 84.5% | 83.7% | 85.8% |
| Terminal-bench 4.0 (general agent) | 19.1% | 11.2% | 51.8% | 37.3% |
| OSWorld-2.0 (computer use) | 59.0% | 50.6% | 75.4% | 62.6% |
| GDPVal-AA v2 (knowledge work, Elo) | 1545 | 1482 | 1824 | 1710 |
Vals Finance Agent v2 asks the model to do the work of a financial analyst, and Harvey's Legal Agent Benchmark asks it to carry a legal workflow through to completion. On those two, 3.8 Flash scores 61.4% and 10.0% against Opus 5's 58.6% and 6.7%. This is the first time a Flash-tier model has posted higher numbers than a frontier model on structured professional work.
The gains reach outside coding. HLE-Verified, an expert reasoning exam spanning many fields, moves from 53.6% to 54.9%, narrowly above Opus 5 at 54.4% and GPT-5.6 Sol at 54.5%. Long-video understanding on LVBench goes from 85.4% to 87.8%, and LABBench2, which uses real biology research tasks, goes from 82.1% to 86.2%; both lead the table. On the hard split of BioMysteryBench, the problems that are difficult for humans too, the score jumps 13 points from 43.5% to 56.5%, ahead of Opus 5 at 49.4%.
The scatter plot Google published puts DeepSWE score on the vertical axis and average cost per task on the horizontal, with cheaper to the right.

Reading the axes off the chart, 3.8 Flash tops out near 74% at roughly $2.40 per task, while Claude Opus 5 reaches the same score band at roughly $11.80 per task. That is about one-fifth the cost for the same score. Against 3.7 Flash, cost per task rises slightly, from roughly $2.20 to roughly $2.40. These are approximations read off a chart, so the decimals are not load-bearing, but the direction is clear.
Where it does not reach Opus 5
On Terminal-bench 4.0, 3.8 Flash scores 19.1%. Opus 5 scores 51.8% and GPT-5.6 Sol 37.3%, a gap of more than double. The two Terminal-bench versions test different things: 2.1 is fixing code at a terminal, while 4.0 is a general task where the model picks its own tools and runs autonomously for long stretches. 3.8 Flash leads the first at 89.4% and sits near the bottom of the second at 19.1%.
Computer use tells the same story. OSWorld-2.0 has the model look at a screen and click and type its way through applications the way a person would. 3.8 Flash improves from 3.7 Flash's 50.6% to 59.0%, still far short of Opus 5 at 75.4%. GDPVal-AA v2, the composite knowledge-work metric, comes in at Elo 1545, below both Opus 5 at 1824 and GPT-5.6 Sol at 1710.
The performance gain has a cost attached. Google writes that 3.8 Flash "works harder" and states plainly that on complex tasks it executes extra reasoning steps and calls tools iteratively, so it may consume more tokens, especially at higher effort levels. The model card carries the same sentence. The announcement then recommends this:
For applications where compute efficiency is the primary constraint, developers can utilize lower effort levels to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads.
It is unusual for a launch post to tell you to keep using the previous model. Identical per-token pricing does not cap the bill if each request burns more tokens. For batch workloads where the per-item cost multiplies straight through, measure token consumption on the same inputs before you switch.
Availability and API changes
Gemini API and Google AI Studio list South Korea among supported countries, and Google AI Pro sells there for 29,000 won per month, so the model is usable today in Korea.
| Item | Detail |
|---|---|
| Who gets it | Developers (Gemini API, AI Studio, Antigravity, Android Studio, Stitch), enterprises (Gemini Enterprise), and consumers (Google AI Pro and Ultra subscribers, via the Gemini app, AI Mode in Google Search, and Google Sheets) |
| Free tier | No charge on input or output, but your input is used to improve Google products |
| Paid tier pricing | Per million tokens: $0.75 input, $3.75 output, $0.075 context caching. Doubles to $1.50, $7.50, and $0.15 on January 1, 2027 |
| Korea | Available. The announcement does not list per-country rollout for the Gemini app |
| Requirements | No application needed for 3.8 Flash. 3.8 Flash Cyber requires Fairwind Program approval |
If you call the API, swapping the model name alone will not work. The thinking_budget parameter used through 3.7 Flash is gone, replaced by thinking_level with three values: low, medium (default), and high. There is no minimal, and passing it raises an error. Sampling parameters such as temperature and top_p have also been removed. The model ID is gemini-3.8-flash, the input window is 1,048,576 tokens, and output is capped at 65,536. The knowledge cutoff is March 2026, with some domains still at January 2025.
The security variant is approval-only
Gemini 3.8 Flash Cyber shipped the same day. It finds vulnerabilities in code and writes the patches to fix them. It is not generally available: access runs through Google's new Fairwind Program, limited to government agencies, critical infrastructure operators, and software maintainers. Academic labs doing defensive benchmarking can also apply.
Only Google's published numbers are reproduced here. On CWE-Bench, a vulnerability patching benchmark, it reaches 47.2% pass@1 against 47.8% for the leading frontier model, at substantially lower cost. On an internal benchmark spanning 20 programming languages, its success rate at finding real vulnerabilities exceeded 70%. The Chrome security team said it produced 2.6 times more correct patches for Chrome vulnerabilities than much larger top commercial models. Wiz reported recall 7.5 to 9.7 points higher on its own penetration testing benchmark at 2.3 to 5.2 times lower cost.
The same week, Anthropic shipped Claude Fable 5.1 and Mythos 5.1 and OpenAI shipped Astra, each gating its security-tuned model behind approval. All three cite the same reason: a model that gets better at finding vulnerabilities gets better at exploiting them. Standard 3.8 Flash is the only one of these an ordinary developer can pick up today.
How to decide whether to switch
Select gemini-3.8-flash in the Google AI Studio free tier and rerun one terminal coding task you currently give 3.7 Flash or Claude Sonnet 5. Free-tier input feeds Google's product improvement, so use a public repository task rather than internal code. Record the token count alongside the result next to 3.7 Flash, and you will also have the arithmetic ready for when the introductory price ends on December 31. If instead you run agents that operate a screen or work autonomously for days, the 19.1% and 59.0% in the table are your answer. That work still belongs to Opus 5.