Gemini 3.8 Flash is Google's fast, low-cost reasoning and coding model, released on September 2, 2026, at $0.75 per million input tokens and $3.75 per million output tokens. It scores about 90% on Terminal-Bench 2.1 and 73.7% on DeepSWE v1.1 — within a fraction of a point of Claude Opus 5 — while costing roughly 6.7 times less. It ships alongside Gemini 3.8 Flash Cyber, a restricted variant for vetted security teams.
Two details should shape how you read every headline about it. First, that price is introductory: it doubles to $1.50/$7.50 on January 1, 2027. Second, on the harder Terminal-Bench 4.0, Gemini 3.8 Flash scores 19.1% against Opus 5's 51.8%. It matches the frontier on short, well-scoped tasks and falls far behind on long, open-ended ones.
This article covers the benchmarks that flatter it and the ones that do not, the real cost once the introductory rate expires, what the Cyber variant is for, and where Flash belongs in a production stack.
Key Takeaways
- Gemini 3.8 Flash launched September 2, 2026 at $0.75/$3.75 per million tokens, rising to $1.50/$7.50 on January 1, 2027.
- It scores 73.7% on DeepSWE v1.1, essentially level with Claude Opus 5 (74.0%), and leads on finance and legal agent benchmarks.
- On Terminal-Bench 4.0 it scores 19.1% versus Opus 5's 51.8% — a large gap on long-horizon agent work.
- It uses roughly 30% more output tokens per task than 3.7 Flash, which erodes part of the price advantage.
- Gemini 3.8 Flash Cyber scores 86.2% on CyberGym and is available only through Google's Fairwind Program.
What is Gemini 3.8 Flash?
Gemini 3.8 Flash is the latest model in Google's Flash tier, the line built for speed and price rather than maximum capability. It has a 1,048,576-token context window, a 65,536-token output limit and three thinking levels that default to medium. Artificial Analysis measures it at about 305 output tokens per second.
It is broadly available. According to Google's announcement, developers get it in Google AI Studio, Android Studio and Google Antigravity; enterprises through Gemini Enterprise; and consumers in the Gemini app, AI Mode in Search and Google Sheets.
The pace is notable on its own. It arrived just weeks after 3.7 Flash, and reviewers counted it as Google's third model release in about six weeks. Google is iterating the cheap tier faster than anyone else in the market, which is a different strategy from the one we covered in Gemini Deep Think and Gemini 3.5 Pro, where the focus was the top of the range.
The output limit is the spec to check against your use case. At 65,536 tokens, it is half of what GPT-6 Sol and Claude Opus 5.5 return (128,000). For chat and tool calls this never matters. For generating a large file or a long report in a single response, it does.
Gemini 3.8 Flash benchmarks: where it wins and where it doesn't
Gemini 3.8 Flash matches far more expensive models on scoped coding and domain-specific agent tasks, and trails them badly on long-horizon general agency. It scores 73.7% on DeepSWE v1.1 against Opus 5's 74.0%, but only 19.1% on Terminal-Bench 4.0 against Opus 5's 51.8%.
Vellum's benchmark breakdown lays the numbers out against the models that were current at its September 2 launch.
| Benchmark | 3.8 Flash | 3.7 Flash | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| DeepSWE v1.1 | 73.7% | 65.3% | 74.0% | 72.7% |
| Terminal-Bench 2.1 | 89.4% | — | 89.1% | 88.8% |
| Vals Finance Agent v2 | 61.4% | 59.0% | 58.6% | 53.8% |
| Harvey Legal Agent (all-pass) | 10.0% | 8.8% | 6.7% | 2.5% |
| BioMysteryBench (hard) | 56.5% | 43.5% | 49.4% | 44.7% |
| LVBench (long video) | 87.8% | — | 75.4% | 82.1% |
| HLE-Verified | 54.9% | 53.6% | 54.4% | 54.5% |
| Terminal-Bench 4.0 | 19.1% | 11.2% | 51.8% | 37.3% |
| OSWorld 2.0 | 59.0% | 50.6% | 75.4% | 62.6% |
| GDPval-AA v2 (Elo) | 1545 | 1482 | 1824 | 1710 |
| Gray Swan injection (lower is better) | 5.5% | 9.2% | 4.8% | 27.0% |
A note on the Terminal-Bench 2.1 figure: Google's launch material cites 90.8%, up from 81.6% for 3.7 Flash, while the comparison table above lists 89.4%. The difference is likely harness or effort settings. Either way it is at or near the top.
The table splits cleanly into two halves, and the split is the most useful thing to know about this model.
Top half — scoped tasks. Give Flash a bounded problem with a clear finish line and it performs like a model that costs several times more. The finance and legal results are genuine leads, not ties. The BioMysteryBench jump from 43.5% to 56.5% in a single release is the largest generational gain on the sheet.
Bottom half — open-ended agency. Terminal-Bench 4.0, OSWorld and GDPval measure long tasks where the model must plan, recover from mistakes and keep going for many steps. Here Flash is not close. The 19.1% versus 51.8% gap on Terminal-Bench 4.0 is the difference between a model that can finish a multi-hour engineering task and one that usually cannot.
Newer releases have widened that gap since launch. Claude Opus 5.5, released September 22, scores 66.4% on Terminal-Bench 4.0.
The unique angle most reviews miss is that both halves are true at once, and which half applies depends on how you decompose work. A harness that breaks a job into small, verifiable steps keeps Flash in the top half of the table. A harness that hands over a vague goal and waits pushes it into the bottom half. Flash rewards orchestration.
One more strong result: prompt-injection resistance. On the Gray Swan benchmark, attacks succeed 5.5% of the time, down from 9.2% for 3.7 Flash and far below GPT-5.6 Sol's 27.0%. For a model that will sit inside browsers and Search, that is the right thing to have improved. We covered why it matters in AI browser agents and agentic browsers.
How much does Gemini 3.8 Flash really cost?
Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, with cached input at $0.075. On January 1, 2027 the price doubles to $1.50 and $7.50. It also uses about 30% more output tokens per task than 3.7 Flash, so the effective cost is higher than the list price suggests.
That is two separate adjustments you need to make before trusting a cost comparison.
| Model | Input / 1M | Output / 1M | Note |
|---|---|---|---|
| Gemini 3.8 Flash (to Dec 31, 2026) | $0.75 | $3.75 | Introductory |
| Gemini 3.8 Flash (from Jan 1, 2027) | $1.50 | $7.50 | Standard |
| GPT-6 Luna | $0.10 | $0.50 | Permanent |
| GPT-6 Sol | $2.00 | $10.00 | Permanent |
| Claude Opus 5.5 | $4.00 | $20.00 | — |
Work through it for a realistic monthly volume:
def monthly(input_tokens, output_tokens, price_in, price_out):
return (input_tokens * price_in + output_tokens * price_out) / 1_000_000
IN, OUT = 500_000_000, 80_000_000 # tokens per month on 3.7 Flash
OUT_38 = int(OUT * 1.30) # ~30% more output tokens per task
intro = monthly(IN, OUT_38, 0.75, 3.75)
standard = monthly(IN, OUT_38, 1.50, 7.50)
sol = monthly(IN, OUT, 2.00, 10.00)
print(f"3.8 Flash, intro rate: ${intro:,.2f}") # $765.00
print(f"3.8 Flash, from Jan 2027: ${standard:,.2f}") # $1,530.00
print(f"GPT-6 Sol: ${sol:,.2f}") # $1,800.00
At the introductory rate, Flash is less than half the cost of GPT-6 Sol. After January 1, with the extra output tokens included, it is about 85% of Sol's cost. The "roughly 6.7 times cheaper than Opus 5" framing in launch coverage is accurate today and will be half as true in three months.
The timing is awkward for Google. Three weeks after this launch, OpenAI released GPT-6 Sol and Luna with prices it explicitly called permanent, and Luna at $0.10/$0.50 undercuts Flash's introductory rate by more than seven times on input. If you are budgeting for 2027, use the $1.50/$7.50 figure. Our LLM API pricing guide tracks the rest of the table.
What is Gemini 3.8 Flash Cyber?
Gemini 3.8 Flash Cyber is a security-specialized variant of the same model, tuned for vulnerability discovery and patching. It is not generally available: access runs through Google's Fairwind Program for governments, critical infrastructure operators and other trusted defenders. Google does not publish a price for it.
Its published results:
- CyberGym: 86.2% pass@1, ahead of GPT-5.5 Cyber (85.6%) and Claude Mythos 5 (83.8%).
- Internal 20-language vulnerability benchmark: 71.0%, up from 58.9% for 3.7.
- CWE-Bench patching: 47.2% pass@1.
Google also cites field results from early users. Chrome Security reported 2.6 times more correct patches than larger commercial models. Wiz reported 7.5 to 9.7% higher recall at 2.3 to 5.2 times lower cost. Google Cloud Vulnerability Research found a critical vulnerability in under two hours.
The strategic point is what kind of model this is. OpenAI's Critical-tier cyber capability lives in GPT-6 Astra at $10/$50. Google put frontier-level vulnerability discovery into its cheap, fast tier. Scanning a large codebase is a volume problem — you want to run thousands of cheap passes, not a few expensive ones. A Flash-class cyber model fits that shape better than a flagship does.
The access model is the same across the industry, though. Google's Fairwind, OpenAI's Daybreak and Anthropic's Cyber Verification Program all gate offensive capability behind vetting. As Vellum puts it, Fairwind is "not a checkout page." If your team needs this, plan for an application, not a credit card.
Should you use Gemini 3.8 Flash?
Use Gemini 3.8 Flash for high-volume, well-scoped work where speed matters: domain agents in finance and legal, long-video analysis, bounded coding tasks and anything user-facing that needs a fast response. Do not use it as the single model behind a long-running autonomous agent, where it scores well below Claude Opus 5.5, GPT-6 Astra and GPT-6 Sol.
| Use case | Verdict |
|---|---|
| Finance / legal document agents | Strong pick — leads its comparison set |
| Long-video understanding | Strong pick — 87.8% on LVBench |
| Scoped coding (single task, clear tests) | Good — near-frontier DeepSWE score |
| Latency-sensitive product features | Good — about 305 tokens/second |
| Multi-hour autonomous engineering | Avoid — 19.1% on Terminal-Bench 4.0 |
| Desktop computer use | Avoid — 59.0% on OSWorld 2.0 |
| Lowest possible cost per token | Consider GPT-6 Luna instead |
In practice the best pattern is a two-model setup. Let a stronger model plan and decompose, then fan the pieces out to Flash. You pay frontier prices for the small share of tokens that need judgment and Flash prices for the rest. If you are comparing tools that do this for you, see our roundup of the best AI coding assistants.
Frequently asked questions
When was Gemini 3.8 Flash released? Gemini 3.8 Flash was released on September 2, 2026. It is generally available through Google AI Studio, Gemini Enterprise, the Gemini app and AI Mode in Search, while the Cyber variant is limited to the Fairwind Program.
How much does Gemini 3.8 Flash cost? Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens until December 31, 2026. From January 1, 2027 the price rises to $1.50 and $7.50. Cached input is $0.075 per million during the introductory period.
Is Gemini 3.8 Flash better than Claude Opus 5? On some tasks, yes. It leads Opus 5 on finance, legal, bioinformatics and long-video benchmarks and nearly ties it on DeepSWE, but trails substantially on Terminal-Bench 4.0 (19.1% vs 51.8%) and OSWorld 2.0 (59.0% vs 75.4%).
What is the context window of Gemini 3.8 Flash? Gemini 3.8 Flash has a context window of 1,048,576 tokens and an output limit of 65,536 tokens. The output limit is about half that of GPT-6 Sol and Claude Opus 5.5.
What is the Fairwind Program? The Fairwind Program is Google's vetted-access scheme for Gemini 3.8 Flash Cyber. It is aimed at governments, critical infrastructure operators and trusted defenders, and access requires an application rather than standard API sign-up.
How fast is Gemini 3.8 Flash? Gemini 3.8 Flash generates about 305 output tokens per second according to Artificial Analysis. That makes it one of the faster reasoning-capable models available for interactive use.
The verdict
Gemini 3.8 Flash is the best cheap model for bounded tasks and a poor choice for unbounded ones. Both statements come from the same benchmark sheet, and any review that quotes only the top half is selling you something.
The pricing deserves the same skepticism. At $0.75/$3.75 it is a bargain; at $1.50/$7.50 with 30% more output tokens, it is merely competitive — and OpenAI's Luna now sits well underneath it. Build your 2027 budget on the standard rate.
The genuinely new idea here is Flash Cyber: frontier vulnerability discovery in a fast, inexpensive model, which is the right shape for scanning at scale. For the wider security picture, read our piece on AI coding agent security.
Fast and cheap used to mean "good enough." With the right harness around it, it now means "as good" — just not for very long stretches on its own.