Gemini 4 Argon is Google DeepMind's new frontier model, announced on September 30, 2026. It can write up to 1 million tokens in a single response, costs $2 per million input tokens and $10 per million output during an introductory period, and posts the lowest hallucination rate of any model at its intelligence level. The catch: almost nobody can use it yet. Google is releasing it first, without cyber guardrails, to vetted security teams in its Fairwind Program.
That combination is unusual. Most frontier launches lead with a benchmark chart and an API page. Argon leads with an access list. Google trained it to find, validate and patch critical vulnerabilities on its own, decided that capability was too sharp to hand out on day one, and gave it to defenders first.
This guide covers what Argon actually is, the benchmarks it wins and the ones it loses, what a 1M-token output limit costs in practice, how its access gate compares with OpenAI's and Anthropic's, and whether you should plan around it.
Key Takeaways
- Gemini 4 Argon was announced September 30, 2026 and is rolling out first to cyber defenders in Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers next.
- Introductory pricing is $2/$10 per million input/output tokens, rising to $4/$20; cached input is 95% off the input price.
- Output is capped at 1 million tokens per response, up from 64,000 on earlier Gemini models.
- Argon leads DeepSWE v1.1 (77.9%) and AutomationBench (51.3%) but trails Claude Opus 5.5 on Terminal-Bench 4.0 (57.4% vs 66.4%) and GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%).
- Its 15% hallucination rate on Artificial Analysis' benchmark compares with 51% for GPT-6 Astra.
What is Gemini 4 Argon?
Gemini 4 Argon is the first model of Google's Gemini 4 generation, built for long, multi-step work: large software projects, professional knowledge work and cybersecurity defense. Its headline feature is a 1-million-token output limit, and Google is releasing it to vetted cyber defenders before anyone else.
According to Google's announcement, Argon is meant to "sustain deep reasoning" across long-horizon workflows rather than answer quick questions. The examples Google chose say a lot about the intended use. Internally, teams used it to:
- migrate C and C++ codebases of up to 800,000+ lines to Rust;
- free 300 TiB of memory across data centers, with 500 TiB to 1 PiB projected;
- make the libgav1 video decoder 2.7x faster with identical output;
- improve a quantum algorithm optimization by 40% over the baseline.
None of those are chat tasks. They are the kind of job you used to hand to a team for a quarter. That framing matters when you read the benchmarks, because Argon is tuned for duration and reliability more than for raw speed on short tasks.
Gemini 4 Argon benchmarks: where it leads and where it doesn't
Gemini 4 Argon leads on long software tasks, automation, long-video understanding and factual reliability. It trails Claude Opus 5.5 on terminal-based agentic coding and GPT-6 Astra on the hardest software-engineering benchmark. It is not a clean sweep, and Google's own chart does not claim one.
| Benchmark | Gemini 4 Argon | Claude Opus 5.5 | GPT-6 Astra |
|---|---|---|---|
| DeepSWE v1.1 (long software tasks) | 77.9% | 74.2% | 74.1% |
| AutomationBench | 51.3% | 40.0% | 41.4% |
| Terminal-Bench 4.0 | 57.4% | 66.4% | 57.9% |
| FrontierSWE v2 | 55.0% | — | 65.5% |
| CWE-bench v1 (vulnerability work) | 68% (tied) | — | 68% (tied) |
| Artificial Analysis Intelligence Index | 53 | — | 53 |
| AA-Omniscience hallucination rate (lower is better) | 15% | — | 51% |
| LVBench (long video) | 91.7% | — | — |
Sources: Google, DevX, Anthropic and OpenAI launch figures. A dash means no comparable published score.
Three readings stand out.
The hallucination number is the real story. On Artificial Analysis' Omniscience test, Argon hallucinates 15% of the time. GPT-6 Astra sits at 51% and GPT-6.1 Sol at 54%. Raw intelligence scores are now tied at the top (Argon and Astra both land at 53), so the differentiator is how often a model confidently makes things up. For a model meant to run unattended for hours, that is the number that decides whether you trust the output.
Terminal work is still Anthropic's. If your agents live in a shell, running builds, reading logs and fixing tests, Claude Opus 5.5's 66.4% on Terminal-Bench 4.0 is nine points ahead. We covered why that benchmark matters in our Claude Opus 5.5 breakdown.
The security tie is with the models that are also gated. Argon's 68% on CWE-bench v1 ties GPT-6 Astra and xAI's Grok 4.7, per SecurityWeek. The leaderboard for vulnerability work is now a three-way tie between models that their makers all restrict.
How much does Gemini 4 Argon cost?
Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens during its introductory period, then $4 and $20. Cached input is discounted 95% off the input price. That makes it far cheaper than GPT-6 Astra and the same list price as Claude Opus 5.5 once the promotion ends.
| Model | Input ($/M) | Output ($/M) | Max output per response |
|---|---|---|---|
| Gemini 4 Argon (intro) | $2 | $10 | 1,000,000 |
| Gemini 4 Argon (standard) | $4 | $20 | 1,000,000 |
| Claude Opus 5.5 | $4 | $20 | 128,000 |
| GPT-6 Astra | $10 | $50 | — |
| GPT-6 Sol | $2 | $10 | — |
The comparison that matters is not the rate card. It is what the 1M output limit lets a single request cost. Here is the arithmetic for the most extreme case, one response that uses the whole output budget:
def response_cost(output_tokens, price_per_million):
return output_tokens / 1_000_000 * price_per_million
max_output = 1_000_000
print(response_cost(max_output, 10)) # Argon, intro pricing -> 10.0
print(response_cost(max_output, 20)) # Argon, standard -> 20.0
print(response_cost(128_000, 20)) # Opus 5.5 at its cap -> 2.56
A single maxed-out Argon response costs $10 today and $20 later. That is not a reason to avoid it. A long refactor delivered in one coherent pass can be cheaper than twenty shorter calls that each re-read the codebase. But it does change how you budget. With a 64K cap, a runaway response cost cents. With a 1M cap, a misconfigured agent loop can spend real money in minutes. Set an explicit output ceiling per call, and track cost per completed task rather than per request.
Who can use Gemini 4 Argon right now?
Right now, only teams in Google's Fairwind Program and Google's own internal teams can use Gemini 4 Argon. Google says paid API customers and Google AI Ultra subscribers are next, followed by broader developer and enterprise access, but it has not given a date.
Fairwind launched in early September as an application-only program for governments, Google Cloud customers and security partners, and it had more than 650 participants at launch, according to BetaNews. It started with Gemini 3.8 Flash Cyber paired with CodeMender, Google's harness that finds, verifies and fixes vulnerabilities. We covered that first wave in our Gemini 3.8 Flash review.
What changes with Argon is the phrase Google used: trusted defenders get it "without cyber guardrails." In practice, that means the model will carry out vulnerability discovery and validation work that a public model would refuse. Wiz, which Google acquired, is already using it in its Scan for Good effort for critical infrastructure, and Google says Argon found a critical vulnerability in hospital software that earlier frontier models had missed.
The broader release keeps safeguards. Google lists refusal training for cyberattack and CBRN requests, monitoring of the model's internal activations for misuse, chain-of-thought and action monitoring for misalignment, and adversarial training that makes it its "most resilient model yet" against indirect prompt injection.
Fairwind vs Daybreak vs Cyber Verification: three gates, one pattern
All three frontier labs have now reached the same conclusion: build the offensive-grade security capability, then restrict who can invoke it. The mechanics differ in ways that matter if your team does security work.
| OpenAI | Anthropic | ||
|---|---|---|---|
| Program | Fairwind | Daybreak | Cyber Verification Program |
| Gated model | Gemini 4 Argon, Gemini 3.8 Flash Cyber | GPT-6 Astra (full capability) | Claude Opus 5.5 on cyber tasks |
| What the public gets | Argon later, with guardrails | Astra with code review and patching, no exploit PoCs | Cyber tasks routed to Opus 4.8 |
| Entry | Application; governments, Cloud customers, partners | Application; vetted defenders | Application; verified security teams |
Google's version is the most aggressive in one direction: it shipped its newest flagship to defenders before the public at all. OpenAI and Anthropic released their models broadly and fenced off the dangerous capability. Google fenced off the whole model.
Our read is that Google's ordering is the more honest one. A model that can autonomously find and patch critical bugs can also find them for someone who doesn't patch. Giving defenders a head start means the most valuable early findings go into fixes rather than exploits. The cost is that ordinary developers wait, and we still do not know how much the public version's guardrails will narrow what Argon can do.
If your team needs this class of capability, the lesson from GPT-6 Astra applies here too: plan for an application process, not a credit card. And whatever model you run, the containment rules from our piece on AI coding agent security apply even more to a model built to find holes.
Is Gemini 4 Argon better than Claude Opus 5.5 or GPT-6 Astra?
Gemini 4 Argon is better for long autonomous jobs where reliability matters most, Claude Opus 5.5 is better for terminal-driven agentic coding, and GPT-6 Astra is better for the hardest software-engineering and math tasks. None of them wins everything, so the right choice depends on the shape of your workload.
| Workload | Best pick today | Why |
|---|---|---|
| Multi-hour refactors and migrations | Gemini 4 Argon (once available) | Leads DeepSWE v1.1, 1M output, lowest hallucination rate |
| Shell-based coding agents | Claude Opus 5.5 | 66.4% on Terminal-Bench 4.0, available now at $4/$20 |
| Hardest SWE and math problems | GPT-6 Astra | 65.5% on FrontierSWE v2, but $10/$50 |
| Business process automation | Gemini 4 Argon | 51.3% on AutomationBench, 10 points clear |
| Everyday coding on a budget | GPT-6 Sol | $2/$10 and available today |
| Vulnerability research | Whichever gate you can get through | All three tie on CWE-bench v1 |
There is a practical point that the benchmark tables hide: you cannot ship on a model you cannot call. Until Argon reaches the paid API, it is a planning input, not a production option. If you are choosing a model this month, pick between what is available, Opus 5.5 and GPT-6 Sol for most teams, and lean on GPT-6 Sol when cost matters more than the last few benchmark points.
What is worth doing now is preparing for the output limit. Pipelines built around 64K or 128K responses usually chunk long jobs into many calls and stitch the results together. A 1M-token response removes much of that scaffolding, but only if your tooling can stream, store and validate output that large. Check your timeouts and your logging before the model arrives, not after.
Frequently Asked Questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind's frontier model announced on September 30, 2026, built for long, multi-step work in software engineering, knowledge work and cybersecurity defense. It can produce up to 1 million output tokens per response and is rolling out first to vetted cyber defenders.
When will Gemini 4 Argon be available to developers?
Google has not announced a date. Argon is available now only through the Fairwind Program and inside Google; paid API customers and Google AI Ultra subscribers are next in line, followed by broader developer and enterprise access.
How much does Gemini 4 Argon cost?
Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens at introductory pricing. Standard pricing is $4 and $20, and cached input tokens get a 95% discount off the input price.
What is the Google Fairwind Program?
Fairwind is Google's application-only early-access program for vetted cybersecurity teams, including governments, Google Cloud customers and security partners. It launched in early September 2026 with more than 650 partners and gives members Argon without the cyber guardrails the public version will have.
Is Gemini 4 Argon better than Claude Opus 5.5?
Not across the board. Argon leads on long software tasks (77.9% vs 74.2% on DeepSWE v1.1) and automation, while Claude Opus 5.5 leads on terminal-based agentic coding (66.4% vs 57.4% on Terminal-Bench 4.0) and is available today.
What does a 1 million token output limit mean?
It means a single Argon response can be up to 1 million tokens long, roughly 15 times the 64,000-token cap on earlier Gemini models. That allows very large outputs, such as an entire migrated codebase, in one pass, at up to $10 per maxed-out response at introductory pricing.
Conclusion
Gemini 4 Argon is Google's strongest argument yet that it belongs at the top of the frontier. It wins long-horizon software work, automation and, most importantly, reliability, with a hallucination rate less than a third of GPT-6 Astra's. It is also priced to undercut Astra badly, at least during the introductory period.
Our verdict: Argon is the model to plan around for long, unattended jobs, and the wrong model to plan a launch on this month, because you probably cannot call it yet. Use Claude Opus 5.5 or GPT-6 Sol for what you ship now, and get your pipelines ready for 1M-token responses so you can switch the day the API opens. For the cheaper, available half of Google's lineup, start with our Gemini 3.8 Flash review.
The frontier's newest trick is not a higher score. It is deciding who gets the model first.