Devlery
Blog/Google

Gemini 4 Argon Leads 13 of 19 Benchmarks, and Only Cyber Defenders Can Call It

Google announced Gemini 4 Argon on September 30. Max output jumps from 64K to 1M tokens and 13 of the 19 rows in its own comparison table are outright wins, but there is no model ID, no launch date, and no public API at $2 per million input tokens.

Gemini 4 Argon Leads 13 of 19 Benchmarks, and Only Cyber Defenders Can Call It
AI 요약
  • Google published Gemini 4 Argon on September 30, leading 13 of the 19 rows in its own table.
  • Max output jumps from 64K to 1M tokens, but Terminal-bench 4.0 puts it last of four.
  • Input is $2 per million tokens, yet there is no model ID and only cyber defenders have access.

Google DeepMind announced Gemini 4 Argon, a new frontier model, on September 30, 2026. The post carries a 19 row benchmark table and a full price sheet, and nowhere to call the model. No API model ID, no general availability date. The only parties running Argon today are Google's own internal teams and cyber defense organizations inside the Fairwind Program, Google's track for handing new models to defenders ahead of public release.

Google names three jobs for Argon: software engineering against real repositories, enterprise knowledge work in domains such as finance and law, and cyber defense that finds and patches vulnerabilities. The announcement comes from Koray Kavukcuoglu, SVP at Google DeepMind and Google's Chief AI Architect, and states that releasing frontier-level capability safely requires a staged approach, noting Google's participation in the US government's voluntary pre-deployment model access process.

What changed from the previous generation: 64K to 1M output tokens

The headline number is how much the model can emit in a single response. Gemini models capped max output at 64,000 tokens; Argon allows 1 million. That is 15.6x. Google's argument is that a model no longer has to truncate and continue in a following turn, so a large refactor or a long report finishes in one pass. For scale, Claude Opus 5.5 and GPT-6.1 Sol, both shipped the same month, cap max output at 128,000 tokens each.

The input context window, on the other hand, appears nowhere in the announcement. Google did not publish it.

The benchmarks arrived as a single 19 row table, comparing Argon against GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5.

The 19 row Gemini 4 Argon benchmark comparison table, scoring Argon alongside GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5

Tallied up: 13 outright wins, 1 tie for first, 5 rows behind. The widest margins sit in knowledge work. On AutomationBench, Zapier's workflow automation benchmark, Argon scores 51.3% against 42.5% for second-place Opus 5.5, a gap of 8.8 points. On Harvey's Legal Agent Benchmark, which measures legal research and drafting, Argon posts 19.6% while Astra manages 5.4% and Opus 5.5 manages 3.8%, a difference of an order of magnitude. On long context, the GraphWalks band running from 256K to 1M tokens puts Argon at 84.2% against Astra's 71.8%, ahead by 12.4 points.

Two rows that matter if you run a coding agent

Where Argon falls behind matters more for coding work than where it wins. The row Google highlights as a new state of the art, DeepSWE v1.1 at 77.9%, and the two rows where Argon finishes last are all agentic coding benchmarks.

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
DeepSWE v1.177.9%74.1%67.4%74.2%
Vibe Code Bench91.9%89.6%90.3%90.3%
FrontierSWE v255.0%65.5%56.3%62.3%
Terminal-bench 4.057.4%58.2%57.9%66.4%

On FrontierSWE v2, Argon's 55.0% is the lowest of the four, 10.5 points behind Astra at 65.5%. On Terminal-bench 4.0, which runs shell commands directly to complete a task, Argon again finishes last at 57.4%, 9 points under Opus 5.5 at 66.4%. Terminal-Bench Science 0.1, on the science and math side, lands at 57.6% against Astra's 68.1%.

The pattern: Argon leads on long tasks that read documents and reason over them, and trails on tasks that run a shell and check the result. If your coding agent lives in CI or a terminal, the "new DeepSWE record" line in the announcement does not carry over to your workload. Every number here is also Google's own measurement, with no independent reproduction yet. Only an evals methodology URL sits under the table.

CWE-bench 68% is a three-way tie, and every model used a different harness

Cyber defense is Argon's flagship use case. Google says the model was trained to find vulnerabilities on its own, confirm they are real, and produce a patch. Cloud security vendor Wiz is running Argon in Scan for Good, its free public-interest scanning program, and found a critical vulnerability exposing personal data in medical software used by hospitals worldwide. The announcement notes that previous frontier models missed it.

Argon scores 68% on CWE-bench v1, which measures vulnerability remediation. The announcement calls this a tie for first and leaves it there. The leaderboard chart published alongside it shows Grok 4.7 at 68% and GPT-6 Astra at 68% as well, so three models are level. Opus 5.5 sits one point back at 67%.

CWE-bench v1 leaderboard. Grok 4.7, Gemini 4 Argon, and GPT-6 Astra are tied at 68%, with Claude Opus 5.5 at 67%

The fine print under that chart says each model ran inside a different coding tool. The scaffolding that actually drives a model through a task is called a harness, and here Argon ran on Antigravity, Grok 4.7 on opencode, Astra on Codex, and Opus 5.5 and Fable 5.1 on Claude Code. That the same model shifts score when the harness changes was already demonstrated when GPT-6 Astra's ARC-AGI-3 result was re-measured under neutral conditions. Nothing in this leaderboard separates a one-point model difference from a one-point tooling difference. Resistance to indirect prompt injection gets the same treatment: the announcement claims a lead on Gray Swan's IPI benchmark and publishes no score.

Can you use it: there is a price, there is no model ID

ItemDetail
Who has accessCyber defense organizations in the Fairwind Program, plus Google internal teams. Developers, enterprises, and consumers get it "as soon as possible"
PricePromotional $2 per million input tokens, $10 output, cached input discounted 95% off the input rate ($0.10). After the promotion, $4 input and $20 output
Promotion end dateNot published
Regional availabilityNot available anywhere yet, Singapore and the rest of APAC included. At general availability, paid API customers and Google AI Ultra subscribers go first
RequirementsFairwind application portal, background vetting, multi-factor authentication, access logging, and use restricted to in-house security teams. Model ID and release date unpublished

Fairwind prioritizes national cyber authorities, critical infrastructure operators in healthcare, telecommunications, energy, and finance, and platform companies maintaining foundational software used by millions. More than 650 organizations participate, and only a subset of them received Argon. The public Fairwind page states no country restrictions, so nothing on record excludes an APAC applicant that meets the criteria. Nothing on record admits one either, because Google has not said which of the 650 got the model.

Google writes that it is shipping Argon without cyber guardrails to trusted defenders and its own internal teams. Finding a vulnerability overlaps with what an attacker does, so safety filters block the work itself. Google's trade is a guardrail-free model in exchange for access controls imposed on the receiving organization. The public release of Argon waits on four things being tightened: refusal of harmful requests, defense against indirect prompt injection, chain-of-thought monitoring, and a sealed sandbox.

The price sheet is early to take at face value too. The promotional $2 and $10 exactly match GPT-6.1 Sol, announced one day earlier at OpenAI DevDay on September 29, down to the $0.10 cached input rate. The post-promotion $4 and $20 match Opus 5.5, which dropped to that price on September 22. The difference is that Sol and Opus 5.5 have model IDs and answer calls today. They are gpt-6.1-sol and claude-opus-5-5.

Nothing here justifies holding a model decision this quarter for Argon. Run your own task through GPT-6.1 Sol, which costs the same $2 and is callable now, and revisit Argon when a model ID and a context window size appear in the Vertex AI or Gemini API docs. The number to compare at that point is not the 77.9% DeepSWE figure the announcement leads with, but the two closest to a real working environment: 57.4% on Terminal-bench 4.0 and 55.0% on FrontierSWE v2.