Google Unveils Gemini 4 Argon, Claiming Benchmark Leads Over GPT-6 Astra and Claude Opus 5.5 but Starting With a Limited Rollout
Google announced Gemini 4 Argon, which it says leads on 12 of 18 disclosed benchmarks. Access begins with trusted cyber defenders, with introductory pricing of $2 and $10 per million tokens.
Overview
Google on September 30 announced Gemini 4 Argon, which VentureBeat describes as a new frontier AI model that the company says leads rival systems on several enterprise-relevant benchmarks, including long-horizon software engineering, cybersecurity vulnerability remediation, business automation and economic-impact knowledge work. The model is not yet broadly available: Google says it is rolling out first to a set of trusted cyber defenders through its Fairwind Program.
What We Know
Benchmarks (Google-disclosed)
Across the 18 benchmarks Google disclosed, Argon leads outright on 12 and ties for first on one, while GPT-6 Astra leads outright on three and ties Argon on one, and Claude Opus 5.5 leads outright on two, according to VentureBeat. These are figures from Google’s own comparison set, not independent evaluations.
Per the same VentureBeat report, Argon scores 77.9% on DeepSWE v1.1, a long-horizon software engineering benchmark, against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. On Zapier’s AutomationBench it scores 51.3%, compared with 42.5% for Claude Opus 5.5 and 41.4% for GPT-6 Astra. On GraphWalks at longer contexts it scores 84.2%, against 71.8% for GPT-6 Astra and 66.8% for Claude Opus 5.5. On Harvey’s Legal Agent Benchmark it scores 19.6%, against 5.4% for GPT-6 Astra and 3.8% for Claude Opus 5.5. VentureBeat also reports that Argon ties GPT-6 Astra at 68% on CWE-bench v1, a vulnerability-remediation benchmark, with Claude Opus 5.5 at 67%.
The results are not a sweep. VentureBeat reports that GPT-6 Astra leads by 10.5 percentage points on FrontierSWE v2 (65.5% to 55.0%) and on Terminal-Bench Science 0.1 (68.1% to 57.6%), and that Claude Opus 5.5 leads on Terminal-bench 4.0, scoring 66.4% to Argon’s 57.4%.
On the Gray Swan indirect prompt-injection benchmark, VentureBeat reports Argon posts a 0.7% attack success rate, compared with 1.0% for Claude Opus 5.5 and 8.5% for GPT-6 Astra.
Capabilities and internal use
Google says it is expanding the model’s output token limit to 1 million tokens, up from the previous 64,000, according to VentureBeat; Google’s blog post calls this an industry-leading limit.
As an example of internal use, Google says Argon agents took an existing Rust port of its libgav1 video decoder and replaced 32,000 lines of SIMD code, producing a memory-safe decoder that runs 2.7 times faster than the Rust port, per VentureBeat. Both VentureBeat and Google’s blog say Wiz is already using Argon through its Scan for Good initiative.
Access and pricing
Google says it is participating in the U.S. government’s voluntary process for pre-release model access while it gradually expands access. According to VentureBeat, broad availability is planned “as soon as possible” for developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers. For trusted defenders and Google’s own internal teams, the company says it will release Argon without cyber guardrails, VentureBeat reports.
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off the input price, per Google. After the introductory period the price becomes $4 and $20. VentureBeat notes that Google has not said how long the introductory period will last, and lists GPT-6 Astra at $10 input and $50 output per million tokens, and Claude Opus 5.5 at $4 and $20.
What We Don’t Know
- How long the introductory pricing will run, which VentureBeat says Google has not specified.
- When general access begins. Google has said only “as soon as possible”, and VentureBeat notes most customers will need to wait for broader API access before they can test whether the benchmark advantages translate into lower production costs.
- Whether the benchmark results hold under independent testing; the comparisons come from Google’s disclosed materials.
Context
VentureBeat reports that Google had not released a new flagship since the Gemini 3 series in November 2025.