Google released its next-generation frontier model, Gemini 4 Argon, on September 30, with first access limited to cybersecurity defense teams enrolled in its Fairwind program. Google has mapped out three use cases for the model: software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. Paid API customers and Google AI Ultra subscribers are next in line, with no timeline given for a broader public rollout.
Giving security teams first access ties directly to how the model was trained. Google says Argon received dedicated training for cyber defense and can autonomously discover, verify, and patch critical software vulnerabilities. Trusted defenders inside Fairwind and Google's own internal teams get Argon without cybersecurity guardrails, while everyone else receives a version with anti-abuse protections, prompt-injection defenses, and alignment monitoring built in.
The official scorecard
Google published several benchmark results:
| Benchmark | Focus | Argon score |
|---|---|---|
| DeepSWE v1.1 | Real-world software engineering | 77.9% |
| CWE-bench v1 | Vulnerability remediation | 68% (tied for first) |
| AutomationBench | Automated tasks | 51.3% (first place) |
| LVBench | Long-video understanding | 91.7% |
Two more results were given only as rankings: Google says Argon leads the Vals Index, which spans finance, coding, law, and tax work, and that it posted the best resistance in Gray Swan's indirect prompt-injection tests. All of these figures are vendor-reported, and only one independent retest was available on launch day.
Output length is another notable change. The previous-generation Gemini capped single-response output at 64,000 tokens; Argon raises that to 1 million tokens, matching its 1-million-token context window.
Pricing and independent testing
During the launch window, the API is priced at a discount: $2 per million input tokens and $10 per million output tokens, with cache hits cutting the input price by another 95%. Once the discount ends, pricing reverts to $4 input and $20 output per million tokens; Google has not said how long the discount period will last.
Independent evaluator Artificial Analysis published its own results the same day. Argon scored 53 on its Intelligence Index, tying GPT-6 Astra's top reasoning tier, edging out GPT-6.1 Sol's top tier by 1 point, and beating the previous-generation Gemini 3.1 Pro Preview's score of 30 by 23 points.
Comparing per-task cost on the same index:
- At the discounted rate, Argon costs about $1.99 per task, versus $3.26 for GPT-6 Astra;
- Once the discount ends, Argon's cost rises to about $3.98, roughly 20% more than Astra;
- GPT-6.1 Sol costs about $0.72 per task while scoring just 1 point lower — making Argon's discounted price about 2.8 times higher.
The gap mainly comes down to token usage. Artificial Analysis found Argon averages about 62,000 output tokens per task, versus roughly 27,000 for GPT-6 Astra — more than double. With the per-token price cut in half but output doubled, the discount's price advantage is largely offset at full price. Argon supports a mechanism Google calls "extended decode continuation," letting its reasoning process write up to 1 million output tokens, which explains its higher token usage on the index test.
Who gets the guardrails
In the same week, OpenAI canceled the release of GPT-6.1 Astra, citing safety standards it hadn't met, and shipped the cheaper Sol instead. Google's approach is to put its strongest version into a controlled list first. The two companies are handling things differently, but facing the same underlying tension: the better a model gets at finding vulnerabilities and writing patches, the greater the risk it could be turned around and used for attacks.
Google's release blog doesn't say which organizations are on the Fairwind list, what criteria are used to vet them, or how the guardrail-free version is kept from leaking out. For developers in mainland China, Argon remains largely out of reach for now: Fairwind is invite-only, there's no timeline for opening up Ultra subscriptions or paid API access, and the Gemini API itself isn't available in mainland China.
Sources: Google's official blog, Artificial Analysis evaluation report, CocoLoop, VentureBeat; benchmark scores are as self-reported by Google, while per-task cost and token usage figures follow Artificial Analysis's Intelligence Index methodology.