Change language to
0:00

Google is beginning a phased rollout of Gemini 4 Argon, its new frontier AI model, with trusted cyber defenders in its Fairwind Program getting access first. It can generate up to one million output tokens in a single run, but developers and consumers cannot use it yet.

Google’s announcement lists introductory rates of $2 per million input tokens and $10 per million output tokens. That is more detail than the model’s limited rollout might suggest: there is a price, even if there is no public access date.

Our earlier report on Gemini 4 being nearly ready covered Google DeepMind chief Koray Kavukcuoglu’s comments about the model’s refinement stage. The release now gives that earlier update a name, a benchmark slate and a cyber-first rollout.

Subscribe to our Newsletters for more Business Stories

Gemini 4 Argon begins with a controlled cyber rollout

Google says Argon is available to a set of trusted defenders through Fairwind, while Google’s own teams are also using it. The company says it is taking part in the US government’s voluntary pre-release model access process and will keep adjusting safeguards based on tester feedback before wider availability.

The first external partner named by Google is Wiz, which is using Argon through its Scan for Good initiative. Google says the model found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide; it says earlier frontier models missed the flaw. Google has not publicly identified the affected software or disclosed technical details, so the finding remains the company’s account rather than an independently described case study.

For Fairwind participants and Google’s internal teams, Google says it is releasing Argon without cyber guardrails. That access is intended for defensive work, including finding, validating and patching vulnerabilities. The phased approach keeps the model with selected security teams while Google works on protections before broader access.

Gemini 4 Argon benchmarks show strengths and gaps

Google reports a 77.9% score on DeepSWE v1.1, a test of long-horizon software engineering. But its own comparisons do not show a clean sweep across coding evaluations. The New Stack’s comparison of Google’s published scores shows Argon behind competitors on two other tests:

BenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
DeepSWE v1.177.9%74.1%67.4%74.2%
FrontierSWE v255.0%65.5%56.3%62.3%
Terminal-Bench 4.057.4%58.2%57.9%66.4%

Google also says Argon tied for first at 68% on CWE-bench v1, which tests vulnerability remediation. The numbers are useful for comparing performance on those specific tasks, not a guarantee that one model will be best at every coding or security job. Ars Technica likewise notes that Google’s broader claims are based on its benchmark results.

Google Gemini 4 Argon CWE-bench v1 cybersecurity benchmark chart

The pricing on Google’s announcement is introductory: $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95%. Google has not said how long the introductory period lasts. This gives developers a published token rate to assess, but not a way to try the model today.

What Gemini 4 Argon’s million-token limit changes

The output ceiling rises to one million tokens from 64,000 on earlier Gemini models, according to Google. A larger output allowance can let a model work through longer tasks in one run, rather than splitting every large analysis or code change into chained requests. It does not mean every task needs that much output, or that a longer response is automatically a better one.

Google says Argon is already being used internally for debugging, research and code migrations. One fleet-telemetry project freed more than 300 TiB of memory once rolled out, with Google estimating total savings of 500 TiB to 1 PiB. Its agents are also helping migrate large C and C++ codebases to Rust, including work on more than 800,000 lines of the Fuchsia Zircon kernel. Google says those changes go through automated and manual audits, emulation testing and review before production.

The model is also a different proposition from more narrowly scoped Gemini releases such as Gemini 3.5 Transcribe, which Google positioned around speech transcription. Argon is aimed at longer engineering, enterprise knowledge and defensive-security workflows, with access deliberately limited while it is tested.

Google has not set a public date for general availability or a UAE-specific release. Ars Technica reports that paid API customers and Google AI Ultra subscribers are expected to be first when the model moves beyond cyber testing; Google has not published a timetable. For now, Argon’s name and prices are public, but access remains with selected defenders and Google teams.

Does the one-million-token limit mean every Gemini 4 Argon reply will be that long?

No. One million tokens is the maximum output allowance for a single run, not a default response length. A larger ceiling can support long tasks without splitting every stage into separate requests, but the model does not need to use the full allowance each time.

Why is Gemini 4 Argon going to cybersecurity teams first?

Google says models with these cyber capabilities need a phased release. It is gathering feedback from trusted defenders and working on safeguards before opening access more widely.

Does Gemini 4 Argon lead every coding benchmark?

No. Google’s published comparison puts Argon first on DeepSWE v1.1, but it trails some rivals on FrontierSWE v2 and Terminal-Bench 4.0. Benchmark results vary by task, so the headline score does not establish a universal coding lead.