Gemini 4 Argon: 1M Output Tokens at $2/$10 Introductory Rate
In this article
Google DeepMind published Gemini 4 Argon on September 30, 2026, positioning it as a frontier reasoning model for sustained, long-horizon workflows in software engineering, enterprise knowledge work, and cybersecurity. The release is staged: initial access is restricted to trusted cyber defenders through Google's Fairwind Program, with broader availability to paid API customers and Google AI Ultra subscribers to follow after a feedback period with early testers.
Output Scale and Benchmark Performance
The architectural decision engineers should focus on is the output token limit. Google DeepMind reports an expansion from 64K to 1 million output tokens — a 15× increase — designed to let the model "think deeply and generate hundreds of thousands of tokens in a single trajectory." That headroom enables long-horizon tasks without chained calls or external state management.
On vendor-reported benchmarks, Google DeepMind claims the following scores for Argon: DeepSWE v1.1 (77.9%), measuring real-world long-horizon software engineering; AutomationBench from Zapier, measuring end-to-end execution across core business functions (51.3%, ranked first); LVBench for long video understanding (91.7%, described as state of the art); and CWE-bench v1 for security vulnerability remediation (68%, tied for first). Google DeepMind also reports top placement on the Vals Index, which weights finance, coding, legal, and tax performance by each sector's contribution to U.S. GDP.
On internal deployments, Google DeepMind reports that Argon agents took an existing Rust port of libgav1 — Google's open-source video decoder — and replaced 32K lines of SIMD code with safe Rust that the compiler vectorizes automatically, achieving a 2.7× throughput improvement with identical video output. Separately, Argon agents autonomously analyzed fleet-wide profiling telemetry and freed over 300 TiB of memory across Google data centers, with an estimated 500 TiB to 1 PiB in total projected savings. Argon agents are also working on codebase migrations scaling up to 800K+ lines for projects such as the Fuchsia Zircon kernel.
Pricing and Competitive Position
Google DeepMind states an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off the input price. After the introductory period, pricing moves to $4 per million input tokens and $20 per million output tokens. No end date for the introductory period is disclosed.
| Model | Provider | Input ($/1M tokens) | Output ($/1M tokens) |
|---|---|---|---|
| Gemini 4 Argon (intro) | $2.00 | $10.00 | |
| Gemini 4 Argon (standard) | $4.00 | $20.00 | |
| GPT-5.6 Sol | OpenAI | $4.00 | $20.00 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 |
| GPT-6 Astra | OpenAI | $10.00 | $50.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 |
Competitor prices from AI Mastery's pricing data (as of 2026-09-14). Gemini 4 Argon prices per Google DeepMind.
At its introductory rate, Argon's output pricing matches Claude Sonnet 5 and undercuts Claude Fable 5 by 5×. Its standard output price lands identically to GPT-5.6 Sol. The 95% cache discount is significant for token-efficient pipeline architectures where repeated context dominates cost.
Cybersecurity Access and Safety Architecture
The Fairwind Program gating is the most consequential deployment decision in this release. Google DeepMind is releasing Argon to trusted cyber defenders without standard cyber guardrails, granting full frontier-level cybersecurity capability to that cohort. Wiz is named as an early participant through its Scan for Good initiative; Google DeepMind reports the model identified a critical vulnerability in healthcare software used by hospitals worldwide that prior frontier models had missed.
On safety, Google DeepMind describes four active mitigation areas: misuse prevention including monitoring of internal model activations; indirect prompt injection robustness, with Google claiming leadership on Gray Swan's IPI benchmark; chain-of-thought misalignment monitoring with a dedicated incident response team; and hardened sandboxed evaluation environments. Google DeepMind explicitly states it avoided feeding misalignment monitoring findings back into training to prevent Argon's reasoning from learning to evade detection — a detail relevant to ongoing debates about gated capability verification.
AI Mastery Analysis
The 1M output token limit is the architectural decision that most changes agent system design. Most current multi-agent frameworks exist partly because output token limits force task decomposition — not because that architecture is optimal, but because no single call can hold a full trajectory. At 1M tokens, some of that decomposition pressure disappears, cutting orchestration complexity but concentrating latency and cost risk into a single call. Engineers planning agentic pipelines around Argon should verify that their timeout handling, streaming infrastructure, and error recovery logic is built for calls at this scale — the failure modes described in production AI architecture become more acute, not less, when a single failed call represents a full work unit.
The tiered pricing structure — introductory at $2/$10 flipping to $4/$20 — creates a planning risk for teams building cost models now. Any production system that prices out economics at the lower rate carries an undated repricing exposure. The 95% cache discount partially offsets this for read-heavy workflows, but only where input context is genuinely reusable. The cybersecurity access model raises a separate concern: releasing without cyber guardrails to a vetted cohort is a defensible posture, but the verification boundary — who qualifies as a trusted defender and how that is audited — is left unspecified in Google DeepMind's announcement, which is precisely the infrastructure governance gap that becomes consequential at deployment scale.
Argon's internal deployments — 800K+ line kernel migrations, fleet-wide memory profiling, multi-round compiler optimisation experiments — are structured arguments that long-trajectory single-call generation is becoming the primitive around which serious engineering workflows are being rebuilt. General API access remains pending, and the model's most capable configurations will not be reachable by most developers in the near term.
Sources
- Primary: Google DeepMind: Gemini 4 Argon: our next era of frontier intelligence
- Competitor prices: AI Mastery token calculator, figures as of 2026-09-14
Frequently asked questions
How much does Gemini 4 Argon cost per million tokens?
At launch, Google DeepMind prices Gemini 4 Argon at $2 per million input tokens and $10 per million output tokens, with cached input tokens at 95% off the input price. After the introductory period, pricing moves to $4 per million input tokens and $20 per million output tokens, though no end date for the introductory rate is disclosed.
What is Gemini 4 Argon's output token limit?
Google DeepMind reports an output token limit of 1 million tokens, up from the previous 64K — a 15× increase. The company states this allows the model to think deeply and generate hundreds of thousands of tokens in a single trajectory.
What benchmark scores does Gemini 4 Argon achieve?
According to Google DeepMind, Argon scores 77.9% on DeepSWE v1.1 (real-world software engineering), 51.3% on AutomationBench ranking first, 91.7% on LVBench for long video understanding, and ties for first on CWE-bench v1 with 68% for security vulnerability remediation.
How does Gemini 4 Argon handle cybersecurity access?
Google DeepMind is initially releasing Argon through its Fairwind Program to trusted cyber defenders without standard cyber guardrails, granting full frontier-level cybersecurity capability to that cohort. Wiz is named as an early participant; broader availability to paid API customers and Google AI Ultra subscribers will follow after a feedback period.
What internal performance improvements did Argon agents achieve at Google?
Google DeepMind reports that Argon agents replaced 32K lines of SIMD code in an existing Rust port of libgav1 with safe Rust, achieving a 2.7× throughput improvement with identical video output. Separately, Argon agents freed over 300 TiB of memory across Google data centers, with an estimated 500 TiB to 1 PiB in total projected savings.
Related Reading
GPT-6.1 Sol Matches Astra on Coding at One-Fifth the Price
OpenAI's GPT-6.1 Sol hits $2/$10 per million tokens—one-fifth of Astra's rate—while matching it on DeepSWE and closing to within 2.1 points on computer use.
Claude Sonnet 5.5 Beats Opus 5.5 on Agentic Coding at Half the Tier Price
Anthropic's Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, outpacing Opus 5.5's 66.4%, at the same $2/$10 per-million-token price as Sonnet 5.
Meta's Muse Zero-Day Let Any Local App Seize Full Account Control
A zero-day in Meta's Muse macOS agent let any unprivileged local app steal the auth token granting full account control. Meta hotfixed it 12+ hours after disclosure.