Gemini 4 Argon: 1M output tokens at $2/$10 — but gated at launch
Google's frontier model 4x's max output and leads DeepSWE and AutomationBench, but ships to cyber defenders first. Plus Magnitude runs local models free.

Copy markdown
1M output tokens in a single call
Argon returns up to 1M output tokens per response — roughly 4x the ~128K ceiling on rival frontier models. Enough to emit a fully refactored module, a long migration diff, or a book-length structured report in one turn, without chunking and stitching.
Intro price $2/$10, cache at $0.10 — then it doubles
Introductory rates are $2 per 1M input and $10 per 1M output, with cached input at $0.10/1M (95% off). After the intro window the standard rate is $4/$20. So the Sonnet-tier price is a limited-time hook, not the steady state — plan your cost model around $4/$20.
Tops DeepSWE and AutomationBench — but not a clean sweep
New SOTA on DeepSWE v1.1 (77.9%) and AutomationBench (51.3%, vs Claude Opus 5.5's 42.5%), plus LVBench long-video at 91.7%. It still trails GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and OSWorld-2.0 (69.2% vs 72.6%). Strong at long-horizon work; benchmark the specific task before switching.
You probably can't call it yet
At launch Argon is gated to cybersecurity defenders via Google's Fairwind Program; standard API and AI Ultra access are "coming next" with no date. The headline model is announced, not shippable for most builders today — watch the API docs before you architect around it.
A no-guardrails cyber variant, already patching real bugs
Trusted defenders get a version without cyber guardrails that can autonomously find, validate and patch vulnerabilities — Wiz used it to catch a critical healthcare flaw prior models missed. Internally, Argon agents rewrote 32K lines of SIMD (2.7x faster) and are porting 800K+ lines of C/C++ to Rust in Fuchsia.
Magnitude: a free, local inference engine for your agents
Launched on HN, Magnitude (Apache-2.0) profiles your machine, picks and tunes a model, then serves it to Claude Code, Codex, Cline and OpenCode via one-click setup — no API keys, token costs, or rate limits, across Apple Silicon, NVIDIA, AMD and CPU. A clean weekend path to taking your agent stack fully offline.