Gemini 4 Argon: 1M output tokens at $2/$10 — but gated at launch

Google's frontier model 4x's max output and leads DeepSWE and AutomationBench, but ships to cyber defenders first. Plus Magnitude runs local models free.

Nowline OCT 1 2:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 1M output tokens in a single call

    Argon returns up to 1M output tokens per response — roughly 4x the ~128K ceiling on rival frontier models. Enough to emit a fully refactored module, a long migration diff, or a book-length structured report in one turn, without chunking and stitching.

  • Intro price $2/$10, cache at $0.10 — then it doubles

    Introductory rates are $2 per 1M input and $10 per 1M output, with cached input at $0.10/1M (95% off). After the intro window the standard rate is $4/$20. So the Sonnet-tier price is a limited-time hook, not the steady state — plan your cost model around $4/$20.

  • Tops DeepSWE and AutomationBench — but not a clean sweep

    New SOTA on DeepSWE v1.1 (77.9%) and AutomationBench (51.3%, vs Claude Opus 5.5's 42.5%), plus LVBench long-video at 91.7%. It still trails GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and OSWorld-2.0 (69.2% vs 72.6%). Strong at long-horizon work; benchmark the specific task before switching.

  • You probably can't call it yet

    At launch Argon is gated to cybersecurity defenders via Google's Fairwind Program; standard API and AI Ultra access are "coming next" with no date. The headline model is announced, not shippable for most builders today — watch the API docs before you architect around it.

  • A no-guardrails cyber variant, already patching real bugs

    Trusted defenders get a version without cyber guardrails that can autonomously find, validate and patch vulnerabilities — Wiz used it to catch a critical healthcare flaw prior models missed. Internally, Argon agents rewrote 32K lines of SIMD (2.7x faster) and are porting 800K+ lines of C/C++ to Rust in Fuchsia.

  • Magnitude: a free, local inference engine for your agents

    Launched on HN, Magnitude (Apache-2.0) profiles your machine, picks and tunes a model, then serves it to Claude Code, Codex, Cline and OpenCode via one-click setup — no API keys, token costs, or rate limits, across Apple Silicon, NVIDIA, AMD and CPU. A clean weekend path to taking your agent stack fully offline.