Gemini 3.7 Flash lands: a coding workhorse at half the price

$0.75/M through year-end, a 1M-token window, and a 9-point coding jump — but no open weights, and the rate doubles in January. Plus GPT-5.6 Sol at 750 tok/s.

Nowline AUG 15 3:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • The cheap model is the agent workhorse now

    Gemini 3.7 Flash ships with a 1M-token context window, up to 64K output tokens, and text, image, audio, and video input. It's live today in the Gemini API, Google AI Studio, Antigravity, and Android Studio — drop it into an existing agent loop without rewriting your harness.

  • $0.75 per 1M input — half the last Flash, under everyone

    Introductory pricing is $0.75/M input and $3.75/M output through Dec 31 — about half Gemini 3.6 Flash, and well below Claude Sonnet 5 ($2/$10) and GPT-5.6 Terra ($2/$12). For high-volume agent runs, it's the cheapest near-frontier tier you can call right now.

  • The intro rate doubles January 1 — budget for it

    On Jan 1, 2027 the price jumps to $1.50/M input and $7.50/M output — a clean 2x. If you're wiring Flash into production, model the post-holiday cost now so a New Year invoice doesn't blindside you.

  • Coding up ~9 points, with agent benchmarks to match

    FrontierCode 1.1 climbed from 34.4% to 43.6%, DeepSWE v1.1 hit 65.3%, Terminal-bench 2.1 reached 85.8%, and WebDev Arena posted a 1588 Elo. Google is pitching it as a workhorse for real coding and agentic work, not just chat.

  • The catch: no open weights, 64K output ceiling

    There's no self-hosting or air-gapped option, so regulated and offline builds are still out. And the 64K output cap means very long single-shot generations need chunking or streaming.

  • Elsewhere: GPT-5.6 Sol hits 750 tokens/sec on Cerebras

    OpenAI previewed an Ultrafast mode for GPT-5.6 Sol running up to 14x faster — as high as 750 output tokens/sec — on Cerebras hardware. It's a limited preview for select customers today; sign up to get in line for latency-critical agents.