Gemini 3.6 Flash cuts output price to $7.50/M, codes better

Google's new Flash tier runs 17% leaner on tokens, gains double digits on coding and computer-use benchmarks, and adds a $0.30 Lite and a security variant.

Nowline JUL 23 4:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Output down to $7.50/M — and 17% fewer tokens to get there

    Gemini 3.6 Flash runs $1.50 input / $7.50 output per million tokens, down from 3.5 Flash's $9 output, and Google says it needs 17% fewer output tokens to finish comparable work — a compounding discount for high-volume agent loops. That's roughly half Claude Sonnet 5's output rate, with a 1M-token context and 65K max output.

  • A real jump on software-engineering and computer-use tasks

    DeepSWE climbs 37% to 49% and MLE-Bench 49.7% to 63.9%; the computer-use benchmark OSWorld-Verified hits 83% (up from 78.4%). It also posts 78% on Terminal-Bench 2.1 and 58.7% on SWE-Bench Pro — Flash-tier pricing doing agentic coding you'd have reached for a Pro model to handle.

  • Flash-Lite lands at $0.30/M for high-volume jobs

    Gemini 3.5 Flash-Lite is priced at $0.30 input / $2.50 output per million tokens for cheap, high-throughput workloads — classification, extraction, routing — where you fan out millions of calls and every fraction of a cent compounds.

  • A security-tuned Flash Cyber, and a Gemini 4 tease

    Gemini 3.5 Flash Cyber is a security-focused variant rolling out through Google's CodeMender pilot (limited access) for code-security work. Google also teased Gemini 4 and said 3.5 Pro remains in restricted testing past its June target.