Gemini 3.7 Flash: the cheap agent workhorse, half price till Dec 31

Google's fastest model yet — 340 tok/s, 1M context, big agentic-coding jumps. Plus GLM-5.3's open-weights claim and Claude Code pulling its todo tools.

Nowline Aug 18 1:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Half the price — but only through year-end

    Gemini 3.7 Flash runs $0.75 / $3.75 per million tokens (in/out) as an intro rate, then doubles to $1.50 / $7.50 on Jan 1, 2027. If you route real agent traffic through it, that cheap window is a budgeting decision, not a footnote.

  • The agentic-coding jumps are the real headline

    Against 3.6 Flash, Google reports DeepSWE coding 49.0% to 65.3%, AutomationBench 17.0% to 30.4%, and WebDev Arena +50 Elo (1538 to 1588). Vendor-run numbers, but the deltas point at fewer broken multi-step runs.

  • Fastest output of any model right now

    At ~340 tokens/sec it ranks #1 for output speed across 186 tracked models, with a 1M-token context and a March 2026 cutoff. For latency-bound agent loops, that speed compounds on every step.

  • Report: GLM-5.3 tops open coding — but you can't download it yet

    Z.ai's GLM-5.3 (743B MoE, 1M context) reportedly leads several open-model coding benchmarks and claims first on CyberGym (84.5%). Catch: the numbers are vendor-run and the weights are only promised ~2 weeks post-launch, so for now it's API / Coding-Plan only.

  • Also new: Claude Code just pulled its todo tools on newer models

    Claude Code 2.1.233 removes the todo/task-tracking tools on Opus 4.8, Sonnet 5, Fable 5 and Mythos 5+ — set CLAUDE_CODE_ENABLE_TODO_TOOLS=1 to bring them back. Same release adds WebFetch response caching and GitLab merge-request support.