Gemini 3.7 Flash: the cheap agent workhorse, half price till Dec 31
Google's fastest model yet — 340 tok/s, 1M context, big agentic-coding jumps. Plus GLM-5.3's open-weights claim and Claude Code pulling its todo tools.

Copy markdown
Half the price — but only through year-end
Gemini 3.7 Flash runs $0.75 / $3.75 per million tokens (in/out) as an intro rate, then doubles to $1.50 / $7.50 on Jan 1, 2027. If you route real agent traffic through it, that cheap window is a budgeting decision, not a footnote.
The agentic-coding jumps are the real headline
Against 3.6 Flash, Google reports DeepSWE coding 49.0% to 65.3%, AutomationBench 17.0% to 30.4%, and WebDev Arena +50 Elo (1538 to 1588). Vendor-run numbers, but the deltas point at fewer broken multi-step runs.
Fastest output of any model right now
At ~340 tokens/sec it ranks #1 for output speed across 186 tracked models, with a 1M-token context and a March 2026 cutoff. For latency-bound agent loops, that speed compounds on every step.
Report: GLM-5.3 tops open coding — but you can't download it yet
Z.ai's GLM-5.3 (743B MoE, 1M context) reportedly leads several open-model coding benchmarks and claims first on CyberGym (84.5%). Catch: the numbers are vendor-run and the weights are only promised ~2 weeks post-launch, so for now it's API / Coding-Plan only.
Also new: Claude Code just pulled its todo tools on newer models
Claude Code 2.1.233 removes the todo/task-tracking tools on Opus 4.8, Sonnet 5, Fable 5 and Mythos 5+ — set CLAUDE_CODE_ENABLE_TODO_TOOLS=1 to bring them back. Same release adds WebFetch response caching and GitLab merge-request support.