Qwen3.8-Max-0902: terminal-coding score up 2.5x, price unchanged
Alibaba's API-only 2.4T flagship gets a coding-focused refresh — new #1 marks on several agent benchmarks, cache reads from $0.17/M, still shy of Opus 5.

Copy markdown
A coding-focused refresh of the 2.4T flagship
Alibaba shipped qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02) on the QwenCloud/DashScope API — same 2.4T-parameter MoE and 1M-token context as before, but post-trained specifically for agentic coding and office/'Cowork' tasks. It's proprietary and API-only; no open weights this time.
Terminal-Bench score jumps from 11.3 to 29
The refresh more than doubles Terminal-Bench 3.0 (29.0 vs 11.3), lifts DeepSWE 1.1 to 69.3 (from 56.6) and repo-level NL2Repo to 64.9 (from 55.9). Several agent benchmarks — CoWorkBench 76.1%, QwenSWEBench V2 70%, JobBench 64% — now top the field it's tracked against.
Same $2/$6 pricing, so the gains are basically free
Prices didn't move: $2/M input, $6/M output, with explicit cache reads at $0.17/M and implicit hits at $0.25/M. If you're already calling Max, the coding bump costs nothing extra. It speaks the OpenAI-compatible API and is live on OpenRouter and the usual aggregators today.
Still a step behind Opus 5 on the hardest tasks
Don't retire your Opus 5 agent: it still leads on Terminal-Bench (42.7 vs 29.0), DeepSWE (73.6 vs 69.3) and NL2Repo. But 0902 edges Opus on repo-level understanding (SWE-Atlas 66.3 vs 63.2) and automation — a cheaper option for high-context, tool-heavy work you'd otherwise run at frontier rates.