Qwen3.8-Max ships: 2.4T MoE, 1M context, $2/$6, weights next week
Alibaba's biggest yet: 95B active of 2.4T, multimodal, live on QwenCloud and Vercel — tops agentic benchmarks but trails Fable 5 on hard coding.

Copy markdown
2.4T total, 95B active — a Max-class MoE that runs sparse
Qwen's new flagship fires just 95B of its 2.4T params per token, defaults to xhigh reasoning, and takes text, image and video as input. Alibaba calls it the most capable Qwen model to date.
Priced to undercut: $2 in, $6 out, $0.25 cached
QwenCloud lists $2.00/MTok input and $6.00/MTok output, with implicit cache reads at $0.25 — 8x cheaper than fresh input. Frontier-class capability at well below frontier US pricing.
Coding reality check: 67.7 on SWE-bench Pro vs Fable 5's 80.0
It posts 73.5 on FrontierSWE and 86.6 on Terminal-Bench 2.1, but trails Fable 5 on the hardest coding bench (67.7 vs 80.0). A strong generalist, not the new coding king.
Where it wins: agents, docs and vision
Qwen3.8-Max tops the models it was benchmarked against on PaperBench (93.0), OSWorld-Verified (86.1), OmniDocBench 1.5 (92.1) and GPQA Diamond (92.6). Agentic, document and reasoning work is its sweet spot.
Use it today: QwenCloud and Vercel AI Gateway
Live now as qwen3.8-max on QwenCloud and alibaba/qwen3.8-max on Vercel AI Gateway, with roughly 1M context (991K in, 131K out, 262K reasoning budget) and a 2M-tokens/min rate limit.
Open weights land next week — a first for Max-class
Qwen says weights for both Qwen3.8-Max and a smaller Qwen3.8-27B hit Hugging Face and ModelScope next week — the first time a Max-tier Qwen ships open. Line up a private, no-API-bill deploy.