Qwen3.8-Max ships: 2.4T MoE, 1M context, $2/$6, weights next week

Alibaba's biggest yet: 95B active of 2.4T, multimodal, live on QwenCloud and Vercel — tops agentic benchmarks but trails Fable 5 on hard coding.

Nowline AUG 3 6:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 2.4T total, 95B active — a Max-class MoE that runs sparse

    Qwen's new flagship fires just 95B of its 2.4T params per token, defaults to xhigh reasoning, and takes text, image and video as input. Alibaba calls it the most capable Qwen model to date.

  • Priced to undercut: $2 in, $6 out, $0.25 cached

    QwenCloud lists $2.00/MTok input and $6.00/MTok output, with implicit cache reads at $0.25 — 8x cheaper than fresh input. Frontier-class capability at well below frontier US pricing.

  • Coding reality check: 67.7 on SWE-bench Pro vs Fable 5's 80.0

    It posts 73.5 on FrontierSWE and 86.6 on Terminal-Bench 2.1, but trails Fable 5 on the hardest coding bench (67.7 vs 80.0). A strong generalist, not the new coding king.

  • Where it wins: agents, docs and vision

    Qwen3.8-Max tops the models it was benchmarked against on PaperBench (93.0), OSWorld-Verified (86.1), OmniDocBench 1.5 (92.1) and GPQA Diamond (92.6). Agentic, document and reasoning work is its sweet spot.

  • Use it today: QwenCloud and Vercel AI Gateway

    Live now as qwen3.8-max on QwenCloud and alibaba/qwen3.8-max on Vercel AI Gateway, with roughly 1M context (991K in, 131K out, 262K reasoning budget) and a 2M-tokens/min rate limit.

  • Open weights land next week — a first for Max-class

    Qwen says weights for both Qwen3.8-Max and a smaller Qwen3.8-27B hit Hugging Face and ModelScope next week — the first time a Max-tier Qwen ships open. Line up a private, no-API-bill deploy.