Qwen3.8-Max lands: 2.4T MoE that tops GPT-5.6 on computer use
Alibaba's biggest model yet: 1M context, $2/$6 pricing, Anthropic-compatible so it slots into Claude Code — and open weights hit Hugging Face next week.

Copy markdown
2.4 trillion parameters, and it's a MoE
Qwen3.8-Max is Alibaba's largest and most capable model to date — a 2.4T-parameter mixture-of-experts (reportedly ~95B active per token) that shipped Aug 3 via the QwenCloud API. It's the new top of the Qwen line.
Open weights hit Hugging Face next week
Qwen says the weights land on Hugging Face and ModelScope next week — the first open-weights release of a Max-class Qwen. If it holds, you'll be able to self-host or fine-tune a frontier-tier model instead of renting one.
It claims the top spot on computer use
Alibaba says Qwen3.8-Max beats GPT-5.6 Sol Max and Fable 5 on agentic computer use, posting 86.1 on OSWorld-Verified. That's the benchmark that matters if you're building agents that click through real apps.
Codes near the frontier — and drops into Claude Code
It scores 86.6 on Terminal-Bench 2.1 and 73.5 on FrontierSWE (up from last gen's 40.7). Endpoints are OpenAI- and Anthropic-compatible, so you can point Claude Code or a Vercel AI Gateway app at it today.
$2 in, $6 out — and cache is 8x cheaper
API pricing is $2 per million input tokens and $6 output, with implicit cached input at $0.25 — eight times cheaper than fresh input. That undercuts the Western flagships it's benchmarking against.
The catch: the flagship is rent-only for now
Until the weights drop, the 2.4T model is API-only, and the activated-param count isn't fully documented — so self-hosting cost is still a guess. On-prem today realistically means the smaller Qwen3.8-27B, not the flagship.