Qwen3.8-Max lands: 2.4T MoE that tops GPT-5.6 on computer use

Alibaba's biggest model yet: 1M context, $2/$6 pricing, Anthropic-compatible so it slots into Claude Code — and open weights hit Hugging Face next week.

Nowline AUG 4 10:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 2.4 trillion parameters, and it's a MoE

    Qwen3.8-Max is Alibaba's largest and most capable model to date — a 2.4T-parameter mixture-of-experts (reportedly ~95B active per token) that shipped Aug 3 via the QwenCloud API. It's the new top of the Qwen line.

  • Open weights hit Hugging Face next week

    Qwen says the weights land on Hugging Face and ModelScope next week — the first open-weights release of a Max-class Qwen. If it holds, you'll be able to self-host or fine-tune a frontier-tier model instead of renting one.

  • It claims the top spot on computer use

    Alibaba says Qwen3.8-Max beats GPT-5.6 Sol Max and Fable 5 on agentic computer use, posting 86.1 on OSWorld-Verified. That's the benchmark that matters if you're building agents that click through real apps.

  • Codes near the frontier — and drops into Claude Code

    It scores 86.6 on Terminal-Bench 2.1 and 73.5 on FrontierSWE (up from last gen's 40.7). Endpoints are OpenAI- and Anthropic-compatible, so you can point Claude Code or a Vercel AI Gateway app at it today.

  • $2 in, $6 out — and cache is 8x cheaper

    API pricing is $2 per million input tokens and $6 output, with implicit cached input at $0.25 — eight times cheaper than fresh input. That undercuts the Western flagships it's benchmarking against.

  • The catch: the flagship is rent-only for now

    Until the weights drop, the 2.4T model is API-only, and the activated-param count isn't fully documented — so self-hosting cost is still a guess. On-prem today realistically means the smaller Qwen3.8-27B, not the flagship.