GLM-5.3-Flash: MIT-licensed open weights, 1M context, cheap flash
Z.ai's 320B/18B MoE self-hosts via vLLM and undercuts paid tiers; Qwen3.8-Flash-Next runs on ~75GB RAM — the open-weight capability premium is collapsing.

Copy markdown
A 320B MoE you can host yourself
GLM-5.3-Flash is 320B total with just 18B active, a 1M-token context, and an MIT license — the weights are on Hugging Face and self-host via SGLang or vLLM. The API runs $0.15/$0.50 per million tokens, with a $0.075/$0.25 launch promo through Sept 9.
Strong agentic-coding scores at flash prices
It posts 63.4 on DeepSWE and 78.4 on Toolathlon Verified, and lands 57 on Artificial Analysis's Intelligence Index at roughly $0.045 a task — flash-tier cost for scores that until recently sat a tier up.
Qwen3.8-Flash-Next fits on one machine
Alibaba's Qwen4-architecture preview pairs a 125B MoE with a 51B N-gram component and a 1M context, and reportedly runs locally on ~75GB of RAM at $0.16/$0.47 per million — a real single-box option for private inference.
The capability premium is collapsing
1M context, native multimodality, and permissive open weights are now table stakes at the cheap tier — so the model becomes the cheap part and your real cost is effort settings, cache-hit rate, and provider choice. Worth re-pricing your stack.