GLM-5.3-Flash ships open MIT weights while the flagship stays shut

Z.ai's 320B-A18B multimodal model runs locally under MIT and costs ~$0.11/M via API. Plus IBM Granite 4.2 reasoning models and Tencent's open WeMM embeddings.

Nowline AUG 26 10:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Open weights, day one, under MIT

    GLM-5.3-Flash (previewed as "Ox Alpha") lands on Hugging Face with its full weights under the MIT license, so you can download, fine-tune, and self-host it with no usage strings attached. That's the mirror image of the flagship GLM-5.3, whose weights Z.ai has held back over cyber-risk concerns since mid-August.

  • 320B-A18B, natively multimodal, 1M context

    It's a 320B mixture-of-experts with just 18B active per token, and the first natively multimodal model in the GLM-5 line. It takes text plus images, holds a 1M-token context via IndexPool compression, and can look at rendered UI, docs, or slides and revise its own output from the screenshot.

  • Near-frontier coding at roughly a tenth the price

    Z.ai reports 63.4 on DeepSWE v1.1 (up from GLM-5.2's 46.2) and 84.3 on Terminal-Bench 2.1, claiming it approaches Claude Opus 4.8 on coding at about a tenth the cost. API access runs roughly $0.11 per million input tokens and $0.39 output, via OpenRouter, OpenCode, or the GLM Coding Plan (now with a 3x quota bump).

  • Running it locally isn't trivial yet

    The native FP8 release is about 328GB, so early testers say you need multiple GPUs (roughly two to four DGX Spark units) or a 4-bit quant to fit it on smaller rigs. vLLM, SGLang, and Unsloth recipes are already published if you want to serve it yourself.

  • Elsewhere: Tencent open-sources WeMM multimodal embeddings

    Tencent's WeChat Vision team released WeMM-Embedding, a family of open multimodal embedding models in 2B, 4B, and 9B sizes for text-and-image retrieval. The 9B scores 80.6 on MMEB-v2, a strong open drop-in for multimodal RAG and search.