GLM-5.3-Flash ships open MIT weights while the flagship stays shut
Z.ai's 320B-A18B multimodal model runs locally under MIT and costs ~$0.11/M via API. Plus IBM Granite 4.2 reasoning models and Tencent's open WeMM embeddings.

Copy markdown
Open weights, day one, under MIT
GLM-5.3-Flash (previewed as "Ox Alpha") lands on Hugging Face with its full weights under the MIT license, so you can download, fine-tune, and self-host it with no usage strings attached. That's the mirror image of the flagship GLM-5.3, whose weights Z.ai has held back over cyber-risk concerns since mid-August.
320B-A18B, natively multimodal, 1M context
It's a 320B mixture-of-experts with just 18B active per token, and the first natively multimodal model in the GLM-5 line. It takes text plus images, holds a 1M-token context via IndexPool compression, and can look at rendered UI, docs, or slides and revise its own output from the screenshot.
Near-frontier coding at roughly a tenth the price
Z.ai reports 63.4 on DeepSWE v1.1 (up from GLM-5.2's 46.2) and 84.3 on Terminal-Bench 2.1, claiming it approaches Claude Opus 4.8 on coding at about a tenth the cost. API access runs roughly $0.11 per million input tokens and $0.39 output, via OpenRouter, OpenCode, or the GLM Coding Plan (now with a 3x quota bump).
Running it locally isn't trivial yet
The native FP8 release is about 328GB, so early testers say you need multiple GPUs (roughly two to four DGX Spark units) or a 4-bit quant to fit it on smaller rigs. vLLM, SGLang, and Unsloth recipes are already published if you want to serve it yourself.
Elsewhere: Tencent open-sources WeMM multimodal embeddings
Tencent's WeChat Vision team released WeMM-Embedding, a family of open multimodal embedding models in 2B, 4B, and 9B sizes for text-and-image retrieval. The 9B scores 80.6 on MMEB-v2, a strong open drop-in for multimodal RAG and search.