GLM-5.3-Flash's launch discount ends Sept 9 — every rate doubles
Z.ai's cheap, multimodal 1M-context model loses its 50% promo and free caching this week — on Z.ai and OpenRouter alike. The MIT weights are the way out.

Copy markdown
Rates double on September 9
The 50% launch discount expires Sept 9 at 24:00 UTC+8. Input goes $0.075 → $0.15 and output $0.25 → $0.50 per million tokens, so anything you route to GLM-5.3-Flash costs 2x the next morning.
Free caching ends the same day
Cached input is free during the promo; on Sept 9 it reverts to $0.03/M. If you cache long system prompts or RAG context, that line goes from zero to real — factor it in before you lock the quarter's budget.
It's OpenRouter too, not just Z.ai
OpenRouter carries the same 50%-off sticker ($0.075/$0.25) over the same $0.15/$0.50 list, so proxying through a gateway won't dodge the hike. Every path to this model doubles at once.
The escape hatch: MIT open weights
GLM-5.3-Flash ships open under MIT (zai-org/GLM-5.3-Flash on Hugging Face), commercial use included, with vLLM and SGLang support. Running real volume? Self-host the 320B/18B-active MoE and pay $0 per token instead of the new list rate.
Even at list, it's still cheap
Don't panic-switch: $0.15/$0.50 is still low for a natively multimodal, 1M-context model that scores 57 on Artificial Analysis's Intelligence Index. What closes Sept 9 is the arbitrage window, not the model — decide whether 2x still pencils for your traffic.