GLM-5.3-Flash's launch discount ends Sept 9 — every rate doubles

Z.ai's cheap, multimodal 1M-context model loses its 50% promo and free caching this week — on Z.ai and OpenRouter alike. The MIT weights are the way out.

Nowline SEP 6 3:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Rates double on September 9

    The 50% launch discount expires Sept 9 at 24:00 UTC+8. Input goes $0.075 → $0.15 and output $0.25 → $0.50 per million tokens, so anything you route to GLM-5.3-Flash costs 2x the next morning.

  • Free caching ends the same day

    Cached input is free during the promo; on Sept 9 it reverts to $0.03/M. If you cache long system prompts or RAG context, that line goes from zero to real — factor it in before you lock the quarter's budget.

  • It's OpenRouter too, not just Z.ai

    OpenRouter carries the same 50%-off sticker ($0.075/$0.25) over the same $0.15/$0.50 list, so proxying through a gateway won't dodge the hike. Every path to this model doubles at once.

  • The escape hatch: MIT open weights

    GLM-5.3-Flash ships open under MIT (zai-org/GLM-5.3-Flash on Hugging Face), commercial use included, with vLLM and SGLang support. Running real volume? Self-host the 320B/18B-active MoE and pay $0 per token instead of the new list rate.

  • Even at list, it's still cheap

    Don't panic-switch: $0.15/$0.50 is still low for a natively multimodal, 1M-context model that scores 57 on Artificial Analysis's Intelligence Index. What closes Sept 9 is the arbitrage window, not the model — decide whether 2x still pencils for your traffic.