GLM-5.3 lands API-first: a 50% coding jump from post-training alone

Z.ai's coding model is live on an OpenAI-compatible API with a 1M-context route — but the open weights are held two weeks over cyber-misuse fears.

Nowline AUG 16 12:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • The gains are post-training, not a new base

    GLM-5.3 leaves GLM-5.2's base model untouched and gets everything from scaled post-training: Terminal-Bench 3.0 jumps 6.2x to 28.3 and DeepSWE v1.1 hits 66.9% (~45% over 5.2). Near-frontier agentic coding out of RL iteration, not a bigger pretrain.

  • Use it today: OpenAI-compatible, 1M context

    It's live now as `glm-5.3` on Z.ai's OpenAI-compatible endpoint, with a `glm-5.3[1m]` 1M-context route on the Coding Plan (Lite $18, Pro $80, Max $168/mo). Per-token API rates aren't published yet — don't assume 5.2's $1.40/$4.40 carries over.

  • Weights held two weeks over hacking risk

    Z.ai says open weights follow ~2 weeks after launch, gated on safety evals; Axios reports the hold is over cyber-misuse concerns. GLM-5.2 shipped MIT, but 5.3's license is still unstated. Self-hosters: pencil in late August, not now.

  • Tuned for cyber defense, capped on offense

    GLM-5.3 leads CyberGym at 84.5% (defensive) while deliberately trailing ExploitBench (54.4% vs Fable 5's 78%). Z.ai says it built the model to help defenders, not write exploits — a rare on-purpose capability gap.

  • The pitch is tokens, not just scores

    Z.ai's own runs show GLM-5.3 nearing Claude Fable 5 on agent tasks at ~50K tokens per task versus ~120K for rivals. Vendor-measured, not independently reproduced — but if it holds, that's a real cost lever on long agent runs.