GLM-5.3 leads open coding benchmarks; Z.ai holds back the weights

Live now on Z.ai's $18/mo plan and 6x better on Terminal-Bench 3.0, but its exploit-finding froze the open release. Plus Meta's $0.10-in coding tier.

Nowline AUG 23 7:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • SOTA on open benchmarks, live on the $18 plan

    GLM-5.3 jumps ~6x on Terminal-Bench 3.0 (4.6 to 28.3) and hits 66.9 on DeepSWE v1.1 — best-in-class among open models by Z.ai's numbers. The 744B, 1M-context model is usable today only through the GLM Coding Plan ($18/mo Lite) and ZCode; no per-token API rate is published yet.

  • Too good at finding exploits to open-source — yet

    In Z.ai's own tests the model surfaced 2,436 vulnerabilities across 269 real projects (1,097 critical/high) and scored 84.5% on CyberGym, so the open weights are frozen behind a ~2-week safety review. Z.ai is instead shipping OpenVuln, a Hugging Face tool that points the same skill at your repos — though the security numbers are self-reported, not yet reproduced.

  • Meta's Muse now costs $0.10 in — if it can train on you

    Meta's new Contributor tier drops the 1M-context Muse Spark 1.2 to $0.10/M input and $0.20/M output — about a tenth of the $1.25/$4.25 standard rate. The catch: your prompts and outputs may be used to improve Meta's products, so it's cheap enough to leave an agent running as long as the code isn't sensitive.