OpenRouter's #1 'stealth' model unmasked: Z.ai's open GLM-5.3-Flash

Anonymous for a week, it outdrew DeepSeek 2:1 before the reveal. The 321B/18B-active weights are MIT on Hugging Face and near Opus 4.8 on coding.

Nowline AUG 28 1:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Update: the 'Stealth' #1 is GLM-5.3-Flash

    The entry listed only as 'Stealth' on OpenRouter since Aug 20 climbed to #1 by usage — over double DeepSeek's volume — before Z.ai confirmed on Aug 26 that it's GLM-5.3-Flash. The weights are now live at huggingface.co/zai-org/GLM-5.3-Flash.

  • Blind usage, not a gamed benchmark

    It topped the chart on real agent traffic — billions of tokens routed through tools like Claude Code — while nobody knew who built it. That's a harder signal of everyday coding quality to fake than any leaderboard you can optimize against.

  • 321B/18B-active, MIT, 1M multimodal context

    Only 18B params fire per token, so it's cheap to serve; MIT lets you ship it inside a product; the 1M-token window takes text and images. API runs $0.15/M in, $0.50/M out, $0.03/M cached — and Z.ai's model card claims it nears Claude Opus 4.8 on coding at a tenth the price.

  • The catch: test it on your own workload

    A $0.15 rate isn't a $0.15 bill — token efficiency varies, so a cheaper sticker can still cost more per finished task. And a Chinese-lab model, self-hosted or not, raises data-governance questions for some stacks. Benchmark it against your current setup before you switch.