GLM-5.3-Flash and Qwen3.8-Flash-Next: two open Flash models, one day

Both under $0.16/M input, open under MIT and Apache, posting frontier-beating coding scores you can self-host. Plus: ChatGPT tasks now fire on webhooks.

Nowline AUG 26 11:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • GLM-5.3-Flash: 320B open weights, MIT, $0.15 in

    Z.ai shipped its first natively multimodal GLM-5 model — 320B total, 18B active, 1M-token context — under the MIT license, with weights already on Hugging Face. The API runs $0.15/$0.50 per million ($0.03 cached input), and Zhipu reports it edging Opus 4.8 on GDPVal-AA v2 (1,773 vs 1,582) and hitting 84.3 on Terminal-Bench 2.1, just shy of Opus 4.8's 85.0.

  • Qwen3.8-Flash-Next: 6B active per token, Apache 2.0

    Alibaba's MoE activates just 6B of 125B parameters per token, pairing a 262K context (1M via YaRN) with an N-gram embedding layer that lives in system RAM instead of VRAM. Apache-2.0 weights are on Hugging Face and ModelScope; the production API is $0.16/$0.47 per million. Qwen reports it beating DeepSeek-V4-Flash and Opus 4.6 on most benchmarks (SWE-bench Pro 62.5), trained at one-ninth its predecessor's cost.

  • The math: agentic coding for pennies

    Both sit in the ~$0.15/M-input Flash tier — roughly an order of magnitude under frontier closed models on the token-hungry agent loops builders actually run. And because both ship open weights, you can pull the GGUFs and run them locally: quantized GLM-5.3-Flash and Qwen3.8-Flash-Next builds are already up on Hugging Face.

  • ChatGPT scheduled tasks now fire on webhooks

    Separately, OpenAI wired ChatGPT's scheduled tasks to webhook triggers from Gmail, Slack, and GitHub, so a task runs when something changes instead of on a clock. Even free accounts get up to three active tasks — enough to build event-driven automations without a polling loop or a paid tier.