GLM-5.3-Flash and Qwen3.8-Flash-Next: two open Flash models, one day
Both under $0.16/M input, open under MIT and Apache, posting frontier-beating coding scores you can self-host. Plus: ChatGPT tasks now fire on webhooks.

Copy markdown
GLM-5.3-Flash: 320B open weights, MIT, $0.15 in
Z.ai shipped its first natively multimodal GLM-5 model — 320B total, 18B active, 1M-token context — under the MIT license, with weights already on Hugging Face. The API runs $0.15/$0.50 per million ($0.03 cached input), and Zhipu reports it edging Opus 4.8 on GDPVal-AA v2 (1,773 vs 1,582) and hitting 84.3 on Terminal-Bench 2.1, just shy of Opus 4.8's 85.0.
Qwen3.8-Flash-Next: 6B active per token, Apache 2.0
Alibaba's MoE activates just 6B of 125B parameters per token, pairing a 262K context (1M via YaRN) with an N-gram embedding layer that lives in system RAM instead of VRAM. Apache-2.0 weights are on Hugging Face and ModelScope; the production API is $0.16/$0.47 per million. Qwen reports it beating DeepSeek-V4-Flash and Opus 4.6 on most benchmarks (SWE-bench Pro 62.5), trained at one-ninth its predecessor's cost.
The math: agentic coding for pennies
Both sit in the ~$0.15/M-input Flash tier — roughly an order of magnitude under frontier closed models on the token-hungry agent loops builders actually run. And because both ship open weights, you can pull the GGUFs and run them locally: quantized GLM-5.3-Flash and Qwen3.8-Flash-Next builds are already up on Hugging Face.
ChatGPT scheduled tasks now fire on webhooks
Separately, OpenAI wired ChatGPT's scheduled tasks to webhook triggers from Gmail, Slack, and GitHub, so a task runs when something changes instead of on a clock. Even free accounts get up to three active tasks — enough to build event-driven automations without a polling loop or a paid tier.