Cerebras serves open-weights Qwen3.8-27B at 1,500 tok/s
Prompt caching keeps agent loops cheap at that speed. Plus: Claude Code lets orgs push MCP servers to every user, and Cursor agents move into Vercel.

Copy markdown
1,500 tokens a second, with caching
Cerebras is now serving Alibaba's open-weights Qwen3.8-27B at roughly 1,500 tokens/sec — testers on Hacker News say the output is genuinely hard to read in real time. Prompt caching is supported, so repeated-context agent runs stay cheap, and it's on pay-as-you-go token pricing today via the Cerebras inference endpoint.
Claude Code hands orgs the MCP keys
Claude Code 2.1.259 lets organizations provision HTTP/SSE MCP servers to every user automatically, and adds unattended headless permission controls plus GitLab merge-request recognition — the missing pieces if you run Claude Code in CI or scheduled jobs.
Cursor's cloud agents move into Vercel
Vercel now runs Cursor Cloud Agents inside Vercel Sandbox on scale-to-zero microVMs with durable orchestration, so you can fan agents out without standing up your own infra. Pro and Enterprise teams also get cheaper 'Basic' build machines at 2 vCPU / 8 GB.
Elsewhere: GLM-5.3 is half-price on Vercel's Gateway
Vercel's AI Gateway is carrying GLM-5.3 via DigitalOcean at 50% off through Sept 7 — a cheap window to A/B a strong non-frontier model against your current stack behind a single API key.