Cerebras CS-4 pushes inference to 4,400 tokens/sec, 30x GPU
Three WSE-3 Turbo wafers and 750 PFLOPS, shipping this quarter — reachable free today. Plus DFlash 2, Gemini free for students, OpenAI's zero-retention play.

Copy markdown
4,400 tokens/sec, 30x a GPU
The CS-4 stacks three WSE-3 Turbo wafers for 750 PFLOPS and 129.6 PB/s of memory bandwidth, clearing 4,400 tokens/sec on GPT-OSS-120B and 1,000+ tokens/sec on 10-trillion-parameter models — up to 30x a GPU rack. First shipments land this quarter.
You can hit that speed today, free
You don't need the silicon to feel it. Cerebras Inference already serves gpt-oss-120b and GLM on an OpenAI-compatible endpoint with 1M free tokens/day and no credit card — swap your base_url and agent loops and tool calls come back near-instantly.
DFlash 2 makes local models 2.7-3.4x faster
Inco's speculative-decoding update adds a path selector and dynamic convolutions to squeeze 20% more tokens out of every verification pass: 2.7-3.4x throughput on Qwen3.8-27B and up to 4.6x on Meta's Muse Glimmer. Weights are on Hugging Face, with vLLM, SGLang and llama.cpp support.
Gemini Pro goes free for US students
Google is handing eligible US college students 12 months of AI Pro at no cost — Gemini's top app tier, 5TB of storage, and higher usage limits, plus a new Student Hub of study tools. A free year to build against a frontier model.
OpenAI previews zero-retention safety scanning
OpenAI is testing Private Safety Processing: it flags misuse patterns across related API calls while keeping zero data retention for paying customers — a jab at Anthropic, which now requires safety logs. If you build on the API, your prompts stay unstored.