Cerebras CS-4 pushes inference to 4,400 tokens/sec, 30x GPU

Three WSE-3 Turbo wafers and 750 PFLOPS, shipping this quarter — reachable free today. Plus DFlash 2, Gemini free for students, OpenAI's zero-retention play.

Nowline AUG 20 9:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 4,400 tokens/sec, 30x a GPU

    The CS-4 stacks three WSE-3 Turbo wafers for 750 PFLOPS and 129.6 PB/s of memory bandwidth, clearing 4,400 tokens/sec on GPT-OSS-120B and 1,000+ tokens/sec on 10-trillion-parameter models — up to 30x a GPU rack. First shipments land this quarter.

  • You can hit that speed today, free

    You don't need the silicon to feel it. Cerebras Inference already serves gpt-oss-120b and GLM on an OpenAI-compatible endpoint with 1M free tokens/day and no credit card — swap your base_url and agent loops and tool calls come back near-instantly.

  • DFlash 2 makes local models 2.7-3.4x faster

    Inco's speculative-decoding update adds a path selector and dynamic convolutions to squeeze 20% more tokens out of every verification pass: 2.7-3.4x throughput on Qwen3.8-27B and up to 4.6x on Meta's Muse Glimmer. Weights are on Hugging Face, with vLLM, SGLang and llama.cpp support.

  • Gemini Pro goes free for US students

    Google is handing eligible US college students 12 months of AI Pro at no cost — Gemini's top app tier, 5TB of storage, and higher usage limits, plus a new Student Hub of study tools. A free year to build against a frontier model.

  • OpenAI previews zero-retention safety scanning

    OpenAI is testing Private Safety Processing: it flags misuse patterns across related API calls while keeping zero data retention for paying customers — a jab at Anthropic, which now requires safety logs. If you build on the API, your prompts stay unstored.