OpenAI previews Ultrafast: GPT-5.6 Sol at up to 750 tokens/sec

A new API tier runs OpenAI's frontier model 14x faster on Cerebras chips — a capacity-gated preview, open to select customers with a public waitlist.

Nowline AUG 14 12:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 14x speed, no quality tax

    Ultrafast serves GPT-5.6 Sol at up to 750 output tokens per second — about 14x OpenAI's standard tier and 5x Opus 4.8's Fast mode — with “no quality degradation” on the GDP-Val benchmark. Same frontier model, far less waiting.

  • Cerebras wafers are the engine

    The speed comes from Cerebras' Wafer-Scale Engine: 44GB of on-chip SRAM per wafer keeps the weights resident instead of streaming them from memory. Cerebras calls it a frontier-model speed record — it cleared Humanity's Last Exam's 2,500 questions in 11h11m versus ~78h for Claude Fable 5.

  • API-only, and gated for now

    This is a limited preview on the OpenAI API — not in ChatGPT, not open enrollment. Access starts with a “select group of customers” and widens as capacity grows; everyone else joins the waitlist at openai.com/form/ultrafast. Pricing wasn't disclosed.

  • What 750 tok/s unlocks

    Sub-second frontier responses change what you can build: real-time voice agents, agentic loops that don't stall step-to-step, live coding assistants, interactive research. As one early user put it, speed “changes what people can realistically use it for.”