OpenAI previews Ultrafast: GPT-5.6 Sol at up to 750 tokens/sec
A new API tier runs OpenAI's frontier model 14x faster on Cerebras chips — a capacity-gated preview, open to select customers with a public waitlist.

Copy markdown
14x speed, no quality tax
Ultrafast serves GPT-5.6 Sol at up to 750 output tokens per second — about 14x OpenAI's standard tier and 5x Opus 4.8's Fast mode — with “no quality degradation” on the GDP-Val benchmark. Same frontier model, far less waiting.
Cerebras wafers are the engine
The speed comes from Cerebras' Wafer-Scale Engine: 44GB of on-chip SRAM per wafer keeps the weights resident instead of streaming them from memory. Cerebras calls it a frontier-model speed record — it cleared Humanity's Last Exam's 2,500 questions in 11h11m versus ~78h for Claude Fable 5.
API-only, and gated for now
This is a limited preview on the OpenAI API — not in ChatGPT, not open enrollment. Access starts with a “select group of customers” and widens as capacity grows; everyone else joins the waitlist at openai.com/form/ultrafast. Pricing wasn't disclosed.
What 750 tok/s unlocks
Sub-second frontier responses change what you can build: real-time voice agents, agentic loops that don't stall step-to-step, live coding assistants, interactive research. As one early user put it, speed “changes what people can realistically use it for.”