GPT-5.6 Sol runs 14x faster on Cerebras: 750 tokens/sec in the API
OpenAI's fastest tier yet is a Cerebras-powered preview, gated to select API customers. What 750 tok/s unlocks, where Fast mode fits, how to get in.

Copy markdown
750 tokens a second, straight from the API
Ultrafast mode pushes GPT-5.6 Sol to up to 750 output tokens/sec — roughly 14x its Standard speed — over the regular OpenAI API. Frontier-quality output now arrives faster than you can read it, erasing the lag that made big models feel unusable in real-time apps.
Cerebras is doing the heavy lifting
The jump comes from Cerebras wafer-scale hardware, billed as the next step in OpenAI's reported ~$10B inference partnership. It's the strongest signal yet that OpenAI is buying latency GPUs can't touch — and that the fastest frontier inference may not run on NVIDIA.
The catch: it's a gated preview
Ultrafast is limited to a select group of API customers, with no public price and no launch date — you fill out a form and wait for capacity. Treat it as a roadmap signal on where per-token latency is heading, not a knob you can flip tonight.
What real-time speed actually unlocks
OpenAI aims this at voice agents, live support, fraud checks, and interactive research — anywhere a round-trip kills the UX. It makes talk-over-you voice and sub-second agent tool-loops feasible on a frontier model, instead of forcing you down to a smaller, dumber one.
Three tiers: trade cost for speed per call
Ultrafast tops a ladder of Standard, Fast, and Ultrafast. Fast reportedly runs about 2.5x quicker at roughly double the price and is the tier most builders can reach today, while Ultrafast capacity ramps up.