OpenAI's Jalapeño inference chip outpaces Nvidia's GB300 in tests
OpenAI's first custom silicon posts up to 1.9x better tokens-per-watt than Nvidia's best — the supply-side lever behind future API cost, speed, and capacity.

Copy markdown
1.9x more tokens per watt than a GB300
On SemiAnalysis's InferenceX benchmark, Jalapeño served GPT-OSS 120B at 85,448 mixed tokens/kW against the GB300's 44,960 — 1.9x the throughput per watt — and returned answers 1.7x faster end-to-end (1.03s vs 1.80s). The gap widens on bigger models: 3.6x lower latency on DeepSeek R1 670B.
Snappier token-by-token streaming
Minimum time-between-tokens — what makes a streamed response feel instant — dropped 2.7x on GPT-OSS 120B (0.69ms vs 1.87ms) and 4.1x on DeepSeek R1. Jalapeño is built only to serve models, not train them.
OpenAI's own silicon, co-designed with Broadcom
It's OpenAI's first in-house inference chip, built with Broadcom and benchmarked on SemiAnalysis's InferenceX. Chip lead Richard Ho called the early numbers 'a very, very significant performance advance over state of the art.'
Live end of 2026 — small volumes, internal only
OpenAI plans to switch it on inside its own datacenters by year-end in 'very small volumes,' scaling through 2027. There's no chip or board to buy: it runs OpenAI's models, with no plan to sell it externally.
Why a chip you can't buy still matters to you
Power-per-token is OpenAI's operating leverage — the same lever behind this week's fight over the 5-hour Codex cap. Cheaper inference at scale is what eventually loosens rate limits and lowers API prices, and it's OpenAI's first real step off total Nvidia dependence. Nothing changes in your stack today; this is a 2027 story to start watching.