Update: OpenAI's Jalapeño hits up to 1.9x better perf/watt vs GB300
Official InferenceX results also show 3-4x lower latency on DeepSeek R1 and Kimi K2.5; the Broadcom-built chip ships inside OpenAI's datacenters by year-end.

Copy markdown
The official numbers
OpenAI's first published Jalapeño results show 1.9x higher performance-per-kilowatt on GPT-OSS 120B — 85,448 vs 44,960 mixed tokens/kW — plus 1.7x lower end-to-end latency, all at a 700W rating that stayed at or below 550W sustained.
The gap widens on bigger models
Against Nvidia GB200/GB300, Jalapeño posted 3.6x lower latency on DeepSeek R1 670B and 3.4x lower on Kimi K2.5 1T. The efficiency lead grows with model size — exactly the trillion-parameter, long-context regime that agent workloads push toward.
You can't buy it, but you'll feel it
Jalapeño is an internal inference ASIC co-designed with Broadcom, not a retail card; OpenAI begins deploying it across its own datacenters by the end of 2026. The builder payoff is downstream — cheaper, faster inference on OpenAI's API and ChatGPT models as capacity scales.
Independent read: ~50% cheaper than Blackwell
SemiAnalysis's teardown estimates Jalapeño's cost per token undercuts Nvidia Blackwell by roughly half across OpenAI's serving mix — reportedly the strongest sign yet that custom inference silicon can bite into Nvidia's margins.