Update: OpenAI's Jalapeño hits up to 1.9x better perf/watt vs GB300

Official InferenceX results also show 3-4x lower latency on DeepSeek R1 and Kimi K2.5; the Broadcom-built chip ships inside OpenAI's datacenters by year-end.

Nowline AUG 26 2:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • The official numbers

    OpenAI's first published Jalapeño results show 1.9x higher performance-per-kilowatt on GPT-OSS 120B — 85,448 vs 44,960 mixed tokens/kW — plus 1.7x lower end-to-end latency, all at a 700W rating that stayed at or below 550W sustained.

  • The gap widens on bigger models

    Against Nvidia GB200/GB300, Jalapeño posted 3.6x lower latency on DeepSeek R1 670B and 3.4x lower on Kimi K2.5 1T. The efficiency lead grows with model size — exactly the trillion-parameter, long-context regime that agent workloads push toward.

  • You can't buy it, but you'll feel it

    Jalapeño is an internal inference ASIC co-designed with Broadcom, not a retail card; OpenAI begins deploying it across its own datacenters by the end of 2026. The builder payoff is downstream — cheaper, faster inference on OpenAI's API and ChatGPT models as capacity scales.

  • Independent read: ~50% cheaper than Blackwell

    SemiAnalysis's teardown estimates Jalapeño's cost per token undercuts Nvidia Blackwell by roughly half across OpenAI's serving mix — reportedly the strongest sign yet that custom inference silicon can bite into Nvidia's margins.