OpenAI's Jalapeño inference chip outpaces Nvidia's GB300 in tests

OpenAI's first custom silicon posts up to 1.9x better tokens-per-watt than Nvidia's best — the supply-side lever behind future API cost, speed, and capacity.

Nowline AUG 26 12:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 1.9x more tokens per watt than a GB300

    On SemiAnalysis's InferenceX benchmark, Jalapeño served GPT-OSS 120B at 85,448 mixed tokens/kW against the GB300's 44,960 — 1.9x the throughput per watt — and returned answers 1.7x faster end-to-end (1.03s vs 1.80s). The gap widens on bigger models: 3.6x lower latency on DeepSeek R1 670B.

  • Snappier token-by-token streaming

    Minimum time-between-tokens — what makes a streamed response feel instant — dropped 2.7x on GPT-OSS 120B (0.69ms vs 1.87ms) and 4.1x on DeepSeek R1. Jalapeño is built only to serve models, not train them.

  • OpenAI's own silicon, co-designed with Broadcom

    It's OpenAI's first in-house inference chip, built with Broadcom and benchmarked on SemiAnalysis's InferenceX. Chip lead Richard Ho called the early numbers 'a very, very significant performance advance over state of the art.'

  • Live end of 2026 — small volumes, internal only

    OpenAI plans to switch it on inside its own datacenters by year-end in 'very small volumes,' scaling through 2027. There's no chip or board to buy: it runs OpenAI's models, with no plan to sell it externally.

  • Why a chip you can't buy still matters to you

    Power-per-token is OpenAI's operating leverage — the same lever behind this week's fight over the 5-hour Codex cap. Cheaper inference at scale is what eventually loosens rate limits and lowers API prices, and it's OpenAI's first real step off total Nvidia dependence. Nothing changes in your stack today; this is a 2027 story to start watching.