OpenAI's Jalapeño chip beats Blackwell on inference efficiency
First public results from OpenAI's custom inference ASIC, made with Broadcom — but real volume is a 2027 story. Plus: Copilot retires 6 models Aug 31.

Copy markdown
More tokens per watt than Nvidia's Blackwell
At the Hot Chips conference, OpenAI showed Jalapeño — its first in-house inference ASIC — serving more tokens per user and more throughput per kilowatt than current state-of-the-art on SemiAnalysis' InferenceX benchmark, run against Nvidia Blackwell. It was co-designed with Broadcom, hit tape-out in about nine months, and OpenAI says its own models helped design the silicon.
Why your OpenAI latency and bill could drop — later
OpenAI frames the chip as its path to faster ChatGPT answers, cheaper API pricing, and steadier access during demand spikes, by serving more work per watt. The catch for your roadmap: only 'very small volumes' arrive by the end of 2026, with meaningful capacity in 2027 — so nothing changes on your invoice this year.
OpenAI edges off Nvidia — days after Nvidia grabbed Hugging Face
Jalapeño is OpenAI's bet to lean less on Nvidia GPUs for inference, arriving the same week Nvidia moved to buy the open-model hub Hugging Face. Samsung is reportedly supplying the HBM4 memory, with Celestica handling boards and racks; deployment targets gigawatt scale with data-center partners across multiple chip generations.
Elsewhere: Copilot retires 6 models Aug 31, ships default model policy
GitHub Copilot will retire six older models on Aug 31 after a 30-day notice, so audit any workflow that pins a specific model ID. GitHub also moved its 'global model policy' to GA: org models are now enabled by default — except open-weight ones like DeepSeek and Kimi K2, and models outside its data-retention agreement, which stay off unless an admin turns them on.