
OpenAI's Jalapeño chip beats Nvidia Blackwell on inference
OpenAI showed the first benchmark results for its custom inference chip, Jalapeño, at the Hot Chips conference on Tuesday, TechCrunch reported.
Tested on Semianalysis's InferenceX benchmark, Jalapeño registered more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors. The comparison ran against an Nvidia Blackwell system, though by the time Jalapeño reaches full deployment, Nvidia's lineup may have moved on.
Those two metrics measure different things a company running inference at scale needs simultaneously. Tokens per user tracks how fast a single response streams back to one person, while throughput per kilowatt tracks how many customers a data center can serve for a given amount of power. A chip that only wins on one of those tends to force a tradeoff between speed and cost; Jalapeño's benchmark claims both.
“The bottom line is that the results show a very, very significant performance advance over state of the art. Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency.”
— Richard Ho, OpenAI's head of hardware
Ho laid out a cautious timeline for the chip on a press call. Jalapeño is set to deploy at the end of 2026 in very small volumes, with meaningful deployment following in 2027.
OpenAI first announced Jalapeño last October, developing it in close collaboration with Broadcom, with OpenAI's own models assisting in the design process. The company frames Jalapeño as the start of a multigenerational platform, one where AI products, models, chips, and memory get built in concert rather than treated as separate layers.
That full-stack approach targets specific friction points in how inference runs. OpenAI singled out the prefill and communication phases of processing as recurring bottlenecks. "We designed Jalapeño to minimize data movement and communication delays," the company wrote in a blog post presenting the results. "This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase."
The chip fits a pattern across the industry: Google runs its own TPUs, Amazon has Trainium, and Microsoft is building Maia, each aimed at cutting reliance on Nvidia for the specific workloads that matter most to them. For OpenAI, that dependence runs deep, and expensive. OpenAI's Ohio data center lease alone puts Nvidia on the hook for guarantees worth up to $105 billion, a scale of commitment that makes an in-house alternative for at least part of the inference workload worth pursuing, even on a slow rollout.
- Chip: Jalapeño, OpenAI's custom inference processor
- Benchmark: InferenceX results beat state-of-the-art Nvidia Blackwell on tokens per user and throughput per kilowatt
- Development: announced October 2025, built with Broadcom, OpenAI's own models assisted in design
- Timeline: small-volume deployment end of 2026, meaningful deployment in 2027
- Design focus: minimizing data movement and communication delays during the prefill phase
For now, Jalapeño stays a small-volume proof of concept. The benchmark numbers give OpenAI something concrete to point to, but the chip's actual weight in OpenAI's infrastructure will not show up until the 2027 rollout Ho described.
This piece is informational, not a recommendation to buy, sell, or hold any asset.

Comments (0)
No comments yet — be the first!
Related news
Most readTop 7
Silicon Valley Workers Are Wearing Noise-Cancelling Masks to Dictate AI Prompts
256AI





