cd /news/ai-chips/openais-jalapeno-chip-is-built-for-f… · home topics ai-chips article
[ARTICLE · art-110323] src=techcrunch.com ↗ pub= topic=ai-chips verified=true sentiment=↑ positive

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

OpenAI's Jalapeño inference chip outperformed Nvidia's Blackwell system on Semianalysis's InferenceX benchmark, delivering more tokens per user and more throughput per kilowatt, according to results presented at the Hot Chips conference on Tuesday. Richard Ho, OpenAI's head of hardware, said the chip shows a 'very, very significant performance advance over state of the art,' with deployment expected in very small volumes at the end of 2026 and more significant deployment in 2027. Developed with Broadcom and OpenAI's own models, Jalapeño is designed to minimize data movement and communication delays during inference.

read2 min views3 publishedAug 25, 2026
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
Image: TechCrunch AI

At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors.

“The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”

Notably, that comparison is against an Nvidia Blackwell system — but by the time Jalapeño reaches full deployment, the competition may have advanced significantly. Ho estimated that Jalapeño would deploy at the end of 2026 “in very small volumes,” with more significant deployment coming in 2027.

First announced last October, Jalapeño was developed by OpenAI in close collaboration with Broadcom, with OpenAI’s own models assisting in the development process. The company plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips and memory all developed in concert.

Because of that full-stack approach, OpenAI was able to address specific phases in the inference process that often cause friction during inference processing. In particular, Jalapeño is designed to minimize delays during the prefill and communication phases of processing, which OpenAI says often act as bottlenecks.

“We designed Jalapeño to minimize data movement and communication delays,” the company said in a blog post presenting the results. “This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.”

── more in #ai-chips 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openais-jalapeno-chi…] indexed:0 read:2min 2026-08-25 ·