# How OpenAI's Jalapeno chip just surprised the AI industry

> Source: <https://www.thedeepview.com/articles/how-openai-s-jalapeno-chip-just-surprised-the-ai-industry>
> Published: 2026-08-25 14:00:00+00:00

OpenAI has built its first AI chip, and the early returns on performance look impressive.

On Tuesday at the Hot Chips conference, OpenAI said that its custom inference chip, Jalapeño, is now up and running on working silicon and is already in testing. The frontier lab revealed the first third-party benchmarks and claims that it's accelerating performance-per-watt, the favorite measurement of efficiency in the AI industry. It's basically a way of showing how much AI work a chip can complete for every unit of power it consumes.

On tests against various types of workloads, OpenAI claims that Jalapeño delivers 1.5 to 4 times the performance-per-watt compared to today's leading AI chips.

OpenAI used [SemiAnalysis's open-source InferenceX](https://inferencex.semianalysis.com/) benchmark to test Jalapeño, using three models: GPT-OSS ([OpenAI's open-source model](https://www.thedeepview.com/articles/why-openai-quietly-embraced-open-models)), DeepSeek R1, and Kimi K2.5. In other words, it didn't just test Jalapeño on its own models fine-tuned to work on its own hardware.

And while its own GPT-OSS model performed the best, the other models also performed remarkably well. The biggest surprise was that Jalapeño outperformed Nvidia's Blackwell, one of the current industry standards for AI chips. However, we don't yet know how it will perform against other inference accelerator leaders like Groq and Cerebres.

[SemiAnalysis characterized it](https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia) this way: "Jalapeño beats Blackwell on perf/W across almost all scenarios without being tuned for any specific point in the curve. It excels not only in low-latency scenarios but also in high-throughput scenarios."

In a briefing with the media, I asked OpenAI's head of hardware, Richard Ho, where OpenAI would use this newfound power in its AI stack (because obviously it won't be running all workloads on its own chips anytime soon). "It's clear that the cost of serving low latency with this device is a lot lower than other low latency devices," said Ho, "and so there may be a good niche there where some of our very latency sensitive products such as Codex Ultra Fast Mode — even faster than the [fastest mode] we have now — may be a good option… We are expecting to ramp up the volume into 2027."

*OpenAI's Jalapeño chip. Image credit: OpenAI*

## Our Deeper *View*

Other than the performance itself, the most impressive thing about Jalapeño may be how fast OpenAI was able to develop it and bring it into production. The project, which is a partnership with Broadcom, began as a concept in late 2024. The fact that it already has working silicon less than two years later is just as surprising as the benchmark performance. OpenAI credits its own AI models to helping the hardware team dramatically accelerate the process. The team reported that OpenAI's models aided in the optimization, programming, and development of the chips and are currently doing the same for the next generations that come after Jalapeño. While the speed is great and will help meet user expectations over time, the real win for OpenAI is in efficiency. If it can design and optimize its own chips then it can spend far less on compute and lower the long-term cost of intelligence to bring its most advanced models and features to more users.
