cd /news/ai-chips/openai-says-its-first-chip-jalapeno-… · home topics ai-chips article
[ARTICLE · art-110402] src=startupfortune.com ↗ pub= topic=ai-chips verified=true sentiment=· neutral

OpenAI Says Its First Chip Jalapeño Beats Nvidia's GB300 on Inference

OpenAI said its Broadcom-built Jalapeño inference chip delivers 1.5 to 1.9 times more AI work per watt than Nvidia's GB200 and GB300 systems on tested models, with 1.7 to 3.6 times lower end-to-end latency, based on benchmarks run on SemiAnalysis' InferenceX platform across GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's Kimi K2.5 1T. OpenAI's head of hardware, Richard Ho, estimated Jalapeño would deploy in very small volumes at the end of 2026, with more meaningful deployment in 2027, and the company has not published full system economics such as chip cost, yields, and rack cost. Nvidia, scheduled to report fiscal second-quarter results on August 26, is expected to post revenue of about $92.2 billion, with data center revenue estimated at $85.7 billion, according to S&P Global Market Intelligence.

read5 min views6 publishedAug 25, 2026
OpenAI Says Its First Chip Jalapeño Beats Nvidia's GB300 on Inference
Image: Startupfortune (auto-discovered)

OpenAI has a real benchmark for Jalapeño now, but you should read the win as a cost signal, not as the end of Nvidia's grip on AI hardware.

OpenAI put fresh numbers behind Jalapeño on August 25, and the claim is sharp enough to get Nvidia's attention. The company says its Broadcom-built inference chip delivers about 1.5 to 1.9 times more AI work per watt than Nvidia's GB200 and GB300 systems on tested models, while cutting end-to-end latency by 1.7 to 3.6 times. The numbers are bold. They also come with limits you shouldn't ignore.

OpenAI said the tests ran on InferenceX, the public benchmark platform from SemiAnalysis, across GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI's Kimi K2.5 1T. SemiAnalysis said OpenAI invited its team to examine the chip and benchmark it with the InferenceX suite, which gives the announcement more weight than a normal company slide deck. That is real. It still isn't the same thing as Jalapeño running ChatGPT at scale in a production data center.

According to OpenAI's own benchmark post, Jalapeño is rated at 700 watts, while measured sustained power stayed at or below 550 watts on the tested workloads. On DeepSeek R1, OpenAI reported roughly 1.7 times higher peak mixed tokens per second per kilowatt than GB300 and 3.6 times lower end-to-end latency. On Kimi K2.5, the company reported a 1.5 times performance-per-watt lead and 3.4 times lower latency. Those are the figures that matter if you're trying to serve millions of AI responses without letting the power bill eat the product.

The catch is timing #

Jalapeño isn't about training OpenAI's next frontier model. Not across the board. It's built for inference, the work of running an already trained model when a user asks ChatGPT, Codex or the API to do something. If you use these tools every day, this is the part of the stack you actually feel: the wait before the first useful answer, the speed of a coding agent, the cost OpenAI has to carry every time demand spikes.

Sam Altman Fears a Handful of Companies Will Control AI OpenAI CEO Sam Altman told podcaster David Senra that his biggest fear is a small number of companies gatekeeping AI, framing it as a choice between "AI authoritarianism or liberty." The remarks land awkwardly given OpenAI's own dependence on Microsoft's cloud and Nvidia's chips, and its pending push for a $1 trillion IPO valuation. - sam altman fears ai companies controlling market access - how many companies will dominate artificial intelligence industry

TechCrunch reported that Richard Ho, OpenAI's head of hardware, estimated Jalapeño would deploy at the end of 2026 in very small volumes, with more meaningful deployment coming in 2027. That sentence should cool the hype. A benchmark result in August 2026 can shape supplier talks and investor expectations, but it won't cut OpenAI's infrastructure bill this quarter. Not this year.

There are also details still missing. OpenAI has published performance-per-watt and latency comparisons, but it hasn't published the full economics of the system: chip cost, yields, rack cost, networking cost, cooling cost and utilization once real workloads are mixed together. Engineers care about benchmarks. Finance teams care about cost per useful token served at acceptable latency. That is the business case.

Nvidia still has the harder moat #

Nvidia is not stuck. The company is scheduled to report fiscal second-quarter results on August 26, and Nvidia itself said the call will cover the quarter ended July 26. S&P Global Market Intelligence, using Visible Alpha consensus, put expected quarterly revenue at about $92.2 billion, with data center revenue estimated at $85.7 billion. You don't unseat that with one ASIC and a good chart.

Still, Jalapeño lands in the most sensitive part of Nvidia's AI business. Training clusters made Nvidia the company every AI lab needed. Inference is where the next fight gets uglier, because the biggest customers have enough volume to ask whether a custom chip can save money over time. OpenAI doesn't need Jalapeño to replace every Nvidia GPU. It needs Jalapeño to handle enough repeatable inference work that Nvidia has less pricing power over OpenAI's fastest-growing operating cost.

Broadcom gets paid either way. OpenAI's June announcement named Broadcom and Celestica as partners, with Broadcom handling silicon implementation, networking and connectivity and Celestica working on board, rack and system integration. Data Center Dynamics reported from Hot Chips that Ho described 128-chip racks and full pods of 2,048 ASICs, with each 128-chip deployment capable of 1.7 exaflops of 4-bit compute and 27.5TB of HBM4. That is not a lab curiosity if it ships.

The honest read is simple. OpenAI has shown that Jalapeño can beat leading Nvidia systems on public inference benchmarks for several large models, and SemiAnalysis involvement makes the claim harder to dismiss. But the market shouldn't confuse a benchmark with a deployed fleet. Scale is the test. Between now and 2027, the number that matters most isn't 1.9x. It's how many Jalapeño units OpenAI can actually put to work.

Also read: ARIA Bans AI-Generated Songs From Australia's Official Music ChartsNavitas Agrees To Pay Up To $232.8 Million For AI Power Startup ClarosJack Ma Buys $77 Million of Alibaba Shares After Its AI Fundraise Sent Stock Lower

Google, Amazon and Meta Are Out-Growing Nvidia With Their Own AI Chips Custom AI chips from Google, Amazon, Microsoft and Meta are growing shipments 44.6% this year, nearly triple Nvidia's GPU growth rate, according to TrendForce. Nvidia's revenue and margins are still climbing, but its own biggest customers are quietly building the competition. - custom AI chips outpacing Nvidia GPU growth rates - why tech giants building proprietary AI accelerators instead

── more in #ai-chips 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-says-its-firs…] indexed:0 read:5min 2026-08-25 ·