{"slug": "openai-says-its-first-chip-jalapeno-beats-nvidia-s-gb300-on-inference", "title": "OpenAI Says Its First Chip Jalapeño Beats Nvidia's GB300 on Inference", "summary": "OpenAI said its Broadcom-built Jalapeño inference chip delivers 1.5 to 1.9 times more AI work per watt than Nvidia's GB200 and GB300 systems on tested models, with 1.7 to 3.6 times lower end-to-end latency, based on benchmarks run on SemiAnalysis' InferenceX platform across GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's Kimi K2.5 1T. OpenAI's head of hardware, Richard Ho, estimated Jalapeño would deploy in very small volumes at the end of 2026, with more meaningful deployment in 2027, and the company has not published full system economics such as chip cost, yields, and rack cost. Nvidia, scheduled to report fiscal second-quarter results on August 26, is expected to post revenue of about $92.2 billion, with data center revenue estimated at $85.7 billion, according to S&P Global Market Intelligence.", "body_md": "*OpenAI has a real benchmark for Jalapeño now, but you should read the win as a cost signal, not as the end of Nvidia's grip on AI hardware.*\n\nOpenAI put fresh numbers behind Jalapeño on August 25, and the claim is sharp enough to get Nvidia's attention. The company says its Broadcom-built inference chip delivers about 1.5 to 1.9 times more AI work per watt than Nvidia's GB200 and GB300 systems on tested models, while cutting end-to-end latency by 1.7 to 3.6 times. The numbers are bold. They also come with limits you shouldn't ignore.\n\nOpenAI said the tests ran on InferenceX, the public benchmark platform from SemiAnalysis, across GPT-OSS 120B, DeepSeek R1 670B and Moonshot AI's Kimi K2.5 1T. SemiAnalysis said OpenAI invited its team to examine the chip and benchmark it with the InferenceX suite, which gives the announcement more weight than a normal company slide deck. That is real. It still isn't the same thing as Jalapeño running ChatGPT at scale in a production data center.\n\nAccording to OpenAI's own benchmark post, Jalapeño is rated at 700 watts, while measured sustained power stayed at or below 550 watts on the tested workloads. On DeepSeek R1, OpenAI reported roughly 1.7 times higher peak mixed tokens per second per kilowatt than GB300 and 3.6 times lower end-to-end latency. On Kimi K2.5, the company reported a 1.5 times performance-per-watt lead and 3.4 times lower latency. Those are the figures that matter if you're trying to serve millions of AI responses without letting the power bill eat the product.\n\n## The catch is timing\n\nJalapeño isn't about training OpenAI's next frontier model. Not across the board. It's built for inference, the work of running an already trained model when a user asks ChatGPT, Codex or the API to do something. If you use these tools every day, this is the part of the stack you actually feel: the wait before the first useful answer, the speed of a coding agent, the cost OpenAI has to carry every time demand spikes.\n\n[Sam Altman Fears a Handful of Companies Will Control AI](https://startupfortune.com/sam-altman-fears-a-handful-of-companies-will-control-ai/)\n\nOpenAI CEO Sam Altman told podcaster David Senra that his biggest fear is a small number of companies gatekeeping AI, framing it as a choice between \"AI authoritarianism or liberty.\" The remarks land awkwardly given OpenAI's own dependence on Microsoft's cloud and Nvidia's chips, and its pending push for a $1 trillion IPO valuation. - [sam altman fears ai companies controlling market access](https://startupfortune.com/sam-altman-fears-a-handful-of-companies-will-control-ai/) - [how many companies will dominate artificial intelligence industry](https://startupfortune.com/sam-altman-fears-a-handful-of-companies-will-control-ai/)\n\nTechCrunch reported that Richard Ho, OpenAI's head of hardware, estimated Jalapeño would deploy at the end of 2026 in very small volumes, with more meaningful deployment coming in 2027. That sentence should cool the hype. A benchmark result in August 2026 can shape supplier talks and investor expectations, but it won't cut OpenAI's infrastructure bill this quarter. Not this year.\n\nThere are also details still missing. OpenAI has published performance-per-watt and latency comparisons, but it hasn't published the full economics of the system: chip cost, yields, rack cost, networking cost, cooling cost and utilization once real workloads are mixed together. Engineers care about benchmarks. Finance teams care about cost per useful token served at acceptable latency. That is the business case.\n\n## Nvidia still has the harder moat\n\nNvidia is not stuck. The company is scheduled to report fiscal second-quarter results on August 26, and Nvidia itself said the call will cover the quarter ended July 26. S&P Global Market Intelligence, using Visible Alpha consensus, put expected quarterly revenue at about $92.2 billion, with data center revenue estimated at $85.7 billion. You don't unseat that with one ASIC and a good chart.\n\nStill, Jalapeño lands in the most sensitive part of Nvidia's AI business. Training clusters made Nvidia the company every AI lab needed. Inference is where the next fight gets uglier, because the biggest customers have enough volume to ask whether a custom chip can save money over time. OpenAI doesn't need Jalapeño to replace every Nvidia GPU. It needs Jalapeño to handle enough repeatable inference work that Nvidia has less pricing power over OpenAI's fastest-growing operating cost.\n\nBroadcom gets paid either way. OpenAI's June announcement named Broadcom and Celestica as partners, with Broadcom handling silicon implementation, networking and connectivity and Celestica working on board, rack and system integration. Data Center Dynamics reported from Hot Chips that Ho described 128-chip racks and full pods of 2,048 ASICs, with each 128-chip deployment capable of 1.7 exaflops of 4-bit compute and 27.5TB of HBM4. That is not a lab curiosity if it ships.\n\nThe honest read is simple. OpenAI has shown that Jalapeño can beat leading Nvidia systems on public inference benchmarks for several large models, and SemiAnalysis involvement makes the claim harder to dismiss. But the market shouldn't confuse a benchmark with a deployed fleet. Scale is the test. Between now and 2027, the number that matters most isn't 1.9x. It's how many Jalapeño units OpenAI can actually put to work.\n\n**Also read:** [ARIA Bans AI-Generated Songs From Australia's Official Music Charts](https://startupfortune.com/aria-bans-ai-generated-songs-from-australias-official-music-charts/) • [Navitas Agrees To Pay Up To $232.8 Million For AI Power Startup Claros](https://startupfortune.com/navitas-agrees-to-pay-up-to-2328-million-for-ai-power-startup-claros/) • [Jack Ma Buys $77 Million of Alibaba Shares After Its AI Fundraise Sent Stock Lower](https://startupfortune.com/jack-ma-buys-77-million-of-alibaba-shares-after-its-ai-fundraise-sent-stock-lower/)\n\n[Google, Amazon and Meta Are Out-Growing Nvidia With Their Own AI Chips](https://startupfortune.com/google-amazon-and-meta-are-out-growing-nvidia-with-their-own-ai-chips/)\n\nCustom AI chips from Google, Amazon, Microsoft and Meta are growing shipments 44.6% this year, nearly triple Nvidia's GPU growth rate, according to TrendForce. Nvidia's revenue and margins are still climbing, but its own biggest customers are quietly building the competition. - [custom AI chips outpacing Nvidia GPU growth rates](https://startupfortune.com/google-amazon-and-meta-are-out-growing-nvidia-with-their-own-ai-chips/) - [why tech giants building proprietary AI accelerators instead](https://startupfortune.com/google-amazon-and-meta-are-out-growing-nvidia-with-their-own-ai-chips/)", "url": "https://wpnews.pro/news/openai-says-its-first-chip-jalapeno-beats-nvidia-s-gb300-on-inference", "canonical_source": "https://startupfortune.com/openai-says-its-first-chip-jalapeo-beats-nvidias-gb300-on-inference/", "published_at": "2026-08-25 15:15:57+00:00", "updated_at": "2026-08-25 15:43:08.946776+00:00", "lang": "en", "topics": ["ai-chips", "ai-infrastructure", "ai-research"], "entities": ["OpenAI", "Broadcom", "Nvidia", "Jalapeño", "SemiAnalysis", "InferenceX", "GPT-OSS 120B", "DeepSeek R1 670B"], "alternates": {"html": "https://wpnews.pro/news/openai-says-its-first-chip-jalapeno-beats-nvidia-s-gb300-on-inference", "markdown": "https://wpnews.pro/news/openai-says-its-first-chip-jalapeno-beats-nvidia-s-gb300-on-inference.md", "text": "https://wpnews.pro/news/openai-says-its-first-chip-jalapeno-beats-nvidia-s-gb300-on-inference.txt", "jsonld": "https://wpnews.pro/news/openai-says-its-first-chip-jalapeno-beats-nvidia-s-gb300-on-inference.jsonld"}}