OpenAI reports gains over Nvidia GB200 and GB300 systems in company-run tests, with production deployment planned by the end of 2026.
By Ryan Merket · Published
Primary source: OpenAI
Why it matters #
A production Jalapeno deployment would give OpenAI direct control over part of its inference bill and another source of capacity beyond outside chip suppliers. If performance under ChatGPT and Codex traffic tracks OpenAI's engineering tests, the chip could reduce serving costs and waiting times, but those gains still require independent validation and production-scale proof.
OpenAI, led by co-founder and CEO Sam Altman, reported 1.7x to 3.6x lower end-to-end latency for Jalapeno than for the Nvidia systems used in its comparisons. The company also reported 2.1x to 4.1x higher interactivity and 1.5x to 1.9x more AI work per watt at peak throughput across three large open models. The live engineering report on Jalapeno's first results displays an August 25th, 2026 date, while the supplied source metadata dates the same page to August 3rd. No public revision history establishes whether August 25th reflects an update or the original publication. OpenAI's X announcement is timestamped August 25th. These are OpenAI's measurements on engineering hardware, with production deployment planned for the end of 2026.
Richard Ho, who leads OpenAI's hardware program, now has to carry those laboratory results into production traffic.
Altman co-founded Loopt from a Stanford computer science project and took the startup through Y Combinator's first batch. He later ran the accelerator from 2014 to 2019. (Y Combinator)
OpenAI co-founder and President Greg Brockman described Jalapeno in June as part of a "long-term full-stack infrastructure strategy to make compute more abundant". Ho said OpenAI designed the architecture around the model kernels, memory movement, networking and serving patterns used in frontier workloads. Broadcom handled silicon implementation and networking technology, while Celestica is supporting boards, racks and system integration.
A benchmark with a narrow 104x headline
OpenAI's published results cover GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T on the InferenceX benchmark. Across those tests, OpenAI reports 2.1x to 4.1x higher interactivity, 1.7x to 3.6x lower end-to-end latency, 8.6x to 104.3x higher performance per watt at a prior-best time between tokens, and 1.5x to 1.9x higher performance per watt at peak throughput. The runs were reported by OpenAI and have not been independently reproduced in full. SemiAnalysis said it observed the InferenceX runs in person but did not run the full benchmark suite.
OpenAI describes InferenceX as a public SemiAnalysis benchmark covering the full process of serving an AI request. OpenAI says the comparisons covered operating points ranging from maximum throughput to low-latency interactive serving.
On GPT-OSS 120B, OpenAI compared Jalapeno's 700-watt package rating with a 1,200-watt GB200 system. OpenAI reported about 1.9x more peak throughput per kilowatt, 1.7x lower end-to-end latency and a 2.7x reduction in minimum time between output tokens. On DeepSeek R1, OpenAI's table compares Jalapeno with a 1,400-watt GB300 system and reports 1.7x more peak throughput per kilowatt, 3.6x lower end-to-end latency and 4.1x faster token delivery per user. The report also lists Kimi K2.5 among the tested models, with its results contributing to the published ranges. Jalapeno's sustained power remained at or below 550 watts during the tests, though the comparisons used its published 700-watt package rating.
The largest number on OpenAI's slide, a 104.3x performance-per-watt improvement, applies to a specific DeepSeek R1 operating point. OpenAI measured how much throughput Jalapeno could deliver while matching the fastest time between tokens previously reached by the comparison system. OpenAI's reported peak-throughput advantage for that model was 1.7x. The distinction matters because agents can be especially sensitive to token latency: delays accumulate as one generated step feeds the next.
InferenceX provides a common framework for examining benchmark configurations and results. An independent reproduction of the configurations, software stack and comparison systems is still needed to determine whether the reported advantages persist outside OpenAI's test environment. Production systems will also introduce scheduling, networking, utilization and reliability conditions that a controlled comparison cannot settle.
The chip is built around agents
In its results report, OpenAI describes prompt processing, or prefill, as compute-intensive and token-by-token decoding as more constrained by memory bandwidth. OpenAI also says communication between cores and chips can add latency. OpenAI says Jalapeno can keep model state, including the KV cache, local and coordinate compute, memory and networking within one connected system to minimize data movement and communication delays.
OpenAI says in the results report that earlier OpenAI model generations helped design and bring up Jalapeno, while newer models are accelerating optimization and programming. OpenAI says AI-assisted engineering helped move Jalapeno from initial design to tapeout in nine months. Using Codex with GPT-Astra, OpenAI says it brought three open-weight models that were outside the original production plan to high performance within two months. On selected GPT-OSS attention and mixture-of-experts blocks, OpenAI reported that AI-generated implementations ran 1.5x to 1.8x faster than existing implementations written by human experts. That comparison covers selected blocks rather than full models, and neither claim has been independently validated in the supplied material.
The financial motivation is direct. OpenAI said on March 31st that it closed $122 billion in committed capital at an $852 billion post-money valuation, anchored by Amazon, Nvidia and SoftBank, with Microsoft continuing to participate. OpenAI reported more than 900 million weekly ChatGPT users and more than 2 million weekly Codex users at the time. Serving that demand turns small improvements in watts, latency and chip utilization into material infrastructure economics.
Owning an inference processor also gives OpenAI another supply option. OpenAI says Nvidia remains the foundation of its infrastructure and plans to keep deploying Nvidia and other partners' accelerators for training and inference. Jalapeno adds first-party silicon to that mix. If it reaches production at scale, OpenAI could route workloads according to speed, availability and cost while retaining bargaining power across suppliers.
Production is the remaining test
OpenAI says engineering samples are running model workloads at their target frequency and power. Production qualification, software work and preparations for a year-end deployment are continuing. OpenAI's August results report says Gen 2 is deep in development and Gen 3 is taking shape. Its June announcement described Jalapeno more broadly as the first step in a multi-generation compute platform.
OpenAI unveiled the chip on June 24th. The product impact will depend on whether the measured gains hold under ordinary ChatGPT and Codex traffic.
The results give Ho's hardware program a set of performance targets for that deployment. Jalapeno is already running real models on engineering samples. Its value will be decided when production traffic tests the speed, efficiency and reliability claims OpenAI has published.