{"slug": "openais-jalapeno-benchmark-promises-faster-cheaper-ai", "title": "OpenAI’s Jalapeño Benchmark Promises Faster, Cheaper AI", "summary": "OpenAI published first benchmark results for its custom AI inference chip Jalapeño, co-developed with Broadcom, showing 1.5 to 1.9 times greater peak performance per watt than Nvidia Blackwell systems in tests using SemiAnalysis' InferenceX benchmark. The chip, delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom CEO Hock Tan and President Charlie Kawwas in June 2026, was tested with GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, demonstrating support for non-OpenAI models. OpenAI also launched an Ultrafast service tier powered by Cerebras that runs GPT-5.6 Sol at up to 14 times the speed of Standard processing, generating up to 750 output tokens per second, as part of a multi-vendor infrastructure strategy.", "body_md": "In June 2026, Broadcom CEO Hock Tan and President Charlie Kawwas delivered Jalapeño to OpenAI CEO Sam Altman and President Greg Brockman. The custom AI inference chip was designed by OpenAI and co-developed with Broadcom.\n\nOpenAI has now published its first benchmark results for Jalapeño following tests using InferenceX, a public AI inference benchmark from SemiAnalysis. The company said the chip delivered 1.5 to 1.9 times greater peak performance per watt than the Nvidia Blackwell systems used for comparison.\n\nThe results were produced by OpenAI rather than an independent testing organization, and the company normalized them using each accelerator’s published power rating.\n\n## What Jalapeño brings\n\nJalapeño was designed to work across different models, not only OpenAI’s systems. The [ASIC (Application-Specific Integrated Circuit)](https://www.lenovo.com/us/en/glossary/application-specific-integrated-circuit/?srsltid=AfmBOor-1cwigbvqvpGkuaXM7VWiv25rXrC4oK38Wr-uEZfPhi1YOtt5) was tested with GPT-OSS 120B and two non-OpenAI models: [DeepSeek](https://www.techrepublic.com/article/news-us-firms-try-deepseek-ai-costs-rise/) R1 670B and [Kimi](https://www.techrepublic.com/article/news-moonshot-ai-kimi-k3-largest-open-source-model-apac/) K2.5 1T. OpenAI said the results demonstrate that the architecture is not restricted to its own models.\n\nAI inference has several phases with different bottlenecks. During prefill, the system processes the user’s prompt, which is compute-intensive. During decode, it generates the response token by token and relies more heavily on memory bandwidth. Communication between cores and chips can add further latency and reduce an AI model’s responsiveness.\n\nJalapeño has been designed to combine both high batch throughput and real-time responsiveness. OpenAI [says that](https://openai.com/index/jalapeno-first-results/) it has built a flexible AI accelerator “that can support changing model architectures, excel at both prefill and decode, and adapt as the balance between them changes, a defining feature of agentic workloads.”\n\nOne architectural feature behind this performance is a localized KV cache. OpenAI says model data can be explicitly placed and kept close to the required compute resources, reducing the time and power spent moving data during inference.\n\nOpenAI says the resulting design can support high batch throughput and real-time responsiveness without making the same trade-off between throughput and latency found in some existing systems.\n\n### More must-read AI coverage\n\n-\n[SS&C Intralinks DealCentre AI vs. Datasite: Which platform is built for the future of dealmaking?](https://www.techrepublic.com/article/dealcentre-ai-vs-datasite/) -\n[SS&C Intralinks FundCentre AI vs. Juniper Square: Which platform better supports modern private markets fund managers?](https://www.techrepublic.com/article/fundcentre-ai-vs-juniper-square/) -\n[Why Data, Not Models, Determines AI Success](https://www.techrepublic.com/sponsored/why-data-not-models-determines-ai-success/) -\n[The Rise of the AI-Native Factory: How Physical AI Is Transforming Manufacturing](https://www.eweek.com/a/artificial-intelligence/the-rise-of-the-ai-native-factory-how-physical-ai-is-transforming-manufacturing/)\n\n## What Jalapeño means for enterprise users\n\nEarlier this month, OpenAI [launched](https://openai.com/index/openai-broadcom-jalapeno-inference-chip/) a limited preview of an Ultrafast service tier that runs [GPT-5.6 Sol](https://www.techrepublic.com/article/news-openai-gpt-5-6-us-security-review/) at up to 14 times the speed of Standard processing. The service, powered by Cerebras, can generate up to 750 output tokens per second.\n\nThe Cerebras announcement came less than two months after OpenAI unveiled Jalapeño. Together, the announcements show that OpenAI is pursuing a multi-vendor infrastructure strategy rather than immediately replacing Nvidia and Cerebras hardware with its own chip.\n\nOpenAI has confirmed that Jalapeño will complement rather than replace its partner-supplied accelerators. The company said: “Meeting growing demand for AI will require more compute from every available source. We will continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads.”\n\nFor enterprise customers, Jalapeño could eventually mean faster AI responses, greater service capacity and lower inference costs. OpenAI plans to begin deploying the chip within its own infrastructure by the end of 2026, but it has not said whether customers will be able to select the hardware directly or how the efficiency gains will affect API pricing. Those details will determine whether the benchmark produces a measurable advantage for businesses.", "url": "https://wpnews.pro/news/openais-jalapeno-benchmark-promises-faster-cheaper-ai", "canonical_source": "https://www.techrepublic.com/article/news-openai-jalapeno-ai-chip-benchmark/", "published_at": "2026-08-27 09:21:05+00:00", "updated_at": "2026-08-28 07:19:30.010262+00:00", "lang": "en", "topics": ["ai-chips", "ai-infrastructure", "ai-products"], "entities": ["OpenAI", "Broadcom", "Nvidia", "SemiAnalysis", "InferenceX", "Jalapeño", "GPT-OSS 120B", "DeepSeek R1 670B"], "alternates": {"html": "https://wpnews.pro/news/openais-jalapeno-benchmark-promises-faster-cheaper-ai", "markdown": "https://wpnews.pro/news/openais-jalapeno-benchmark-promises-faster-cheaper-ai.md", "text": "https://wpnews.pro/news/openais-jalapeno-benchmark-promises-faster-cheaper-ai.txt", "jsonld": "https://wpnews.pro/news/openais-jalapeno-benchmark-promises-faster-cheaper-ai.jsonld"}}