AMD and Cerebras Partner on Disaggregated AI Inference System Pairing Helios Racks with Wafer-Scale Chips AMD and Cerebras Systems announced a partnership to build a disaggregated AI inference system combining AMD's Helios rackscale solution with Cerebras's Wafer-Scale Engine, claiming up to 5x higher tokens per second per watt versus Cerebras WSE-only configurations. The joint solution will be available through Cerebras Cloud in the second half of 2026, with Cerebras deploying AMD Helios systems in its data centers. Cerebras shares rose about 4% on the news, while AMD traded near $553 during its Advancing AI 2026 event. AMD and Cerebras Partner on Disaggregated AI Inference System Pairing Helios Racks with Wafer-Scale Chips - AMD and Cerebras will combine Helios rackscale systems with Cerebras's Wafer-Scale Engine for a disaggregated AI inference architecture 1 https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution - The joint solution claims up to 5x higher tokens per second per watt compared to Cerebras WSE-only configurations 1 https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution - Cerebras will deploy AMD Helios in its data centers, with the system available through Cerebras Cloud in H2 2026 1 https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution - Cerebras shares gained about 4% on the announcement; AMD traded near $553 during its Advancing AI 2026 event 2 https://www.cnbc.com/2026/07/23/cerebras-stock-gains-on-amd-partnership.html 3 https://www.fxleaders.com/news/2026/07/22/amd-stock-analysis-553-at-advancing-ai-2026-helios-pricing-and-the-700-target/ - AMD Helios racks pack 72 MI455X GPUs with 31 TB of HBM4 memory and are priced at $5-$5.5 million per rack 4 https://www.amd.com/en/products/rackscale-solutions/helios.html AMD and Cerebras Systems on Wednesday announced a technical partnership to build a disaggregated AI inference system that pairs AMD's Helios rackscale solution with Cerebras's Wafer-Scale Engine. The companies said the architecture delivers up to 5x higher tokens per second per watt compared to Cerebras WSE-only configurations 1 . The announcement came during AMD's Advancing AI 2026 event at San Francisco's Moscone Center, where CEO Lisa Su also unveiled the commercial launch of Zen 6 EPYC Venice processors on TSMC's 2nm node and confirmed Helios rack pricing at $5 to $5.5 million 3 . Cerebras shares rose approximately 4% on the news, trading around $207, while AMD held near $553 — more than double its price at the start of 2026 2 https://www.cnbc.com/2026/07/23/cerebras-stock-gains-on-amd-partnership.html . 3 https://www.fxleaders.com/news/2026/07/22/amd-stock-analysis-553-at-advancing-ai-2026-helios-pricing-and-the-700-target/ Under the partnership, Cerebras will deploy AMD Helios systems in its data centers, with the joint inference solution expected to become available through Cerebras Cloud in the second half of 2026 1 . How the System Works The partnership splits the AI inference workload across two specialized compute engines. AMD Helios handles the compute-intensive prompt-processing stage — ingesting user inputs, processing large context windows, and generating internal representations. Cerebras's Wafer-Scale Engine then takes over for the memory-bandwidth-intensive token-generation phase, where ultra-low latency is critical for real-time applications 1 . By disaggregating these two stages, each system operates in the regime where it is strongest rather than forcing a single architecture to handle both. AMD's Helios racks bring 72 Instinct MI455X GPUs, approximately 2.9 exaFLOPS of FP4 compute, and 31 TB of HBM4 memory per rack, connected via UALink over Ethernet with up to 260 TB/s of aggregate scale-up bandwidth 4 . Cerebras contributes its wafer-scale chip, which the company has positioned as the fastest single-device inference engine available . 1 https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution Strategic Context The deal extends the competitive landscape beyond the binary Nvidia-vs.-everyone framing that has dominated AI infrastructure. AMD has been steadily building out its AI ecosystem, with OpenAI and Meta combining for 12 gigawatts of AMD accelerator capacity and Microsoft Azure and Oracle named as early Helios customers 3 . Cerebras, which went public on Nasdaq under ticker CBRS, has carved out a niche with its wafer-scale approach to AI chips and recently secured a chip supply deal with OpenAI . 5 https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html For Cerebras, the partnership addresses a key gap: while its Wafer-Scale Engine excels at low-latency token generation, it benefits from a high-throughput front end for prompt processing. For AMD, the deal adds another channel for Helios rack sales and validates its open-standards approach — Helios is built on OCP Open Rack Wide, UALink, and Ultra Ethernet Consortium specifications 4 . Market Reaction Cerebras shares gained roughly 4% on the announcement, reaching an intraday high of $214.97 before settling near $207 2 . AMD, whose stock has already more than doubled in 2026, traded near $553 — about 14% off its June 30 high of $584.73 . Wall Street analysts have raised AMD price targets into the $600-$700 range heading into the Advancing AI event 3 https://www.fxleaders.com/news/2026/07/22/amd-stock-analysis-553-at-advancing-ai-2026-helios-pricing-and-the-700-target/ . 3 https://www.fxleaders.com/news/2026/07/22/amd-stock-analysis-553-at-advancing-ai-2026-helios-pricing-and-the-700-target/ The partnership is notable for what it signals about the AI inference market's trajectory: as models grow more complex and latency-sensitive applications like autonomous agents and robotics proliferate, the companies are betting that heterogeneous, disaggregated architectures will outperform monolithic GPU-only solutions 1 . What's Next Cerebras said it will begin deploying AMD Helios systems in its data centers, with the joint solution initially available through Cerebras Cloud in H2 2026 1 . No pricing for the combined inference service has been disclosed. The companies did not specify whether the partnership includes any financial terms, exclusivity provisions, or minimum purchase commitments. The deal's commercial success will hinge on whether the 5x efficiency claim holds across real-world workloads and whether enterprise customers adopt disaggregated inference architectures over the integrated rack solutions offered by Nvidia and others. Companies mentioned Further sources 1 AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughpu… ↗ https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution 2 Cerebras stock gains on AMD partnership — CNBC ↗ https://www.cnbc.com/2026/07/23/cerebras-stock-gains-on-amd-partnership.html 3 AMD Stock Analysis: $553 at Advancing AI 2026; Helios Pricing and the $700 Targ… ↗ https://www.fxleaders.com/news/2026/07/22/amd-stock-analysis-553-at-advancing-ai-2026-helios-pricing-and-the-700-target/ 4 AMD Helios Rackscale Solution specifications — AMD product page and Next Platfo… ↗ https://www.amd.com/en/products/rackscale-solutions/helios.html 5 OpenAI chip deal with Cerebras adds to roster of Nvidia, AMD, Broadcom — CNBC ↗ https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html The stories that matter, in one email. Free — unsubscribe anytime.