AMD, Cerebras partner on joint Helios rack-scale AI inference platform AMD and Cerebras Systems announced a technical partnership to develop a joint Helios rack-scale AI inference platform, integrating Cerebras' Wafer-Scale Engine with AMD's Helios system to deliver up to five-times higher tokens per second per watt. The solution, available through Cerebras Cloud in the second half of 2026, targets ultra-low-latency inference workloads. Cerebras CEO Andrew Feldman said the partnership will bring ultra-fast inference to more customers. SAN FRANCISCO – AMD continued its Helios customer wins with a doozy of a partnership: wafer-scale chip maker Cerebras Systems. Having already secured a tie-up with Anthropic https://www.datacenterdynamics.com/en/news/amd-to-invest-5bn-in-anthropic-with-ai-company-to-deploy-2gw-of-instinct-mi450-gpus/ at the start of the week, the chip giant revealed at this week's Advancing AI event that a technical partnership had been struck with the company that makes giant chips. The deal will see Cerebras deploy AMD Helios rack-scale systems in its data centers. The pair also plan on developing a joint solution that would bring Cerebras’ Wafer-Scale Engine integrated as part of Helios, a combined platform they claim would provide up to five-times higher tokens per second per watt T/s/W . Set to become available initially through Cerebras Cloud in the second half of 2026, the joint offering would see Helios act as the high-throughput rack-scale system needed to process sizable complex requests, while Cerebras’s mammoth SRAM-based hardware https://www.sdxcentral.com/analysis/cerebras-spins-nvidias-groq-tieup-as-proof-its-waferscale-bet-was-right/ would bring the ultra-low-latency performance needed to return tokens in what is essentially real time. The resulting solution was described as being “designed specifically for the ultra-low-latency segment of the inference market, with AMD Helios as the foundation for high-throughput and balanced inference workloads across the data center.” “The demand for ultra-fast inference is growing at an unprecedented pace. Cerebras delivers the world’s fastest, ultra-low-latency inference,” said Andrew Feldman, CEO and co-founder, Cerebras. “Partnering with AMD gives us an incredible opportunity to bring that performance to even more customers.” Cerebras would join an ever-growing list of firms signed on to make use of AMD’s Nvidia NVL72 rival Helios, which already counts Meta, OpenAI, Oracle, and Tata Consultancy Services TCS . Earlier this week, Microsoft https://www.datacenterdynamics.com/en/news/microsoft-to-deploy-amd-helios-rackscale-solution-to-support-inference-workloads-and-azure-services/ put pen to paper on a Helios rack-scale deal, while AMD revealed a $5 billion investment https://www.datacenterdynamics.com/en/news/amd-to-invest-5bn-in-anthropic-with-ai-company-to-deploy-2gw-of-instinct-mi450-gpus/ in Anthropic, with the Claude developer in turn making use of up to two gigawatts of AMD Instinct MI450 Series graphic processing units GPUs via Helios.