Cerebras and AMD partner to build the world’s fastest disaggregated AI inference solution
Disaggregated AI inference is proving to be more than a complementary answer to the prefill and decode bottleneck slowing enterprise AI at scale, and Cerebras and AMD just announced a partnership to build the fastest version of it in the world.
The recent collaboration pairs AMD’s Helios rack-scale architecture for the compute-intensive pre-fill phase with the Cerebras Wafer-Scale Engine for ultra-low-latency decode, according to Julie Choi (pictured), chief marketing officer at Cerebras. The resulting combination delivers 5x higher tokens per second per watt compared to existing solutions. Later this year, Cerebras will bring AMD Helios systems into its own data centers to power the pre-fill layer of the production deployment.
“Lisa and Andrew both announced how AMD and Cerebras are collaborating on the world’s most powerful disaggregated inference solution,” Choi said. “It’s a one plus one equals five X in this case.”
Choi spoke with theCUBE’s John Furrier at the Neo4j GraphTalk event during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the technical architecture behind disaggregated AI inference, along with what workloads are driving the fastest demand and why the partnership extends to deploying AMD Helios inside Cerebras data centers before the end of 2026.
Disaggregated AI inference pairs AMD Helios with Cerebras for maximum throughput and minimum latency
The architectural logic is fairly straightforward. Pre-fill is computationally intensive, but manageable with optimized GPU infrastructure like AMD Helios. Decode is constrained by memory bandwidth. The Cerebras Wafer-Scale Engine carries roughly 2,000 times the memory bandwidth of competing Nvidia GPUs, making it purpose-built for the decode bottleneck that limits large-scale AI inference in production.
“On the decode portion, this is a memory bandwidth constrained problem,” Choi said. “The Cerebras Wafer-Scale Engine has the largest amount of memory bandwidth. 2,000 times more than Nvidia GPUs.”
The workloads driving the most demand for disaggregated AI inference are agentic coding, real-time voice and multimodal generation. These are all categories where response speed is a functional requirement. The 5x throughput gain means the same infrastructure can serve dramatically more concurrent users, with joint go-to-market efforts expected before the end of the year, Choi noted. Cerebras will also bring AMD Helios systems into its own data centers this year to power the pre-fill layer, making the partnership both a product collaboration and a production infrastructure commitment.
“Our vision is to really provide this speed and max intelligence, no trade-off, to every developer on Earth,” Choi said.
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the Neo4j GraphTalk event:
( Disclosure: TheCUBE is a paid media partner for the Neo4j GraphTalk event. Neither Neo4j, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*
Photo: SiliconANGLE
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.
About SiliconANGLE Media
theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.