{"slug": "cerebras-and-amd-partner-to-build-the-worlds-fastest-disaggregated-ai-inference", "title": "Cerebras and AMD partner to build the world’s fastest disaggregated AI inference solution", "summary": "Cerebras and AMD have partnered to build what they claim will be the world's fastest disaggregated AI inference solution, combining AMD's Helios rack-scale architecture for the pre-fill phase with the Cerebras Wafer-Scale Engine for decode. The resulting system delivers 5x higher tokens per second per watt compared to existing solutions, according to Julie Choi, chief marketing officer at Cerebras. Cerebras will deploy AMD Helios systems in its own data centers by the end of 2026 to power the pre-fill layer of production deployments.", "body_md": "### Cerebras and AMD partner to build the world’s fastest disaggregated AI inference solution\n\nDisaggregated AI inference is proving to be more than a complementary answer to the prefill and decode bottleneck slowing enterprise AI at scale, and Cerebras and AMD just announced a partnership to build the fastest version of it in the world.\n\nThe recent collaboration pairs AMD’s Helios rack-scale architecture for the compute-intensive pre-fill phase with the [Cerebras Wafer-Scale Engine](https://cdn.sanity.io/files/e4qjo92p/production/2d7fa58e3b820715664bcf42097e86c05070c161.pdf) for ultra-low-latency decode, according to [Julie Choi](https://www.linkedin.com/in/julieshinchoi/) (pictured), chief marketing officer at Cerebras. The resulting combination delivers 5x higher tokens per second per watt compared to existing solutions. Later this year, Cerebras will bring AMD Helios systems into its own data centers to power the pre-fill layer of the production deployment.\n\n“Lisa and Andrew both announced how AMD and Cerebras are collaborating on the world’s most powerful disaggregated inference solution,” Choi said. “It’s a one plus one equals five X in this case.”\n\nChoi spoke with theCUBE’s [John Furrier](https://www.linkedin.com/in/furrier) at the [Neo4j GraphTalk event](https://www.thecube.net/events/neo4j/thecube-nyse-wired-data-ai-turning-data-into-knowledge-for-autonomous-systems) during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the technical architecture behind disaggregated AI inference, along with what workloads are driving the fastest demand and why the partnership extends to deploying AMD Helios inside Cerebras data centers before the end of 2026.\n\n### Disaggregated AI inference pairs AMD Helios with Cerebras for maximum throughput and minimum latency\n\nThe architectural logic is fairly straightforward. Pre-fill is computationally intensive, but manageable with optimized [GPU infrastructure](https://siliconangle.com/2026/06/15/gpu-infrastructure-management-startup-hydra-host-raises-100m/) like AMD Helios. Decode is constrained by memory bandwidth. The Cerebras Wafer-Scale Engine carries roughly 2,000 times the memory bandwidth of competing Nvidia GPUs, making it purpose-built for the decode bottleneck that limits large-scale AI inference in production.\n\n“On the decode portion, this is a memory bandwidth constrained problem,” Choi said. “The Cerebras Wafer-Scale Engine has the largest amount of memory bandwidth. 2,000 times more than Nvidia GPUs.”\n\nThe workloads driving the most demand for disaggregated AI inference are [agentic coding](https://siliconangle.com/2026/06/29/exclusive-agentic-coding-startup-baz-brings-code-reviews-planning-stage-extends-seed-funding-17m/), real-time voice and multimodal generation. These are all categories where response speed is a functional requirement. The 5x throughput gain means the same infrastructure can serve dramatically more concurrent users, with joint go-to-market efforts expected before the end of the year, Choi noted. Cerebras will also bring AMD Helios systems into its own data centers this year to power the pre-fill layer, making the partnership both a product collaboration and a production infrastructure commitment.\n\n“Our vision is to really provide this speed and max intelligence, no trade-off, to every developer on Earth,” Choi said.\n\nHere’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the [Neo4j GraphTalk event](https://www.thecube.net/events/neo4j/thecube-nyse-wired-data-ai-turning-data-into-knowledge-for-autonomous-systems):\n\n*(* Disclosure: TheCUBE is a paid media partner for the Neo4j GraphTalk event. Neither Neo4j, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*\n\n##### Photo: SiliconANGLE\n\n# A message from John Furrier, co-founder of SiliconANGLE:\n\nSupport our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.\n\n**15M+ viewers of theCUBE videos**, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.\n\n# Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links.\n\n**About SiliconANGLE Media**\n\n[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),\n\n[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),\n\n[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),\n\n[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),\n\n[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.\n\nFounded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.", "url": "https://wpnews.pro/news/cerebras-and-amd-partner-to-build-the-worlds-fastest-disaggregated-ai-inference", "canonical_source": "https://siliconangle.com/2026/07/29/disaggregated-ai-inference-cerebras-amd-amdadvancingai/", "published_at": "2026-07-29 22:09:51+00:00", "updated_at": "2026-07-29 22:22:54.332955+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-chips"], "entities": ["Cerebras", "AMD", "AMD Helios", "Cerebras Wafer-Scale Engine", "Julie Choi", "Nvidia", "Neo4j", "John Furrier"], "alternates": {"html": "https://wpnews.pro/news/cerebras-and-amd-partner-to-build-the-worlds-fastest-disaggregated-ai-inference", "markdown": "https://wpnews.pro/news/cerebras-and-amd-partner-to-build-the-worlds-fastest-disaggregated-ai-inference.md", "text": "https://wpnews.pro/news/cerebras-and-amd-partner-to-build-the-worlds-fastest-disaggregated-ai-inference.txt", "jsonld": "https://wpnews.pro/news/cerebras-and-amd-partner-to-build-the-worlds-fastest-disaggregated-ai-inference.jsonld"}}