AMD takes a shot at Nvidia by betting on AI's next big shift AMD CEO Lisa Su announced a partnership with chip startup Cerebras to pursue disaggregated inference, splitting AI workloads across different hardware types, as AMD claims its Helios server system delivers up to 30% more inference tokens per dollar than Nvidia's Vera Rubin NVL72 rack. The collaboration, which will bring Helios into Cerebras' data centers later this year, reflects a broader industry shift toward disaggregated inference that analysts at UBS say is already underway. AMD is making a bet about the future of AI: one chip shouldn't rule them all. CEO Lisa Su https://www.businessinsider.com/amd-lisa-su-ceo-nvidia-competition-2025-3 announced Thursday that AMD is teaming up with chip startup Cerebras on a new approach to AI inference https://www.businessinsider.com/nvidia-pushing-into-intel-amd-cpu-turf-with-meta-deal-2026-2 , which is the process of generating responses from AI models. Increasingly, chipmakers are pursuing "disaggregated inference," which splits workloads across different types of hardware. AMD's partnership with Cerebras shows the company is betting big on this approach. Traditionally, the same hardware handled both processing a prompt and generating an answer. AMD argues those are fundamentally different jobs. Helios, its latest server system, is designed to process huge volumes of requests, whereas Cerebras' giant, wafer-sized chip specializes in generating near-instantaneous responses. The partnership will bring Helios into Cerebras' data centers later this year. Demand for chips from companies like AMD, Nvidia, and Broadcom https://www.businessinsider.com/nvidia-amd-broadcom-chipmakers-employee-retention-ai-boom-2025-10 has skyrocketed in the AI boom. Nvidia dominates chip design for AI training, and the competition has intensified as AI companies shift focus from training models to putting them to work. The AMD and Cerebras pact aligns with a broader shift that analysts say is already underway, with UBS writing in June that the limitations of current architectures "are driving a shift toward disaggregated inference." UBS wrote that Nvidia — through its integration of AI hardware startup Groq https://www.businessinsider.com/nvidia-gtc-ai-system-groq-technology-inference-2026-3 — and Amazon Web Services are also pursuing similar setups to improve efficiency and lower costs. That said, UBS wrote that disaggregated inference presents new challenges around "orchestration" — or getting different chips to work together seamlessly. At Advancing AI, AMD unveiled Helios, its latest server system that bundles several types of AI chips, which is its answer to Nvidia's Vera Rubin NVL72 rack. AI labs and cloud giants using AMD's infrastructure include OpenAI, Meta, Microsoft, Oracle, and Anthropic, with which AMD announced a multibillion-dollar infrastructure partnership on Wednesday. AMD also used the event to take direct aim at Nvidia, claiming that Helios delivers up to 30% more inference tokens per dollar than Nvidia's Vera Rubin NVL72 rack. "Every Helios can deliver more performance for the largest models, more capacity for longer context, and the bandwidth to scale across thousands of racks," Su said Thursday at AMD's Advancing AI event. Have a tip? Contact this reporter via email at gweiss@businessinsider.com or Signal at @geoffweiss.25. Use a personal email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing information securely.