AMD answers Nvidia’s Vera Rubin with AMD Helios AI platform AMD CEO Lisa Su unveiled the Helios AI platform at a San Francisco event to compete with Nvidia's Vera Rubin, claiming AMD has an edge in efficiency and cost. The platform includes the new Instinct MI455X GPU and Venice family of CPUs, with AMD stating its Venice CPU will have 2.2 times the throughput of Nvidia's Vera CPU. Su said the AI accelerator market is growing at 45% CAGR from $200 billion in 2025 to $1.4 trillion by 2030. Become a member of GB MAX to gain exclusive access to the industry and to the most influential global B2B leadership community in the business of gaming, entertainment, and tech. Join now https://go.gamesbeat.com/gb-max/ and also get a VIP ticket to GamesBeat Next Nov 2-3, SF .Lisa Su, CEO of Advanced Micro Devices, marched out AMD’s Helios platform to compete with Nvidia’s Vera Rubin https://gamesbeat.com/welcome-to-the-house-that-built-vera-rubin-the-next-foundation-of-ai-factories-and-the-modern-economy/ platform for AI computing coming this fall. Speaking to thousands of engineers at the Moscone West convention center in San Francisco, Su tried to convince the crowd that AMD has an edge in efficient and cost over Nvidia, which has become the world’s most valuable company worth almost $5 trillion. She said Helios is in production today. Su said the AI accelerator market is growing at a 45% cagr, from $200 billion the size of the game industry in 2025 to $1.4 trillion the size of the whole chip industry by 2030. Server CPUs will grow to at a 50% CAGR $200 billion by 2030. No one company can service that whole market, Su said. Su also unveiled a new building block for Helios, the AMD Instinct MI455X graphics processing unit GPU , which will serve as the parallel computing part of the system that helps AI crunch enormous amounts of data to arrive at answers to user AI queries. And Su introduced the Venice family of CPUs to orchestrate so the GPUs get enough data; the Zen 6-based Venice CPU is 3.3 times faster in performance than general-purpose CPUs, with up to 256 cores. Its battle with Intel in CPUs is going well; AMD said it has 46% of server CPU revenue share. Now called the AMD 9006X SP7 CPU, the first version of Venice has 96 cores running at up to 5.15GHz. It has 18 times the throughput of AMD’s first Epyc CPU in 2017. Competition with Nvidia AMD’s answer to Nvidia’s Vera Rubin is AMD Helios, which aggregates different parts of AI components into a single hardware platform. Within the Helios platform, AMD’s CPU handles orchestration, the GPU handles compute, the Pensando networking connects the system via high-speed networking, and the ROCm.AI software stack ties everything together. AMD’s strategy is to leverage all of these things together, said Alan Smith, corporate fellow of graphics architecture at AMD in a press briefing. Helios has four AMD Instinct MI455X graphics processing units GPUs . The Pensando networking links that connect 72 GPUs as if they were one GPU, and the CPU host AMD Epyc 9006 SP7 server CPU orchestrates the computing work. AMD executives said they were pleased when Nvidia revealed benchmark tests for the upcoming Vera Rubin platform, as AMD believes they can beat Nvidia on those benchmarks. Nvidia released performance data this week for its Vera CPUs in the Vera Rubin platform coming this fall. And AMD said Venice with up to 256 cores at 600 watts will have 2.2 times the throughput of Vera with 88 cores and 450 watts. And a 96-core AMD Epyc “Venice” will have 1.2 times the per core performance of Vera, AMD said. “We were very happy when Nvidia published their per core performance,” Kuppuswamy said. “We have the world’s best server CPU portfolio for the agentic era. It truly is the best CPU for cloud, enterprise and AI. We are delivering up to 30% more tokens per dollar.” In a surprise on stage, Su announced that AMD would partner with Cerebras on ultra low-latency computing. Andrew Feldman, CEO of Cerebras, came out on stage to offer support for the alliance. Big iron with a lot of density Each server rack has 18 compute trays and six switch trays. Overall, there are 72 GPUs functioning as one rack unit. A rack of servers also supports four Switch trays, which combine the compute, memory and networking of the system. Krishna Doddapaneni, corporate vice president of software development, said in a press briefing that AI demands a new networking foundation as it moves from training to distributed inference to agentic AI. The chips, trays, racks and rows of servers all have redundancy to prevent failures. With this redundancy, users may see bandwidth loss but an application won’t die, said Krishna Doddapaneni, corporate vice president of software development at AMD, in a press briefing. There are 72 GPUs in a rack, and each is connected as if it were right next to the other. Across the rack, the aggregate bandwidth is 260 terabytes a second. Krishna Doddapaneni, corporate vice president of software development, said AI demands a new networking foundation as move from training to distributed inference to agentic AI. It’s fault tolerant by design with the aim of isolating failures. It’s resilient by design, Chubb said. It’s a platform for the future of AI infrastructure, Chubb said. AMD’s networking is dubbed Pensando and its Salina is a data processing unit, or DPU, which goes up against Nvidia’s BlueField DPUs. “The future of AI runs on AMD,” he said. The open standards war The Helios platform and its ecosystem is surprisingly similar to Nvidia’s, but AMD insists that Helios is based on open standards that give customers a choice beyond AMD’s closed system, which focuses on its CUDA programming language, which AI programmers know. But AMD has engineered its ecosystem around ways to abstract around Nvidia’s system so customers don’t have to worry about it, AMD executives said. Madhu Rangarajan, corporate vice president for compute and AI enterprise products, said AMD is getting ready for the world of agentic AI, where AI computers create agents who can automate tasks on behalf of users. In this agentic world, agents do real work, and concurrency is the real scaling metric. One CPU profile does not fit all, CPUs hosting CPUs do not equal agentic, and the winning platform lowers costs at scale, he said. And he said AMD has the best server CPU for the agentic era in 6th Gen AMD Epyc “Venice,” which will be accompanied by Venice X and Verano 72 cores . The main function of the CPU is to make sure the GPU is well fed as it crunches data. Venice is world’s best server CPU for cloud, enterprise and AI, said Ravi Kuppuswamy, AMD senior vice president for compute and enterprise AI solutions, in a press briefing. It is the 6th Gen AMD Epyc. Venice X is coming out early next year. Verano is “Vera No,” Kuppuswamy joked. An Epyc 9006 SP8 Server CPU is 8 to 128 cores. Venice X has 96 cores, 16 channels, and it runs at 5.15 GHz. Verano is an Epyc 9006 LP Server CPU with 72 cores, running at 5 GHz, and it comes out early next year. Two years from now, agents could be doing something very different, Rangarajan said. And yes, we know that, as Open AI CEO Sam Altman revealed that its research AI software violated security by autonomously breaking out of the company and hacking rival Hugging Face to solve a problem. Altman called it an “unprecedented cyber incident.” That probably is poor timing as AMD tries to sell the next generation of its AI chips. “The goal is not to build the biggest GPU anymore,” Smith said. “Although we’ve built a pretty big GPU, the goal is now really to build a system that matches the communication and execution patterns of these AI workloads. The AI models and the infrastructure are now co-evolving. Each generation of infrastructure enables new model capabilities, and each new class of models inspires new innovations in compute, memory, communication, and system architecture. So the next leap in AI performance comes from co-designing models and infrastructure as a unified system.” The aim is to improve the tokens per second coming out of the GPU and lowering the cost. The AMD Instinct MI455X GPU is 1.5 times more memory than the predecessor Instinct MI355X GPU, with peak memory bandwidth at 2.9 times, peak MXFP8 at 4X, peakMXFP6 of 2X and peak MXFP4 at 4X. The aggregate bandwidth of the new GPU is three times that of the prior chip, Smith said. Compared to the previous model, it has 34 times higher token throughput and 18 times lower token costs. To improve efficiency, AMD is using clusters to get control over locality and advanced barriers to do synchronization. Part of the goal is to reduce memory traffic and idle cycles to increase delivered performance. AMD also has new DMA engine structure to accelerate data movement. The DMA front ends and back ends allows a programmer to schedule work and then spread the workload across the GPU. Mark Chubb, corporate vice president of platform architecture, talked about the rack that AMD is building to house multiple Helios trays. With Helios, AMD is creating a rack-scale architecture, where Helios creates one unified 72-GPU system. Each GPU can communicate with other GPU without worrying about the time gap between the GPUs. More open source robots Kirk Saban, corporate vice president of products, software and solutions in AMD’s adaptive & embedded group, said in a press briefing that AMD offsets Nvidia’s focus on proprietary technology by taking an open approach with more open source software and standards. “We hear time and again from our customers they don’t want to be locked in to a single vendor,” Saban said. He said AMD is launching its Kria AI System on Module, powered by X100 Series. It has a focus on open standards in robotics. AMD is also launching the AMD Kria AI Robotics Developer Platform, the AMD Robotics Software Suite as an open software stack, and the AMD partner network. “Customer interest and traction we are having is tremendous,” Saban said. Local AI advances Michael Nordquist, corporate vice president of product marketing in the computing and graphics group said in a press briefing that AMD is seeing benefits from putting more AI in processors used on premises at companies or inside homes to reduce the cost of AI tokens for inference computing. It’s also more secure and private, and you can continue working without a connection. Small AI models are becoming very efficient at limited tasks. And they’re closing the game with big frontier models like Open AI. AMD showed off its further push from inference in the home with its Gorgon Halo platform, which expands unified memory from 128GB to 192GB, extending models in the home for AI from 200B to 300B parameters. In closing, Su said that AMD will have new CPU generations in 2028 code-named Florence and 2030. It will have new GPU generations every year, including MI500 in 2027 and MI600 in 2028, as well as a new version of Helios every year. She said MI500 would be a major step up in performance, with a 2,000-times leap over four years of GPUs.