# AMD challenges Nvidia’s networking dominance with Helios racks boasting 50% higher bandwidth

> Source: <https://www.sdxcentral.com/news/amd-challenges-nvidias-networking-dominance-with-helios-racks-boasting-50-higher-bandwidth/>
> Published: 2026-07-23 16:45:00+00:00

SAN FRANCISCO – AMD finally showcased the networking stack behind its NVL72 rival Helios, unveiling a programmable front-end, a single hopscale-up fabric, and high-bandwidth, protocol-flexible AI network interface cards (NICs).

The firm’s Advancing AI event saw it reveal concrete stats behind its rack-scale platform that it [teased last year](https://www.sdxcentral.com/news/amd-debuts-helios-open-rack-system-for-next-gen-ai-workloads/). Anthropic joined a growing list of AI labs, neoclouds, and hyperscalers signed on to deploy Helios, looking to make full use of densely packed pods tied together by a high‑bandwidth, Ethernet‑based fabric.

At the front-end, Helios leverages AMD’s Salina 400 Gb/s (400G) data processing unit (DPU). The third-generation function accelerator chip is programmable and can be upgraded live without traffic disruption. It provides a security layer for the front-end, with users able to encrypt all traffic while also offloading infrastructure tasks from the central processing unit (CPU) to free up valuable cores.

The scale-up fabric inside a Helios pod leans on UALink over Ethernet (UALoE), pooling 72 MI455X graphic processing units (GPUs) to form what Krishna Doddapaneni, corporate VP for AMD Pensando, described as “a humungous amount of bandwidth” – 260 Tb/s per pod.

The scale-up design encompasses a single-tier one-hop Ethernet fabric, supporting 31 terabytes of high-bandwidth memory four (HBM4). AMD also tapped underlying features to double down on fixed latency, congestion avoidance, and resiliency.

Helio’s scale-out – or rack-to-rack network – is where it really begins to take on Nvidia. Built on the second-generation Pensando AI NIC, dubbed "Vulcano," it packs three NICs per GPUs, providing a highly dense configuration capable of 43 Tb/s; that’s 50% more bandwidth than Nvidia’s Vera Rubin NVL72, according to AMD.

On the scale-out side, the Helios platform also lends its open-centric and programmability support for a variety of scaling approaches to allow operators to balance network cost and protocol compatibility, including switch-based packet spraying, NIC-based packet spraying, and NIC-based source routing.

## Simplicity in design

The networking underpinning Helios sees AMD banking on simplicity, functionality, and performance.

The scale-up UALoE’s single-tier design, for example, sees the vendor look to reduce key value cache issues, and instead every GPU has a fixed latency when going through a switch because it's a single hop.

AMD also make use of what Doddapaneni described as “the simplest set of features in Ethernet,” such as layer-two static media access control (MAC) programming that dictates which interface should receive traffic destined for a specific device, a concept the exec reminded has “existed [for] like 20 years,” along with flow control and “all the telemetry features that typical switch silicon provides.”

This deliberately simple stack, the VP suggests, helps cluster operators avoid multitier congestion and variable latency to ensure a 72‑GPU pod behaves like one big, predictable HBM pool rather than a fragmented, incoherent mass of nodes.

Orchestrating the entire "keep it simple, stupid" approach is rack management software built specifically for Helios. The aptly named AMD Fabric Manager (AFM) runs a Kubernetes-based, highly available controller cluster directly on the silicon found in the platform’s six switch trays. It handles zero‑touch bring-up, wiring validation, fabric layout, and failure remediation, while the switches themselves run the open SONiC network operating system (NOS), a NOS that’s effectively hidden behind AFM so operators manage the rack as a single fabric, not as a fleet of individual switches.

“The NOS is kind of hidden from the user," Doddapaneni said. "[A] user doesn't need to know about it; it's all hidden by AFM. But if they wanted to log into that for telemetry reasons, they can. But from the user's point of view, NOS is not a visible construct.”

The networking VP described the underlying software component providing insights into the platform as a “single pane of glass,” adding: “You can manage the whole fabric using AFM. It tracks the GPU utilization, it tracks the compute utilization, it tracks the network failures [and network health], and it also gives you an idea of when some faults happen, how we are remediating, and also gives events and alerts for operators to manage this fabric.”

## DPUs in the spotlight

Speaking in an earlier press briefing, Soni Jiandani, SVP and GM of AMD’s networking technology and solutions group, contended that networking today was as critical as the GPU itself. The exec argued that the shift from large language model (LLM) training to inferencing and agentic workflows means that data movement and latency on the network are core determinants of throughput, GPU utilization, and cost per token.

“While agentic AI is driving even more interactions between the compute, memory, and storage, the common denominator in all of this is data movement at scale, where the network is fundamental to deliver the performance in a deterministic manner and have the ability to do it at a system level, not just looking at the performance of compute alone and on its own,” Jiandani added.

The rise of agentic AI is fundamentally putting the DPU front and center, the exec contended, with Helios taking full advantage of Pensando technology AMD picked up [back in 2022](https://www.sdxcentral.com/analysis/why-amd-spent-19b-for-pensandos-dpu-biz/).

To emphasize this, Jiandani brought on stage Raj Subramaniyan, SVP and head of networking at Oracle Cloud Infrastructure. The long-time customer framed DPUs as a foundation, not just a plumbing nice-to-have, and contended the function accelerator card also aids security, storage acceleration, and tenant/dependency isolation.

Subramaniyan outlined that OCI makes use of a converged DPU architecture, in which the chips handle storage, security, and network virtualization acceleration on its own, enabling disaggregated storage, KV cache expansion, and stricter isolation while significantly boosting SDN performance – having made use of AMD Pensando cards across “multiple generations.”

“As we progress toward agentic AI, now the data movement becomes critical. So now we are seeing a lot of storage access required, disaggregated storage access, key value (KV) cache expansions, and these kinds of things have to come in,” Subramaniyan explained. “It becomes even more important for us to keep loading onto the DPUs, add functionalities, offload many of these functions, and accelerate them.”
