cd /news/artificial-intelligence/microsoft-taps-amd-for-at-scale-ai-c… · home topics artificial-intelligence article
[ARTICLE · art-65838] src=nextplatform.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Microsoft Taps AMD For At Scale AI CPU And GPU Clusters

Microsoft announced a partnership with AMD to deploy large-scale AI clusters based on AMD's Helios rack design, featuring MI455X GPUs and Venice Epyc CPUs, for inference workloads on Azure. The deployment, expected to cost between $5 billion and $10 billion, includes hundreds of thousands of GPUs and positions AMD as a key alternative to Nvidia in the AI hardware market.

read4 min views1 publishedJul 20, 2026
Microsoft Taps AMD For At Scale AI CPU And GPU Clusters
Image: Nextplatform (auto-discovered)

We live in a weird time. Every GPU and every CPU that AMD and Nvidia can make for the rest of 2026 has already been long since sold even if they have not all been manufactured as yet. And given that, if either company just wanted to stop doing product launches, they could just do that.

All we would hear would be crickets. . . .

But, alas, all of the free press and analysis that comes from major product reveals, such as AMD will be doing at its Advancing AI 2026 in Silicon Valley this Thursday, is a great thing for any IT supplier. And with GPU accelerated systems or all-CPU clusters, AMD is just about the only company applying direct pressure on GenAI hardware behemoth Nvidia in the two key places it is minting piles of coin stretching to the Moon and back.

Ahead of that Advancing AI event, Microsoft and AMD decided to announce that the two were teaming up for large scale AI systems based on the AMD “Helios” rack design and including a slew of hardware and software technologies from AMD, including the its “Altair” MI455X GPUs, its “Venice” Epyc 9006 CPUs, its Pensando DPUs, and the ROCm software stack.

The double-wide Helios rack has 4,600 Zen 6 CPU cores across 18 compute trays with one CPU and four GPUs each. That works out to 256 cores per Venice processor, and it also works out to 18,000 GPU compute units across 72 GPUs per rack, with the GPUs delivering 2.9 exaflops at FP4 precision. The GPUs have a combined 31 TB of HBM 4 stacked memory and deliver 43 TB/sec of aggregate bandwidth through the Pensando DPUs, which are programmable on the P4 language. This is a key feature, and one that Microsoft took a shining to as it has been perhaps the dominant deployer of Pensando DPUs to date.

Neither Microsoft nor AMD are talking about how much money is being ponied up by Big Bill for its Helios racks, but we expect it to be for hundreds of thousands of GPUs and between $5 billion and $10 billion. What we do know for sure is that this is a very large scale deployment and that it is very specifically targeted at running inference for frontier models on the Azure cloud.

The scale out networking for the Helios racks is no doubt Ethernet, and it is very likely that Arista Networks is the supplier of the switches for this part of the network.

For the scale up network inside the rack and gluing together the HBM memory of the MI455X GPUs, the story is murkier. Both Microsoft and Meta Platforms are enthusiastic supporters of the ESUN coherent memory fabric protocol that is emerging as an alternative to Nvidia’s NVLink/NVSwitch combo as well as the UALink protocol that was launched by AMD and friends back in May 2024. (We did a detailed piece on UALink the following April, and also talked to Upscale AI in January of this year about its desire to launch the first high radix, high bandwidth UALink Switch.) Microsoft may be starting with ESUN running on low latency Ethernet switches in its initial Helios racks and then moving over to UALink switches further down the road. But as I said, this is unclear and will remain so until Microsoft decides to brag about whatever it is doing. Perhaps we will get some more insight at the Advancing AI event.

In addition to this large scale Helios cluster rollout, Microsoft has committed to rolling out clusters of Venice Epyc CPUs that will underpin in two different instances on the Azure cloud. The Azure HDv2 instances will be aimed at agentic AI workloads and their data pipeline processing, while the HXv2 instances (presumably with more cores) will be aimed at electronic design automation for chips.

Additionally, Microsoft is working with AMD to move its Azure Boost acceleration software to the Pensando DPUs. Microsoft had created its own DPU for the initial deployment of Azure Boost back in November 2024, and now it looks like it is porting its homegrown networking and storage virtualization to P4 routines and the Pensando DPUs.

Having said all of this, we still believe that the distribution of GPU acquisitions between Nvidia and AMD at Microsoft is probably on the order of 70 percent to 30 percent or 75 percent to 25 percent. It is all about who gets what allocations of HBM memory, and when. If Nvidia has 80 percent of the HBM, it sells 80 percent of the XPUs, more or less. It really is that simple. But we do think that AMD’s share of the AI training and inference pie is growing, and that must mean its allocations for HBM are also growing. They better be, or there are going to be some massive lawsuits about antitrust and restraint of trade.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @microsoft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/microsoft-taps-amd-f…] indexed:0 read:4min 2026-07-20 ·