Via nvidia.com
The tech giant becomes one of the earliest adopters of Nvidia's next-generation AI architecture, targeting massive performance gains over the previous Blackwell platform
Microsoft’s data centers have taken delivery of the first production Vera Rubin systems from Nvidia, marking a pivotal moment in the next chapter of the AI hardware arms race. CEO Satya Nadella confirmed the milestone on August 21, 2026, placing Microsoft among the very first companies to get its hands on what Nvidia is billing as a generational leap in AI computing.
What’s inside the Vera Rubin NVL72 #
The Vera Rubin NVL72 architecture is not a modest refresh. Each system pairs 88-core Vera CPUs with Rubin GPUs packing up to 288 GB of HBM4 memory per chip. That memory figure alone represents a substantial jump, giving models far more room to breathe during training runs that previously required aggressive optimization just to fit into GPU memory.
The systems also incorporate NVLink networking and liquid cooling solutions, addressing two of the biggest bottlenecks in modern AI infrastructure: data movement between chips and the sheer thermal output of running thousands of GPUs at full tilt.
Nvidia’s performance claims are aggressive but specific. The company says Vera Rubin delivers up to 5x faster inference and 3.5x improved training capabilities compared to Blackwell. If those numbers hold up in production workloads, the cost-per-token economics of running large language models could shift dramatically, making previously expensive AI applications suddenly viable at scale.
CoreWeave, another early adopter, has reported a 10x increase in AI throughput per megawatt using the new systems.
Microsoft’s broader AI infrastructure play #
Microsoft is targeting specialized deployments at next-generation superfactory sites planned for Wisconsin and Atlanta, purpose-built facilities designed around the thermal and power requirements of hardware like Vera Rubin.
Microsoft joins a short list of early Vera Rubin adopters that also includes Google Cloud, CoreWeave, and OpenAI.
Nvidia formally moved the Vera Rubin platform into full production by June 2026, making the timeline from production start to first customer deployment roughly two months.
The competitive landscape shifts again #
Amazon Web Services, conspicuously absent from early Vera Rubin announcements, may be leaning more heavily on its own custom silicon efforts with Trainium chips.
Nvidia itself benefits from the spectacle of major customers racing to deploy its newest platform. The company has been articulating a vision of trillion-GPU architectures designed for what it calls “agentic AI factories,” essentially autonomous systems that can reason, plan, and execute complex tasks without constant human oversight. Vera Rubin is positioned as a building block toward that vision.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our