cd /news/artificial-intelligence/nvidia-doubles-down-on-ai-factories-… · home topics artificial-intelligence article
[ARTICLE · art-67137] src=siliconangle.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains

Nvidia Corp. revealed performance benchmarks for its next-generation Vera Rubin platform, showing a 10x improvement in tokens per watt on DeepSeek's R1 model compared to the previous GB200 NVL72 system, as demonstrated by partner CoreWeave Inc. The company also reported 1.9x faster agentic performance and a six-fold latency improvement over x86-based alternatives for its Vera CPUs, and highlighted infrastructure optimizations enabling 40% more GPUs within the same power envelope.

read4 min views1 publishedJul 21, 2026
Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains
Image: Siliconangle (auto-discovered)

Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains

Artificial intelligence chip king Nvidia Corp. today revealed a fresh trove of performance benchmarks and architectural milestones for its next-generation Vera Rubin platform as it edges closer to global availability.

The new numbers are impressive, but on a higher level they also underscore the potency of Nvidia’s approach that combines custom silicon with its broader hardware ecosystem and aggressive software optimization to squeeze even greater efficiency gains for AI workloads.

The latest milestones showcase the effort the chipmaker has made toward fine-tuning its ecosystem to tackle the increased performance demands and exploding infrastructure costs of next-generation “agentic” AI workloads.

Impressive gains

One of the key revelations came from Nvidia’s partner CoreWeave Inc., the neocloud company that provides rented, cloud-based access to its graphics processing unit platforms. In early production runs, CoreWeave showed that it was able to squeeze out an incredible 10 times more tokens per watt when running DeepSeek Ltd.’s R1 model on the new Vera Rubin platform, compared with the previous generation GB200 NVL72 system based on Blackwell.

It’s an impressive leap, but CoreWeave showed that the optimizations made to Vera Rubin can also be applied to the GB200 NVL72 system as well, improving throughput per megawatt by more than four times over a three-month span. Nvidia said these gains were validated across 250,000-plus distinct configurations and more than 1.4 million GPU hours of testing.

Nvidia also wheeled out some new numbers for its Vera central processing units for running autonomous AI agents that can perform work on behalf of humans with minimal supervision. The Vera CPUs were built on a custom microarchitecture specification code-named Olympus core, and are highly optimized to process the irregular control flows of AI agents. In the latest benchmarks, Nvidia demonstrated 1.9 times faster agentic performance and a six-fold improvement in latency over x86-based alternatives. The company also showed that Vera surpassed Advanced Micro Devices Inc.’s flagship EPYC Turin CPU by almost 100% on selected industry benchmarks, underscoring the importance of custom silicon in eliminating CPU-side bottlenecks.

Infra optimizations

There’s no doubt that Nvidia’s processors are among the best in the business, but the company made it clear that these gains could only be achieved thanks to its strategy that’s focused on hardware and software co-design. The company also develops the required software and networking systems needed to optimize the performance of its chips so as to maximize power and cooling efficiency.

Nvidia revealed that by dynamically optimizing the full infrastructure and energy stack, it could deploy 40% more GPUs within the same power envelope. At the same time, it showed how it can also minimize the environmental impact of its new processors by cooling them in a 45°C closed-loop liquid-cooled system that saves about 4 million gallons of water per megawatt annually over standard cooling methods.

The company also talked about its latest generation of networking fabrics, showing how they continue to outperform the best generic Ethernet architectures. The sixth-generation NVLink 6 interconnect delivered 2.3 times higher simulated decode throughput for massive large language models than Ethernet-based networks.

Meanwhile, the company’s Spectrum-X platform enabled 1.6 times faster remote direct memory access bandwidth with 1.7 times fewer switches. This integrated network results in five times greater optical power efficiency and 10 times more reliability, Nvidia said. In particular, the company revealed, its latest Spectrum-X platform, Spectrum-6, is arriving at AI factories from the likes of CoreWeave, Microsoft Corp., Nebius B.V., SpaceXAI Corp. and Tesla Inc.

Nvidia is currently racing to ship the new Vera Rubin systems to customers and partners, including the likes of Google Cloud, Microsoft Azure, Meta Platforms Inc., Oracle Cloud Infrastructure, Dell Technologies Inc., OpenAI Group PBC and CoreWeave. As those platforms inch closer to general availability, the release of the latest benchmarks is clearly a strategic move by Nvidia, aiming to show that it now has all of the pieces in the puzzle for its customers to build the “AI factories” of the future.

With reporting from Robert Hof

Images: Nvidia

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

About SiliconANGLE Media

SiliconANGLE,

theCUBE Network,

theCUBE Research,

CUBE365,

theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia corp. 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nvidia-doubles-down-…] indexed:0 read:4min 2026-07-21 ·