cd /news/ai-infrastructure/vera-rubin-nvl72-is-hitting-3-7x-the… · home topics ai-infrastructure article
[ARTICLE · art-131813] src=promptcube3.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Vera Rubin NVL72 is hitting 3.7x the throughput of GB300 NVL72

NVIDIA's Vera Rubin NVL72 delivered up to 3.7x higher throughput than the GB300 NVL72 on the Qwen3-VL model in MLPerf Inference v6.1 preview submissions, using vLLM with the NVIDIA Dynamo framework, while DeepSeek-R1 throughput rose up to 2.5x with TensorRT-LLM. A 288-GPU submission spanning four racks reached 99% scaling efficiency, aided by sixth-generation NVLink and NVLink Switch reporting 10x higher packet rates and 3x lower latency than standard Ethernet. The v6.1 submissions also showed up to 1.6x higher performance over v6.0, with NVIDIA saying optimizations continued after the submission deadline.

by read2 min views1 publishedSep 16, 2026
Vera Rubin NVL72 is hitting 3.7x the throughput of GB300 NVL72
Image: Promptcube3 (auto-discovered)

The MLPerf Inference v6.1 results just dropped, and the gap between the Gundefined and the new Vera Rubin NVL72 is pretty massive. If you're tracking inference economics, the takeaway is that the Rubin rack is pushing significantly more tokens per unit, which directly lowers the cost per token. The 6.1-0106 and 6.1-0074 entries show that this isn't just a marginal bump; we're talking about a different league of throughput for the same rack footprint.

How the Vera Rubin NVL72 performed in MLPerf v6.1 #

The preview submissions focused on two heavy hitters: DeepSeek-R1 and Qwen3-VL. The performance delta depends on the model and the software stack:

  • Qwen3-VL: Using vLLM with the NVIDIA Dynamo open source framework, the Vera Rubin NVL72 hit up to 3.7x higher throughput than the Gundefined NVL72. This was consistent across interactive, server, and offline scenarios.
  • DeepSeek-R1: Using the TensorRT-LLM library, the throughput increase was up to 2.5x compared to the Gundefined NVL72.

It's also worth noting that software velocity is playing a huge role here. The v6.1 submissions showed up to 1.6x higher performance over v6.0, and NVIDIA claims optimizations continued even after the submission deadline.

Scaling efficiency and hardware bottlenecks #

One of the biggest pain points in scaling inference is the "scaling wall" where adding more GPUs doesn't actually result in linear throughput gains. However, the Gundefined NVL72 seems to handle this well. In a 288-GPU submission across four racks, they hit 99% scaling efficiency. That means throughput grew almost linearly from the single-rack baseline, which is rare when you're dealing with that much hardware.

The secret sauce here seems to be the sixth-generation NVLink and NVLink Switch. They're reporting 10x higher packet rates and 3x lower latency than standard Ethernet. Without that interconnect, the disaggregated serving they used for the Rubin submissions—separating prefill and decode stages—wouldn't be nearly as effective.

The technical drivers behind the throughput #

To get these numbers, NVIDIA is leaning hard into full-stack codesign. A few specific technical details stand out:

  • NVFP4 Precision: They're using this to shrink the memory footprint for model weights, the KV cache, and attention, which boosts throughput without killing the output quality.
  • Expert Parallelism: Since DeepSeek-R1 and Qwen3-VL rely on mixture-of-experts (MoE) layers, the Rubin system uses large-scale expert parallelism to keep efficiency high.
  • Hardware acceleration: The updated Transformer Engine and Tensor Cores are specifically targeting the prefill and decode stages to remove bottlenecks.

Next Which of these DEV community posts actually hit the mark this week? →

All Replies (3) #

I want to try this tonight. My current H100 cluster is choking on Llama 3.1 405B, but maybe 3.7x fixes it?

Curious if this is real world or just synthetic. Does anyone know if this holds up with vLLM or just TensorRT?

Stressed about my current budget, so that jump is huge. I'm wondering if it helps with Triton kernels specifically...

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/vera-rubin-nvl72-is-…] indexed:0 read:2min 2026-09-16 ·