cd /news/artificial-intelligence/nvidia-shares-rubin-gpu-deep-dive-an… · home topics artificial-intelligence article
[ARTICLE · art-67480] src=techpowerup.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

NVIDIA Shares "Rubin" GPU Deep-Dive and Die Annotation

NVIDIA detailed its new "Rubin" GPU architecture, which delivers 50 PetaFLOPS at sparse NVFP4 operations, features 224 streaming multiprocessors, 288 GB of HBM4 memory, 336 billion transistors, and 896 Tensor Cores with a third-generation Transformer Engine. The GPU includes NVLink 6 with 3,600 GB/s fabric bandwidth and NVLink-C2C at 1,800 GB/s, along with a PCIe Gen 6 switch at 256 GB/s.

read1 min views1 publishedJul 21, 2026

NVIDIA's latest "Rubin" GPU architecture represents one of the company's most significant advancements in scaling computing across multiple racks, systems, and data centers. The company today showcased the underpinning of an accelerator capable of delivering 50 PetaFLOPS of compute power at sparse NVFP4 operations. In its full configuration, NVIDIA's "Rubin" GPU includes up to 224 streaming multiprocessors (SMs) and 288 GB of HBM4 memory. With 336 billion transistors, NVIDIA has also integrated 896 Tensor Cores with a third-generation Transformer Engine. To achieve massive raw compute power, NVIDIA organizes GPU resources into Graphics Processor Clusters (GPCs) featuring a centralized L2 cache. Working alongside the GPCs is the GigaThread Engine, which coordinates workflows and optimizes GPU resource utilization. There are also MIG Control partitions that enable the GPU to be divided into multiple virtual GPUs, similar to how modern CPUs handle virtual cores.

On the memory side, NVIDIA has collaborated with its memory partners to provide 288 GB of HBM4 memory using 12-high modules. Dedicated HBM controllers and physical layers utilize HBM4 with up to 22 TB/s peak memory bandwidth, which will be competitive across the industry. The Tensor Memory Accelerator has been updated to enhance data movement efficiency, while scaling up and communication to the entire GPU is managed by NVLink 6. With NVLink 6, all-to-all GPU communications operate on 3,600 GB/s of fabric bandwidth, while NVLink-C2C chip-to-chip enables CPU-GPU communication at a bandwidth of 1,800 GB/s. NVIDIA also includes a PCIe Gen 6 switch with 256 GB/s bandwidth for communication with external devices.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nvidia-shares-rubin-…] indexed:0 read:1min 2026-07-21 ·