NVIDIA's latest "Rubin" GPU architecture represents one of the company's most significant advancements in scaling computing across multiple racks, systems, and data centers. The company today showcased the underpinning of an accelerator capable of delivering 50 PetaFLOPS of compute power at sparse NVFP4 operations. In its full configuration, NVIDIA's "Rubin" GPU includes up to 224 streaming multiprocessors (SMs) and 288 GB of HBM4 memory. With 336 billion transistors, NVIDIA has also integrated 896 Tensor Cores with a third-generation Transformer Engine. To achieve massive raw compute power, NVIDIA organizes GPU resources into Graphics Processor Clusters (GPCs) featuring a centralized L2 cache. Working alongside the GPCs is the GigaThread Engine, which coordinates workflows and optimizes GPU resource utilization. There are also MIG Control partitions that enable the GPU to be divided into multiple virtual GPUs, similar to how modern CPUs handle virtual cores.
On the memory side, NVIDIA has collaborated with its memory partners to provide 288 GB of HBM4 memory using 12-high modules. Dedicated HBM controllers and physical layers utilize HBM4 with up to 22 TB/s peak memory bandwidth, which will be competitive across the industry. The Tensor Memory Accelerator has been updated to enhance data movement efficiency, while scaling up and communication to the entire GPU is managed by NVLink 6. With NVLink 6, all-to-all GPU communications operate on 3,600 GB/s of fabric bandwidth, while NVLink-C2C chip-to-chip enables CPU-GPU communication at a bandwidth of 1,800 GB/s. NVIDIA also includes a PCIe Gen 6 switch with 256 GB/s bandwidth for communication with external devices.