NVIDIA's "Rubin CPX" GPU Reborn with HBM4 Memory, No More GDDR7 NVIDIA has revived its 'Rubin CPX' AI accelerator, now featuring 168 GB of HBM4 memory instead of the previously planned 128 GB of GDDR7, according to supply chain analyst Ming Chi Kuo. The GPU, dedicated to prefill workloads for large language models, will be housed in racks with 64 to 256 standalone CPX GPUs, with each eight-GPU tray handling 1.34 TB of long-context prefill and KV cache. The redesign likely uses TSMC's CoWoS-S or CoWoS-L packaging. NVIDIA has reportedly revived its "Rubin CPX" AI accelerator project, which had previously been put on hold. Today, well-known supply chain analyst Ming Chi Kuo reported that NVIDIA is reintroducing the "Rubin CPX," now featuring HBM4 memory instead of the previously planned GDDR7. Late last year, NVIDIA announced https://www.techpowerup.com/340818/nvidia-unveils-rubin-cpx-gpu-single-die-30-petaflops-and-128-gb-of-gddr7-memory its plans to develop a specialized accelerator derived from the "Rubin" GPU family, designed for large-scale agentic AI workloads. After the announcement, the project was reportedly paused, but NVIDIA has now redesigned the memory architecture. Previously, the "Rubin CPX" was slated to have 128 GB of GDDR7 memory, but it will now include 168 GB of HBM4 memory. This GPU SKU is housed in a separate rack with configurations ranging from 64 to 256 standalone CPX GPUs per rack. This GPU is dedicated to prefill workloads, powering racks that manage input context for LLMs and KV cache processing. Essentially, NVIDIA is separating prefill from decode tasks, with CPX GPUs handling prefill and regular "Rubin" GPUs managing decode. A rack tray with eight CPX GPUs can handle 1.34 TB of long-context prefill and the associated KV cache, all utilizing HBM4 memory for faster processing and increased bandwidth. Since the previous design did not require HBM4 integration, NVIDIA has redesigned the package, likely using TSMC's CoWoS-S or CoWoS-L for packaging.