# Nvidia’s ultra-low-latency AI inference LPX racks hit full production

> Source: <https://www.sdxcentral.com/news/nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production/>
> Published: 2026-08-25 08:41:03+00:00

Nvidia’s LPX racks for AI inference accelerators have entered full production, the company confirmed.

Unveiled at GTC [back in March](https://www.sdxcentral.com/news/nvidia-bets-big-on-bandwidth-with-groq-3-lpu-to-complement-gpus/), the rack-scale platform came about following Nvidia’s [acqui-hire](https://www.datacenterdynamics.com/en/news/nvidia-to-license-tech-from-ai-inference-chip-company-groq-hire-its-leadership/) of the eponymous startup. LPX is liquid-cooled and houses some 256 Groq 3 language processing units (LPUs) interconnected through some 640 terabits per second (Tb/s) of scale-up bandwidth.

Inside the LPX rack itself are [BlueField-4 data processing units](https://www.sdxcentral.com/analysis/nvidias-bluefield-4-a-first-look-at-the-dpu-built-to-run-ai-factories/), [Vera central processing unit (CPU)](https://www.sdxcentral.com/news/nvidia-vera-cpu-enters-full-production-pitched-at-agentic-ai-workloads/) racks, and [STX storage servers](https://www.sdxcentral.com/news/nvidia-unveils-bluefield-4-stx-architecture-to-power-next-gen-ai-storage/) all tied together with the recently debuted [Spectrum-6 Ethernet](https://www.sdxcentral.com/news/nvidia-debuts-spectrum-6-switches-to-power-next-gen-ai-networking/) networking tech.

The platform is not a replacement for Nvidia’s flagship NVL72 platform, but rather a complementary add-on for operators wanting to power ultra-low-latency AI inference workloads.

During the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second. The chip giant claims it can be used to drastically reduce the time it takes to perform agentic-related tasks from hours to mere minutes, offering 4-times faster responsiveness for agents and latency-sensitive workloads.

Nvidia founder and CEO Jensen Huang said the LPX will “transform how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness.”

“Inference is the growth engine of AI. Nvidia Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” Huang said. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”

The LPX platform is expected to be available later this year.

Among its early adopters is neocloud darling Nebius, which plans to bring the Groq 3 LPX platform into its Token Factory offering. Danila Shtan, Nebius’ chief technology officer, said the move will make “every step of an agent’s loop feel instant.”
