{"slug": "nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production", "title": "Nvidia’s ultra-low-latency AI inference LPX racks hit full production", "summary": "Nvidia's LPX racks for AI inference accelerators have entered full production, housing 256 Groq 3 language processing units (LPUs) interconnected with 640 terabits per second (Tb/s) of scale-up bandwidth. The platform, which supports 3,400 output tokens per second and offers 4-times faster responsiveness for agentic AI workloads, is expected to be available later this year, with Nebius among early adopters planning to integrate it into its Token Factory offering.", "body_md": "Nvidia’s LPX racks for AI inference accelerators have entered full production, the company confirmed.\n\nUnveiled at GTC [back in March](https://www.sdxcentral.com/news/nvidia-bets-big-on-bandwidth-with-groq-3-lpu-to-complement-gpus/), the rack-scale platform came about following Nvidia’s [acqui-hire](https://www.datacenterdynamics.com/en/news/nvidia-to-license-tech-from-ai-inference-chip-company-groq-hire-its-leadership/) of the eponymous startup. LPX is liquid-cooled and houses some 256 Groq 3 language processing units (LPUs) interconnected through some 640 terabits per second (Tb/s) of scale-up bandwidth.\n\nInside the LPX rack itself are [BlueField-4 data processing units](https://www.sdxcentral.com/analysis/nvidias-bluefield-4-a-first-look-at-the-dpu-built-to-run-ai-factories/), [Vera central processing unit (CPU)](https://www.sdxcentral.com/news/nvidia-vera-cpu-enters-full-production-pitched-at-agentic-ai-workloads/) racks, and [STX storage servers](https://www.sdxcentral.com/news/nvidia-unveils-bluefield-4-stx-architecture-to-power-next-gen-ai-storage/) all tied together with the recently debuted [Spectrum-6 Ethernet](https://www.sdxcentral.com/news/nvidia-debuts-spectrum-6-switches-to-power-next-gen-ai-networking/) networking tech.\n\nThe platform is not a replacement for Nvidia’s flagship NVL72 platform, but rather a complementary add-on for operators wanting to power ultra-low-latency AI inference workloads.\n\nDuring the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second. The chip giant claims it can be used to drastically reduce the time it takes to perform agentic-related tasks from hours to mere minutes, offering 4-times faster responsiveness for agents and latency-sensitive workloads.\n\nNvidia founder and CEO Jensen Huang said the LPX will “transform how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness.”\n\n“Inference is the growth engine of AI. Nvidia Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” Huang said. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”\n\nThe LPX platform is expected to be available later this year.\n\nAmong its early adopters is neocloud darling Nebius, which plans to bring the Groq 3 LPX platform into its Token Factory offering. Danila Shtan, Nebius’ chief technology officer, said the move will make “every step of an agent’s loop feel instant.”", "url": "https://wpnews.pro/news/nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production", "canonical_source": "https://www.sdxcentral.com/news/nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production/", "published_at": "2026-08-25 08:41:03+00:00", "updated_at": "2026-08-25 08:43:16.697656+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-chips", "ai-products"], "entities": ["Nvidia", "Groq", "LPX", "BlueField-4", "Vera CPU", "STX", "Spectrum-6", "Nebius"], "alternates": {"html": "https://wpnews.pro/news/nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production", "markdown": "https://wpnews.pro/news/nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production.md", "text": "https://wpnews.pro/news/nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production.txt", "jsonld": "https://wpnews.pro/news/nvidias-ultra-low-latency-ai-inference-lpx-racks-hit-full-production.jsonld"}}