cd /news/ai-infrastructure/nvidias-ultra-low-latency-ai-inferen… · home topics ai-infrastructure article
[ARTICLE · art-109890] src=sdxcentral.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Nvidia’s ultra-low-latency AI inference LPX racks hit full production

Nvidia's LPX racks for AI inference accelerators have entered full production, housing 256 Groq 3 language processing units (LPUs) interconnected with 640 terabits per second (Tb/s) of scale-up bandwidth. The platform, which supports 3,400 output tokens per second and offers 4-times faster responsiveness for agentic AI workloads, is expected to be available later this year, with Nebius among early adopters planning to integrate it into its Token Factory offering.

read2 min views2 publishedAug 25, 2026
Nvidia’s ultra-low-latency AI inference LPX racks hit full production
Image: Sdxcentral (auto-discovered)

Nvidia’s LPX racks for AI inference accelerators have entered full production, the company confirmed.

Unveiled at GTC back in March, the rack-scale platform came about following Nvidia’s acqui-hire of the eponymous startup. LPX is liquid-cooled and houses some 256 Groq 3 language processing units (LPUs) interconnected through some 640 terabits per second (Tb/s) of scale-up bandwidth.

Inside the LPX rack itself are BlueField-4 data processing units, Vera central processing unit (CPU) racks, and STX storage servers all tied together with the recently debuted Spectrum-6 Ethernet networking tech.

The platform is not a replacement for Nvidia’s flagship NVL72 platform, but rather a complementary add-on for operators wanting to power ultra-low-latency AI inference workloads.

During the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second. The chip giant claims it can be used to drastically reduce the time it takes to perform agentic-related tasks from hours to mere minutes, offering 4-times faster responsiveness for agents and latency-sensitive workloads.

Nvidia founder and CEO Jensen Huang said the LPX will “transform how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness.”

“Inference is the growth engine of AI. Nvidia Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” Huang said. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”

The LPX platform is expected to be available later this year.

Among its early adopters is neocloud darling Nebius, which plans to bring the Groq 3 LPX platform into its Token Factory offering. Danila Shtan, Nebius’ chief technology officer, said the move will make “every step of an agent’s loop feel instant.”

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nvidias-ultra-low-la…] indexed:0 read:2min 2026-08-25 ·