Nvidia’s dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents
Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the world of AI compute.
The new chip, announced today at Hot Chips 2026, is described as a purpose-built extension to Nvidia’s flagship Vera Rubin data center platform. According to Nvidia, it’s designed to deliver ultra-fast token generation speeds, which are necessary to run highly responsive agentic AI workloads. The chipmaker said the neocloud provider Nebius Group N.V. has already signed on as the first customer to commit to using the new chip.
Inference is the AI industry’s lingo for the process of running fully trained AI models in production, and it’s increasingly focused on autonomous AI agents that can perform tasks on behalf of humans. These agents must be able to do everything from reason and plan, write and execute code, inspect system files and use third-party tools in continuous loops. They can quickly crunch through thousands of tokens across these complex chains, but this sometimes results in massive “decode latency” that can create frustrating delays.
To prevent this from happening, AI data centers need more specialized compute architectures that can disaggregate the enormous context processing from token generation, in order to increase the speed at which AI agents can reason and work.
This is where the Vera Rubin NVL72 rack-scale platform comes in. It’s powered by dozens of Nvidia’s most powerful Vera Rubin graphics processing units, which are designed to handle the large-scale context ingestion and processing. Now, with the addition of Groq 3 LPX, it can offload those decode workloads onto the new accelerators.
According to Nvidia, Groq 3 LPX makes it possible for a full rack-scale deployment to harness up to 256 LP30 accelerators, linked by its ultra-high-bandwidth chip interconnects. It means the LPUs can work in tandem with the GPUs to compute every step in an AI agent’s chain of reasoning, acting as a unified inference engine designed for enterprise scale.
Nvidia said this will help to eliminate the tradeoff between throughput and response times. Groq 3 LPX gives AI agents the ability to scan long context windows, verify data, call third-party tools and iterate on the most complex, multistep tasks in real-time, without creating delays for users. It has the third-party data to back up this claim, too. In benchmark tests performed by Artificial Analysis, Groq 3 LPX was shown to output a record-breaking 3,400 tokens per second when running the open-source Gemma 4 31B agentic model with a 100,000-token context window. The chipmaker reckons this makes it four times more responsive for latency-sensitive workloads compared to rival platforms, which means multistep agentic tasks can get done in minutes instead of hours.
The new chip was built using technology licensed from a smaller chipmaker called Groq Inc. Nvidia paid the startup a stunning $20 billion in December to be allowed to access its tech, and also hired its founder Jonathan Ross and President Sunny Madra as part of that deal. Groq, not to be confused with SpaceX Corp.’s Grok AI model, develops processors specifically focused on inference rather than AI training.
Nvidia founder and Chief Executive Jensen Huang said his company had already revolutionized AI inference performance and efficiency with its Grace Blackwell and NVL72 platforms. “Vera Rubin extends this with workload-optimized AI factory configurations designed for the era of agentic AI,” he said. “We’re advancing the performance frontier with LPX for ultra-fast token generation. This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness.”
Nebius is one of the first companies to agree to deploy the Groq 3 LPX chips. It said it will use them in the Nebius Token Factory, its production inference platform, to provide customers with more extreme token generation speeds for their most responsive agentic applications.
“Generation is the phase of inference that determines how responsive an AI system actually is, and that’s exactly what Groq 3 LPX is built to accelerate,” said Nebius Chief Technology Officer Danila Shtan. “As the first AI cloud to bring it to production via Nebius Token Factory, we’re making sure every step of an agent’s loop feels instant.”
Nvidia had plenty more to say at Hot Chips. It revealed it has signed up SpaceX as its latest flagship customer. The space rocket and AI company plans to build its next-generation AI architecture around the Vera Rubin platform. Specifically, it said it will deploy Nvidia’s Vera central processing units for CPU-intensive orchestration, tool execution and simulation tasks that span terrestrial data centers and orbital satellites.
The chipmaker also showcased several complementary technologies for AI factories, including Spectrum-X Multiplane, which is a new AI-optimized Ethernet architecture that divides server connections into parallel paths to scale enormous clusters of up to 512,000 GPUs. Then there’s Nvidia Scale-In, which is a new infrastructure software platform that offloads and accelerates security, networking and data management from the main host compute nodes, and Nvidia NVLink Fusion, which enables custom CPUs and data processing units to connect to its sixth-generation NVLink rack systems.
Images: Nvidia
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.