With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA announced that its NVIDIA Groq 3 LPX rack-scale system is now in full production, delivering 3,400 output tokens per second on a Gemma 4 31B benchmark for 100,000-token long-context use cases, …