NVIDIA Groq 3 LPX Enters Full Production: 3,400 Tokens per Second at 100K Context, 256 LP30s per Rack
NVIDIA announced the production release of the NVIDIA Groq 3 LPX, a dedicated interactive inference accelerator for the Vera Rubin NVL72 platform, achieving 3,400 tokens per second on the Gemma 4 31B …