16:55
2026-09-10
developer.nvidia.com
ai-infrastructure
How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra
NVIDIA's NIM 2.0.12 optimized serving stack delivered up to 2.5x higher output-token throughput than the open-source baseline when serving the Nemotron 3 Ultra model on four B200 GPUs, reaching 1,997 …