18:12
2026-08-17
developer.nvidia.com
artificial-intelligence
Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer
NVIDIA released the Nemotron 3.5 Lightning NVFP4 checkpoint, compressed from 66 GB to 22 GB via 4-bit quantization, achieving up to 4x faster throughput while preserving accuracy. The model was develoβ¦