NVIDIA Nemotron TwoTower: Run LLMs 2.42x Faster Now
NVIDIA open-sourced Nemotron-Labs-TwoTower, a diffusion language model that generates text 2.42x faster than its autoregressive counterpart without retraining original weights. The model achieves 98.7% quality retention …