NVIDIA Nemotron-Labs TwoTower: 2.42x Faster Inference, No Retraining Required
NVIDIA released the 30B-parameter Nemotron-Labs TwoTower language model on July 1 that achieves 2.42x faster inference than its autoregressive baseline without retraining from scratch, by splitting a β¦