# Softmax Reparameterization for Output-Head Quantization

> Source: <https://aiflash.com/news/127850/>
> Published: 2026-09-28 14:30:00+00:00

Large vocabularies make output heads a substantial inference cost in small language models. We propose softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from eve
