# Pretraining Transformers with Quantized Softmax in Attention

> Source: <https://aiflash.com/news/129048/>
> Published: 2026-09-30 05:00:17+00:00

Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interv
