Low-precision Transformer systems increasingly quantize attention matrix multiplications, while softmax often remains at higher precision. During pretraining, an approximate softmax changes the gradients that train the model as well as its forward computation. We study this interaction with K-interv
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?