Keeping the Attention Sink Exact in Activation-Whitened SVD Compression
A zero-cost reweighting technique that keeps the attention sink at token position 0 exact during activation-whitened SVD compression lowered the loss increase in 18 of 20 model-and-budget cells across…