20:05
2026-10-08
discuss.huggingface.co
large-language-models
Keeping the Attention Sink Exact in Activation-Whitened SVD Compression
A zero-cost reweighting technique that keeps the attention sink at token position 0 exact during activation-whitened SVD compression lowered the loss increase in 18 of 20 model-and-budget cells acrossβ¦