{"slug": "keeping-the-attention-sink-exact-in-activation-whitened-svd-compression", "title": "Keeping the Attention Sink Exact in Activation-Whitened SVD Compression", "summary": "A zero-cost reweighting technique that keeps the attention sink at token position 0 exact during activation-whitened SVD compression lowered the loss increase in 18 of 20 model-and-budget cells across six model families up to 32B and raised zero-shot accuracy in 21 of 28 configurations, according to the method's author. The author reports the calibration covariance used by SVD-LLM and KFAC-SVD weights every token position equally, giving the position-0 attention sink a weight of about 0.001 even though that position carries most of the gradient energy in key and value projections. The author states they have stopped trusting perplexity for this class of method, citing cases such as Qwen3-4B at 20% compression where loss worsened slightly while accuracy gained two points, and released code at github.com/aemonalgiz-dev/sink_deflation.", "body_md": "There seem to be some odd relationships between perplexity and accuracy for whitened-SVD models, and I found a near-no-cost algorithm to improve most whitened-SVD compressions along the way.\n\nThe short version: the calibration covariance that SVD-LLM and KFAC-SVD whiten against treats every token position the same, so the attention sink at position 0 gets a weight of about 0.001. In the key and value projections, that position carries most of the gradient energy, and standard truncation reproduces it badly while fitting everything else well. Giving the sink its proper weight, or simply constraining the factors to be exact on it (no tuning parameter, same storage, same inference cost), lowers the loss increase in 18 of 20 model-and-budget cells across six families up to 32B, and moves zero-shot accuracy up in 21 of 28 configurations, significantly so on Qwen3-4B, Qwen3-32B and Llama-3.1-8B.\n\nThe odd part is how weakly the two move together. Loss improves on Qwen3-1.7B, and accuracy does nothing; loss gets slightly worse on Qwen3-4B at 20%, and accuracy gains two points; on top of LEMS allocation, perplexity ties while accuracy still improves; and there is a case where standard truncation beats the dense model on perplexity by damaging the sink and loses accuracy doing it. I have stopped trusting perplexity for this class of method, and I am curious whether others see the same.\n\nGitHub Code: [GitHub - aemonalgiz-dev/sink_deflation: A zero cost reweighting for Whitened-SVD LLMs · GitHub](https://github.com/aemonalgiz-dev/sink_deflation)", "url": "https://wpnews.pro/news/keeping-the-attention-sink-exact-in-activation-whitened-svd-compression", "canonical_source": "https://discuss.huggingface.co/t/keeping-the-attention-sink-exact-in-activation-whitened-svd-compression/190159#post_1", "published_at": "2026-10-08 20:05:42+00:00", "updated_at": "2026-10-08 20:17:40.315687+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research"], "entities": ["SVD-LLM", "KFAC-SVD", "Qwen3-4B", "Qwen3-32B", "Qwen3-1.7B", "Llama-3.1-8B", "LEMS", "sink_deflation"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/keeping-the-attention-sink-exact-in-activation-whitened-svd-compression", "markdown": "https://wpnews.pro/news/keeping-the-attention-sink-exact-in-activation-whitened-svd-compression.md", "text": "https://wpnews.pro/news/keeping-the-attention-sink-exact-in-activation-whitened-svd-compression.txt", "jsonld": "https://wpnews.pro/news/keeping-the-attention-sink-exact-in-activation-whitened-svd-compression.jsonld"}}