cd /news/large-language-models/keeping-the-attention-sink-exact-in-… · home › topics › large-language-models › article
[ARTICLE · art-147815] src=discuss.huggingface.co ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Keeping the Attention Sink Exact in Activation-Whitened SVD Compression

A zero-cost reweighting technique that keeps the attention sink at token position 0 exact during activation-whitened SVD compression lowered the loss increase in 18 of 20 model-and-budget cells across six model families up to 32B and raised zero-shot accuracy in 21 of 28 configurations, according to the method's author. The author reports the calibration covariance used by SVD-LLM and KFAC-SVD weights every token position equally, giving the position-0 attention sink a weight of about 0.001 even though that position carries most of the gradient energy in key and value projections. The author states they have stopped trusting perplexity for this class of method, citing cases such as Qwen3-4B at 20% compression where loss worsened slightly while accuracy gained two points, and released code at github.com/aemonalgiz-dev/sink_deflation.

read1 min views2 publishedOct 8, 2026

There seem to be some odd relationships between perplexity and accuracy for whitened-SVD models, and I found a near-no-cost algorithm to improve most whitened-SVD compressions along the way.

The short version: the calibration covariance that SVD-LLM and KFAC-SVD whiten against treats every token position the same, so the attention sink at position 0 gets a weight of about 0.001. In the key and value projections, that position carries most of the gradient energy, and standard truncation reproduces it badly while fitting everything else well. Giving the sink its proper weight, or simply constraining the factors to be exact on it (no tuning parameter, same storage, same inference cost), lowers the loss increase in 18 of 20 model-and-budget cells across six families up to 32B, and moves zero-shot accuracy up in 21 of 28 configurations, significantly so on Qwen3-4B, Qwen3-32B and Llama-3.1-8B.

The odd part is how weakly the two move together. Loss improves on Qwen3-1.7B, and accuracy does nothing; loss gets slightly worse on Qwen3-4B at 20%, and accuracy gains two points; on top of LEMS allocation, perplexity ties while accuracy still improves; and there is a case where standard truncation beats the dense model on perplexity by damaging the sink and loses accuracy doing it. I have stopped trusting perplexity for this class of method, and I am curious whether others see the same.

GitHub Code: GitHub - aemonalgiz-dev/sink_deflation: A zero cost reweighting for Whitened-SVD LLMs · GitHub

── more in #large-language-models 4 stories · sorted by recency
── more on @svd-llm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/keeping-the-attentio…] indexed:0 read:1min 2026-10-08 · —