There seem to be some odd relationships between perplexity and accuracy for whitened-SVD models, and I found a near-no-cost algorithm to improve most whitened-SVD compressions along the way.
The short version: the calibration covariance that SVD-LLM and KFAC-SVD whiten against treats every token position the same, so the attention sink at position 0 gets a weight of about 0.001. In the key and value projections, that position carries most of the gradient energy, and standard truncation reproduces it badly while fitting everything else well. Giving the sink its proper weight, or simply constraining the factors to be exact on it (no tuning parameter, same storage, same inference cost), lowers the loss increase in 18 of 20 model-and-budget cells across six families up to 32B, and moves zero-shot accuracy up in 21 of 28 configurations, significantly so on Qwen3-4B, Qwen3-32B and Llama-3.1-8B.
The odd part is how weakly the two move together. Loss improves on Qwen3-1.7B, and accuracy does nothing; loss gets slightly worse on Qwen3-4B at 20%, and accuracy gains two points; on top of LEMS allocation, perplexity ties while accuracy still improves; and there is a case where standard truncation beats the dense model on perplexity by damaging the sink and loses accuracy doing it. I have stopped trusting perplexity for this class of method, and I am curious whether others see the same.
GitHub Code: GitHub - aemonalgiz-dev/sink_deflation: A zero cost reweighting for Whitened-SVD LLMs · GitHub