cd /news/large-language-models/grrr-the-geometry-of-reshaping-rotat… · home topics large-language-models article
[ARTICLE · art-136639] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training

A study of 12 post-training chains using supervised fine-tuning (SFT) and reinforcement learning (RL) found that removing the diagonal component of weight updates — the part that reshapes singular values — usually preserves most of the gains from post-training on a math evaluation suite. The research, posted as arXiv:2609.22146v1, expresses each weight update in the pretrained matrix's singular value decomposition (SVD) frame, separating changes into diagonal values, off-diagonal values that rotate the coupling between pretrained input and output directions, and null-space values that route outside the matrix's original nonzero SVD core. The authors conclude that post-training gains are carried primarily by reconfiguring and extending pretrained pathways rather than by substantially changing the singular values of pretrained models.

by read1 min views1 publishedSep 22, 2026

arXiv:2609.22146v1 Announce Type: new Abstract: We study how post-training changes the weights of Large Language Models (LLMs) relative to their pretrained weights. Across 12 post-training chains with supervised fine-tuning (SFT) and reinforcement learning (RL), we express each weight update in the pretrained matrix's singular value decomposition (SVD) frame. This decomposition separates the changes of three geometrically distinct components: diagonal values, which reshapes singular values; off-diagonal values, which rotates the coupling between pretrained input and output directions; and null-space values, which routes outside the matrix's original nonzero SVD core. On a math evaluation suite, we find that removing the diagonal component usually preserves most of the gains from post-training. These results suggest that post-training gains are carried primarily by reconfiguring and extending pretrained pathways rather than by substantially changing singular values of pre-trained models.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grrr-the-geometry-of…] indexed:0 read:1min 2026-09-22 ·