cd /news/artificial-intelligence/thermodynamic-weight-decay-exploring… · home topics artificial-intelligence article
[ARTICLE · art-71452] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Thermodynamic Weight Decay: Exploring Grokking Acceleration via Attention Specific Heat

Researchers introduced CvAdamW, a variant of the AdamW optimizer that monitors attention logit variance as specific heat to detect and accelerate grokking in neural networks, enabling generalization at epoch 2802 in a 4000-epoch budget on modular arithmetic where the baseline never groks. A cold-start variant reduced mean grokking latency by 257 epochs (6.0%; median 166 epochs; Wilcoxon p=0.049, Cohen's d=0.68) across 10 seeds, improving 8 of 10 seeds. The study, published on arXiv, suggests that physically motivated interventions can facilitate generalization within fixed compute budgets.

read1 min views1 publishedJul 24, 2026

arXiv:2607.20552v1 Announce Type: new Abstract: Grokking -- the delayed generalization of neural networks long after they have memorized their training data -- wastes thousands of training epochs and is notoriously unpredictable. Building on the recent result that Transformer attention is formally isomorphic to a thermodynamic system, we treat the variance of attention logits as a specific heat Cv and show that its peak reliably precedes the generalization transition. We introduce CvAdamW, a drop-in AdamW variant that monitors Cv online and injects thermal energy by dynamically scaling weight decay when a phase transition is detected. Through a strictly iterative development process we identify three failure modes -- initialization noise, mini-batch micro-ripples, and slingshot blinding -- and resolve them with a memorization gate and an exponential-moving-average shock absorber. On modular arithmetic (a+b mod 97), CvAdamW enables grokking at epoch 2802 in a 4000-epoch budget where the baseline never groks. We further propose a scale-invariant z-score reformulation that removes task-specific hyperparameters, and evaluate it across 10 paired seeds. A paired analysis shows the cold-start variant reduces mean grokking latency by 257 epochs (6.0%; median 166 epochs; Wilcoxon p=0.049, Cohen's d=0.68, bootstrap 95% CI [53,489]), improving 8 of 10 seeds; on this single task Cv peaks before grokking in all 10 seeds. Our results indicate that neural networks may expose detectable precursors of impending generalization transitions, and that a physically motivated, proportional intervention can facilitate generalization within a fixed compute budget. Code and data are public.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cvadamw 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/thermodynamic-weight…] indexed:0 read:1min 2026-07-24 ·