cd /news/machine-learning/quantifying-the-memorization-to-gene… · home topics machine-learning article
[ARTICLE · art-127424] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

A study posted to arXiv (2609.10657v1) mapped the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for grokking onset time of T_grok ∝ H^-0.27 D^-2.04 η^-0.50 λ^-0.64 (R² = 0.732; 0.821 with interactions). The exponent hierarchy shows data complexity (D^-2.04) dominates the regime transition over model capacity (H^-0.27): doubling data accelerates generalization by roughly 4x, while doubling width yields only about 1.2x. A sharp phase boundary at weight decay λ ≳ 1.0 separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions.

by read1 min views3 publishedSep 12, 2026

arXiv:2609.10657v1 Announce Type: new Abstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: $T_{\mathrm{grok}} \propto H^{-0.27}, D^{-2.04}, \eta^{-0.50}, \lambda^{-0.64}$ ($R^2 = 0.732$; $0.821$ with interactions). The exponent hierarchy reveals that data complexity ($D^{-2.04}$) is the dominant driver of regime transition, not model capacity ($H^{-0.27}$): doubling data accelerates generalization by ${\sim}4\times$, while doubling width yields only ${\sim}1.2\times$. A sharp phase boundary at weight decay $\lambda \gtrsim 1.0$ separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions. These results provide a quantitative foundation for predicting and controlling regime transitions in overparameterized networks.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/quantifying-the-memo…] indexed:0 read:1min 2026-09-12 ·