{"slug": "quantifying-the-memorization-to-generalization-transition-scaling-laws-and-phase", "title": "Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking", "summary": "A study posted to arXiv (2609.10657v1) mapped the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for grokking onset time of T_grok ∝ H^-0.27 D^-2.04 η^-0.50 λ^-0.64 (R² = 0.732; 0.821 with interactions). The exponent hierarchy shows data complexity (D^-2.04) dominates the regime transition over model capacity (H^-0.27): doubling data accelerates generalization by roughly 4x, while doubling width yields only about 1.2x. A sharp phase boundary at weight decay λ ≳ 1.0 separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions.", "body_md": "arXiv:2609.10657v1 Announce Type: new \nAbstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \\emph{why} this transition occurs, the quantitative structure of \\emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: $T_{\\mathrm{grok}} \\propto H^{-0.27}\\, D^{-2.04}\\, \\eta^{-0.50}\\, \\lambda^{-0.64}$ ($R^2 = 0.732$; $0.821$ with interactions). The exponent hierarchy reveals that data complexity ($D^{-2.04}$) is the dominant driver of regime transition, not model capacity ($H^{-0.27}$): doubling data accelerates generalization by ${\\sim}4\\times$, while doubling width yields only ${\\sim}1.2\\times$. A sharp phase boundary at weight decay $\\lambda \\gtrsim 1.0$ separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions. These results provide a quantitative foundation for predicting and controlling regime transitions in overparameterized networks.", "url": "https://wpnews.pro/news/quantifying-the-memorization-to-generalization-transition-scaling-laws-and-phase", "canonical_source": "https://arxiv.org/abs/2609.10657", "published_at": "2026-09-12 04:00:00+00:00", "updated_at": "2026-09-12 04:26:46.026303+00:00", "lang": "en", "topics": ["machine-learning", "neural-networks", "ai-research"], "entities": ["arXiv", "grokking", "MLPs"], "alternates": {"html": "https://wpnews.pro/news/quantifying-the-memorization-to-generalization-transition-scaling-laws-and-phase", "markdown": "https://wpnews.pro/news/quantifying-the-memorization-to-generalization-transition-scaling-laws-and-phase.md", "text": "https://wpnews.pro/news/quantifying-the-memorization-to-generalization-transition-scaling-laws-and-phase.txt", "jsonld": "https://wpnews.pro/news/quantifying-the-memorization-to-generalization-transition-scaling-laws-and-phase.jsonld"}}