{"slug": "deriving-neural-scaling-laws-from-the-statistics-of-natural-language", "title": "Deriving neural scaling laws from the statistics of natural language", "summary": "A paper posted to arXiv on 7 Feb 2026 by Francesco Cagnetta claims the first theory that quantitatively predicts data-limited neural scaling exponents for modern LLMs trained on natural language. The theory isolates two statistical properties of language — the decay of pairwise token correlations with time separation and the decay of next-token conditional entropy with conditioning-context length — and derives a parameter-free formula that matched experimentally measured scaling laws from GPT-2 and LLaMA style models trained from scratch on TinyStories and WikiText.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 7 Feb 2026 (\n\n[v1](https://arxiv.org/abs/2602.07488v1)), last revised 3 Jul 2026 (this version, v3)]\n# Title:Deriving Neural Scaling Laws from the statistics of natural language\n\n[View PDF](https://arxiv.org/pdf/2602.07488)\n\n[HTML (experimental)](https://arxiv.org/html/2602.07488v3)\n\nAbstract:Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.\n    \n\n## Submission history\n\nFrom: Francesco Cagnetta [\n[view email](https://arxiv.org/show-email/c6400098/2602.07488)]\n\n**Sat, 7 Feb 2026 10:40:28 UTC (2,072 KB)**\n\n[\\[v1\\]](https://arxiv.org/abs/2602.07488v1)\n**Thu, 12 Feb 2026 11:54:22 UTC (2,072 KB)**\n\n[\\[v2\\]](https://arxiv.org/abs/2602.07488v2)\n**[v3]** Fri, 3 Jul 2026 02:02:48 UTC (3,083 KB)\n\n### Current browse context:\n\ncs.LG\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/deriving-neural-scaling-laws-from-the-statistics-of-natural-language", "canonical_source": "https://arxiv.org/abs/2602.07488", "published_at": "2026-09-17 09:08:19+00:00", "updated_at": "2026-09-17 09:25:43.575871+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["arXiv", "Francesco Cagnetta", "GPT-2", "LLaMA", "TinyStories", "WikiText"], "alternates": {"html": "https://wpnews.pro/news/deriving-neural-scaling-laws-from-the-statistics-of-natural-language", "markdown": "https://wpnews.pro/news/deriving-neural-scaling-laws-from-the-statistics-of-natural-language.md", "text": "https://wpnews.pro/news/deriving-neural-scaling-laws-from-the-statistics-of-natural-language.txt", "jsonld": "https://wpnews.pro/news/deriving-neural-scaling-laws-from-the-statistics-of-natural-language.jsonld"}}