cd /news/artificial-intelligence/scaling-laws-for-neural-language-mod… · home › topics › artificial-intelligence › article
[ARTICLE · art-139406] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Scaling Laws for Neural Language Models (first "scaling laws" paper from 2020)

Researchers Samuel McCandlish and colleagues submitted "Scaling Laws for Neural Language Models" to arXiv on 23 January 2020, showing that language model cross-entropy loss falls as a power-law with model size, dataset size, and training compute across trends spanning more than seven orders of magnitude. The paper reports that network width and depth have minimal effect within a wide range, and that larger models are significantly more sample-efficient, so optimally compute-efficient training uses very large models on a relatively modest amount of data and stops significantly before convergence.

read2 min views1 publishedSep 25, 2026
Scaling Laws for Neural Language Models (first "scaling laws" paper from 2020)
Image: source
  [Submitted on 23 Jan 2020]


[View PDF](https://arxiv.org/pdf/2001.08361)

[HTML (experimental)](https://arxiv.org/html/2001.08361v1)

Abstract:We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude. Other architectural details such as network width or depth have minimal effects within a wide range. Simple equations govern the dependence of overfitting on model/dataset size and the dependence of training speed on model size. These relationships allow us to determine the optimal allocation of a fixed compute budget. Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.

Submission history #

From: Samuel McCandlish [
[view email](https://arxiv.org/show-email/c9f9e554/2001.08361)]

**[v1]** Thu, 23 Jan 2020 03:59:20 UTC (1,520 KB)

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?) alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?) Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) IArxiv Recommender

(What is IArxiv?) arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scaling-laws-for-neu…] indexed:0 read:2min 2026-09-25 · —