Scaling Laws for Neural Language Models (first "scaling laws" paper from 2020) Researchers Samuel McCandlish and colleagues submitted "Scaling Laws for Neural Language Models" to arXiv on 23 January 2020, showing that language model cross-entropy loss falls as a power-law with model size, dataset size, and training compute across trends spanning more than seven orders of magnitude. The paper reports that network width and depth have minimal effect within a wide range, and that larger models are significantly more sample-efficient, so optimally compute-efficient training uses very large models on a relatively modest amount of data and stops significantly before convergence. Computer Science Machine Learning Submitted on 23 Jan 2020 Title:Scaling Laws for Neural Language Models View PDF https://arxiv.org/pdf/2001.08361 HTML experimental https://arxiv.org/html/2001.08361v1 Abstract:We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude. Other architectural details such as network width or depth have minimal effects within a wide range. Simple equations govern the dependence of overfitting on model/dataset size and the dependence of training speed on model size. These relationships allow us to determine the optimal allocation of a fixed compute budget. Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence. Submission history From: Samuel McCandlish view email https://arxiv.org/show-email/c9f9e554/2001.08361 v1 Thu, 23 Jan 2020 03:59:20 UTC 1,520 KB Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender IArxiv Recommender What is IArxiv? https://iarxiv.org/about arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .