{"slug": "scaling-domain-data-repetition-in-llm-pretraining", "title": "Scaling Domain Data Repetition in LLM Pretraining", "summary": "A new arXiv preprint (2608.14071) finds that for large language models, the optimal number of times to repeat high-quality domain data during pretraining increases mildly with model size at a fixed tokens-per-parameter ratio, and is strongly negatively correlated with the domain's final validation loss. The authors, who studied practical LLM scaling where training-token budgets grow proportionally with model size, suggest that repetition counts tuned on smaller proxy models with the same TPP can provide a practical estimate for larger models.", "body_md": "# Computer Science > Artificial Intelligence\n\n[Submitted on 14 Aug 2026]\n\n# Title:Scaling Domain Data Repetition in LLM Pretraining\n\n[View PDF](/pdf/2608.14071)\n\nAbstract:As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\\(\\mathrm{TPP}\\)). However, high-quality domain data is much harder to scale than general web data. As model size and the training-token budget increase, its fraction in the training mixture tends to decrease. Repeating the available high-quality data provides an effective way to counteract this dilution, but excessive repetition may lead to overfitting. We study this trade-off under practical LLM scaling, where the training-token budget grows proportionally with model size. For a fixed domain, we first find that, surprisingly at a fixed \\(\\mathrm{TPP}\\), the optimal repetition count mildly increases with model size. Across different domains, we find that the optimal repetition count is strongly negatively correlated with the final validation loss of a domain: domains with lower loss can generally benefit from more repetitions. In contrast, the amount of unique domain data is only weakly related to the optimal repetition count. These findings suggest that repetition counts tuned on smaller proxy models with the same \\(\\mathrm{TPP}\\) can provide a practical estimate for larger models.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/scaling-domain-data-repetition-in-llm-pretraining", "canonical_source": "https://arxiv.org/abs/2608.14071", "published_at": "2026-08-30 14:38:32+00:00", "updated_at": "2026-08-30 14:52:16.112393+00:00", "lang": "en", "topics": ["large-language-models", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/scaling-domain-data-repetition-in-llm-pretraining", "markdown": "https://wpnews.pro/news/scaling-domain-data-repetition-in-llm-pretraining.md", "text": "https://wpnews.pro/news/scaling-domain-data-repetition-in-llm-pretraining.txt", "jsonld": "https://wpnews.pro/news/scaling-domain-data-repetition-in-llm-pretraining.jsonld"}}