{"slug": "scaling-laws-for-looped-mixture-of-experts", "title": "Scaling Laws for Looped Mixture of Experts", "summary": "A paper submitted to arXiv on 30 Sep 2026 introduces Loop Scaling Laws, which its authors describe as the first scaling law to jointly model recurrence and sparsity alongside model size and data. The fitted laws predict held-out loss of looped models more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases; downstream evaluations show sparsity delivers roughly 3x active-parameter efficiency and recurrence roughly 2x total-parameter efficiency on reasoning. At trillion-token scale and matched training compute, a looped MoE using law-derived recurrence matches a roughly 2x larger non-looped MoE on reasoning benchmarks while enabling test-time scaling through recurrence.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 30 Sep 2026]\n\n# Title:Scaling Laws for Looped Mixture of Experts\n\n[View PDF](https://arxiv.org/pdf/2609.40316)\n\n[HTML (experimental)](https://arxiv.org/html/2609.40316v1)\n\nAbstract:Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.\n    \n\n### Current browse context:\n\ncs.LG\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/scaling-laws-for-looped-mixture-of-experts", "canonical_source": "https://arxiv.org/abs/2609.40316", "published_at": "2026-10-02 04:13:59+00:00", "updated_at": "2026-10-02 04:46:17.565343+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "artificial-intelligence"], "entities": ["arXiv", "Loop Scaling Laws", "Mixture-of-Experts", "Looped transformers"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/scaling-laws-for-looped-mixture-of-experts", "markdown": "https://wpnews.pro/news/scaling-laws-for-looped-mixture-of-experts.md", "text": "https://wpnews.pro/news/scaling-laws-for-looped-mixture-of-experts.txt", "jsonld": "https://wpnews.pro/news/scaling-laws-for-looped-mixture-of-experts.jsonld"}}