{"slug": "smelt-scaling-laws-for-compute-matched-moe-looped-transformers", "title": "SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers", "summary": "A new research paper introduces SMELT, a study of scaling laws for compute-matched Mixture-of-Experts (MoE) looped Transformers, finding that looping provides architectural advantages beyond mere FLOPs. The authors compare looped and non-looped models at matched per-token FLOPs and total non-embedding parameters, demonstrating improved performance for looped MoE Transformers.", "body_md": "Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding p", "url": "https://wpnews.pro/news/smelt-scaling-laws-for-compute-matched-moe-looped-transformers", "canonical_source": "https://aiflash.com/news/112582/", "published_at": "2026-09-02 03:30:00+00:00", "updated_at": "2026-09-02 03:51:42.874680+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["SMELT"], "alternates": {"html": "https://wpnews.pro/news/smelt-scaling-laws-for-compute-matched-moe-looped-transformers", "markdown": "https://wpnews.pro/news/smelt-scaling-laws-for-compute-matched-moe-looped-transformers.md", "text": "https://wpnews.pro/news/smelt-scaling-laws-for-compute-matched-moe-looped-transformers.txt", "jsonld": "https://wpnews.pro/news/smelt-scaling-laws-for-compute-matched-moe-looped-transformers.jsonld"}}