{"slug": "moba-mixture-of-block-attention-for-long-context-llms", "title": "MOBA: Mixture of Block Attention for Long-Context LLMs", "summary": "MoonshotAI researchers submitted a paper on 18 Feb 2025 introducing Mixture of Block Attention (MoBA), an attention mechanism that applies Mixture of Experts principles to let long-context large language models select where to attend without predefined structural biases. MoBA has been deployed to support Kimi's long-context requests and can switch between full and sparse attention, with code released on GitHub.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 18 Feb 2025]\n\n# Title:MoBA: Mixture of Block Attention for Long-Context LLMs\n\n[View PDF](/pdf/2502.13189)\n\n[HTML (experimental)](https://arxiv.org/html/2502.13189v1)\n\nAbstract:Scaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI). However, the quadratic increase in computational complexity inherent in traditional attention mechanisms presents a prohibitive overhead. Existing approaches either impose strongly biased structures, such as sink or window attention which are task-specific, or radically modify the attention mechanism into linear approximations, whose performance in complex reasoning tasks remains inadequately explored.\n\nIn this work, we propose a solution that adheres to the ``less structure'' principle, allowing the model to determine where to attend autonomously, rather than introducing predefined biases. We introduce Mixture of Block Attention (MoBA), an innovative approach that applies the principles of Mixture of Experts (MoE) to the attention mechanism. This novel architecture demonstrates superior performance on long-context tasks while offering a key advantage: the ability to seamlessly transition between full and sparse attention, enhancing efficiency without the risk of compromising performance. MoBA has already been deployed to support Kimi's long-context requests and demonstrates significant advancements in efficient attention computation for LLMs. Our code is available at[this https URL](https://github.com/MoonshotAI/MoBA).\n    \n\n### Current browse context:\n\ncs.LG\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/moba-mixture-of-block-attention-for-long-context-llms", "canonical_source": "https://arxiv.org/abs/2502.13189", "published_at": "2026-09-13 19:12:21+00:00", "updated_at": "2026-09-13 19:51:27.124434+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "artificial-intelligence", "ai-research", "natural-language-processing"], "entities": ["MoonshotAI", "Mixture of Block Attention", "MoBA", "Kimi", "Mixture of Experts", "GitHub", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/moba-mixture-of-block-attention-for-long-context-llms", "markdown": "https://wpnews.pro/news/moba-mixture-of-block-attention-for-long-context-llms.md", "text": "https://wpnews.pro/news/moba-mixture-of-block-attention-for-long-context-llms.txt", "jsonld": "https://wpnews.pro/news/moba-mixture-of-block-attention-for-long-context-llms.jsonld"}}