{"slug": "xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public", "title": "Xiaomi MiMo-V2.6 Breaks Cover: A 1T-Class Chinese Lab Trains in Public", "summary": "Xiaomi's MiMo team began publicly streaming the reinforcement learning training run for its MiMo-V2.6 model, with the MiMo-V2.6-Pro run that started September 15 consuming roughly $432,000 per day, or $5 per second, according to a live dashboard at mimo.xiaomi.com/rl. Team head Luo Fuli, formerly of DeepSeek, ended nearly half a year of silence on September 16 to announce the run, which uses about 2 billion tokens per step across 1,568 prompts and 16 fully asynchronous rollouts. MiMo-V2.6-Pro has reached 65.97% on the DeepSWE v1.1 benchmark, a 47-point gain over the MiMo-V2.5 baseline's 19%, trailing GPT-6 Astra at 74% and Claude Fable 5 at 70% but ahead of Grok 4.6 at 67%.", "body_md": "Xiaomi has initiated a live, public stream of its [reinforcement learning](https://forkast.news/glossary/reinforcement-learning/) training run for the MiMo-V2.6 model, a level of operational exposure currently absent from major US-based frontier labs. By providing a [live dashboard](https://mimo.xiaomi.com/rl/) that tracks real-time costs, token throughput, and benchmark performance, the Xiaomi MiMo team is challenging the industry standard of closed-door development cycles.\n\nLuo Fuli, head of the Xiaomi MiMo team and formerly of DeepSeek, [signaled the end of a quiet period](https://x.com/_LuoFuli/status/2100296686719610932) on September 16. \n\nNearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now.\n\nThis public disclosure marks a departure from the opaque, milestone-based announcements typical of the current AI race.\n\n## Scaling the RL Infrastructure\n\nThe technical scope of the MiMo-V2.6 run focuses on three specific dimensions of scaling. First, the team has scaled compute to approximately 2 billion tokens per step, utilizing 1,568 prompts across 16 fully asynchronous rollouts. Second, they have integrated multi-task [agentic AI](https://forkast.news/what-is-agentic-ai/) environments, mixing various harnesses into a single training run. Finally, the team has scaled grader compute, implementing agentic in-group credit assignment that relies on test-case and rubric-based rewards.\n\nThe financial commitment required for this transparency is substantial. The MiMo-V2.6-Pro run, which began on September 15, is currently consuming resources at a rate of approximately $432,000 per day, or $5 per second. As of the latest data, the Pro model has processed 32.5 billion tokens, while the concurrent MiMo-V2.6-Flash run has processed 49.4 billion tokens at a cost of $512,000.\n\n## Performance and Benchmarking\n\nThe performance gains reported during this mid-training phase are significant. The MiMo-V2.6-Pro model has reached a score of 65.97% on the [DeepSWE v1.1 benchmark](https://deepswe.datacurve.ai/blog/deepswe-v1-1). This represents a 47-point jump from the MiMo-V2.5 baseline, which previously scored 19%. When placed against current frontier models on the same benchmark, MiMo-V2.6-Pro sits in a competitive tier: GPT-6 Astra leads at 74%, followed by Claude Fable 5 at 70%, Kimi K3 at 69%, and Grok 4.6 at 67%.\n\nThis performance data is presented alongside a suite of metrics including step timings and dataset composition across 23 distinct categories. Luo Fuli has stated that the team intends to open-source the technical details incrementally in the coming weeks, following the precedent set by the MIT-licensed MiMo-V2.5.\n\n## Transparency as a Counter-Narrative\n\nThis move by Xiaomi aligns with a broader 2026 trend among Chinese labs, including Kimi K3 and Qwen 3.8, to prioritize compute and training transparency. This approach stands in direct contrast to the [Amodei pacing framework](https://forkast.news/amodeis-pacing-framework-is-not-a-pause-its-an-operating-model-for-the-frontier/), which posits that safety must function as an operating model requiring strict coordination and, by extension, controlled information flow. While US labs often cite safety as a justification for secrecy, Xiaomi’s public dashboard suggests an alternative philosophy where accountability is derived from visibility.\n\nThe strategy also intersects with ongoing discussions regarding [safety accountability architecture](https://forkast.news/openai-formalizes-misalignment-disclosure-six-incident-reports-mark-a-shift-from-promises-to-infrastructure/) and the integration of [agent-native research tools](https://forkast.news/stanfords-paper2agent-turns-research-papers-into-deployable-ai-agents/). By exposing the training process, Xiaomi is forcing a debate on whether the industry’s current reliance on closed-door development is a necessary safety measure or a barrier to collective progress.\n\n## Limits of the Dashboard\n\nDespite the technical detail, the transparency play has met with skepticism. Community members on [Hacker News](https://news.ycombinator.com/item?id=49732270) have noted that the dashboard numbers may reset or replay upon page refresh, challenging the authenticity of the real-time data. Furthermore, eagle-eyed observers noted a line item labeled “Claude Distill Requests: hidden” on the dashboard. This suggests that despite the open nature of the training run, the model may still rely on distillation from external proprietary models, complicating the narrative of independent scaling.\n\nWhether this transparency is a genuine shift in research culture or a strategic performance remains an open question. For investors and industry professionals, the MiMo-V2.6 run serves as a test case for whether public, high-cost training runs can coexist with the competitive pressures of the current AI landscape.", "url": "https://wpnews.pro/news/xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public", "canonical_source": "https://forkast.news/xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public/", "published_at": "2026-09-17 20:53:25+00:00", "updated_at": "2026-09-17 21:23:28.193338+00:00", "lang": "en", "topics": ["ai-research", "machine-learning", "large-language-models", "ai-safety", "ai-policy"], "entities": ["Xiaomi", "MiMo-V2.6", "Luo Fuli", "DeepSeek", "MiMo-V2.6-Pro", "MiMo-V2.6-Flash", "DeepSWE v1.1", "MiMo-V2.5"], "alternates": {"html": "https://wpnews.pro/news/xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public", "markdown": "https://wpnews.pro/news/xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public.md", "text": "https://wpnews.pro/news/xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public.txt", "jsonld": "https://wpnews.pro/news/xiaomi-mimo-v2-6-breaks-cover-a-1t-class-chinese-lab-trains-in-public.jsonld"}}