{"slug": "xiaomi-spent-3m-to-train-mimo-v2-6-pro-1t-a42b-and-it-is-now-out-as-an-open", "title": "Xiaomi spent $3M to train MiMo-V2.6-Pro 1T-A42B and it is now out as an open weights model.", "summary": "Xiaomi released MiMo-V2.6-Pro, a 1T-A42B open-weights omnimodal model family that the company says cost $3M to train, alongside Flash and Pro-UltraSpeed variants, with Pro-UltraSpeed claimed to deliver 20x faster output without quality loss. Xiaomi's technical report details an RL phase scaled across three axes: a fully asynchronous architecture handling 1,568 samples per update at 1M context length and 3.5 to 3.7B tokens per step, a multi-task suite mixing coding, visual, cyber and general agent tasks, and increased grader compute using relative comparison within each group. Xiaomi is open sourcing tooling and environments — including coding recipes, the ARVO cyber environment, general knowledge tools, web-development grading, music data preparation and agent configurations — while withholding the full 7k+ task datasets.", "body_md": "# Xiaomi spent $3M to train MiMo-V2.6-Pro 1T-A42B and it is now out as an open weights model.\n\nI have been digging into the technical report for this release because the transparency on their RL training runs is actually wild. They've released a whole family here: the Pro version is the heavy hitter, the Flash version is for efficiency, and there is a Pro-UltraSpeed version that supposedly hits 20x faster output without losing quality. The big thing is that these are natively omnimodal, which puts Xiaomi in a very interesting spot compared to the usual \"AI Tigers\" in China.\n\n## How did they scale the RL compute?\n\nThe most interesting part of the MiMo-V2.6-Pro development is how they handled the Reinforcement Learning phase. They didn't just throw GPUs at it; they scaled across three specific axes that I think other devs should look at if they are trying to optimize their own training loops.\n\nFirst, they pushed throughput with a fully asynchronous architecture. We are talking about 1,568 samples per update, training at a 1M context length, and hitting 3.5 to 3.7B tokens per step. That is a massive amount of data moving through the system.\n\nSecond, they used a multi-task training suite. Instead of just focusing on one area, they mixed coding, visual, cyber, and general agent tasks across several harnesses. The goal here was to make sure gains in one area actually reinforced the other capabilities.\n\nFinally, they increased their grader compute. By using relative comparison within each group, they managed to get more precise reward signals for long-horizon RL tasks. This effectively closed a self-improvement loop that steers the model toward shorter paths and uses fewer tokens to complete a task.\n\n## What actually gets open sourced?\n\nXiaomi is promising to open source the tooling and environments, though they are holding onto the full 7k+ task datasets for now. If you are looking for the specific recipes, here is what is coming:\n\n- **Coding/Software Engineering:** These include the code recipes, dataset loader, and rewards.\n- **Cyber/Vulnerability:** They are releasing the ARVO environment and the training recipe.\n- **General Knowledge:** This covers the general environment, tools, and training recipe.\n- **Visual/Web Dev:** You get the web-development environment and the grading system.\n- **Music:** Data preparation and the music scorer are included.\n- **Infrastructure:** They are releasing the agent configurations (composable mini-harnesses) and the mimoagent environments (shared environment adapters).\n\n[Next Gradual Disempowerment is a more realistic existential risk than most people realize →](https://promptcube3.com/en/threads/9572/)\n\n## All Replies （1）\n\nWant a live back-and-forth? [Join the global AI chat room](https://promptcube3.com/en/chat/) — login to talk.\n\nRelieved they finally dropped this. I've been waiting for the Pro-UltraSpeed version since my last project lagged out on slower inference.", "url": "https://wpnews.pro/news/xiaomi-spent-3m-to-train-mimo-v2-6-pro-1t-a42b-and-it-is-now-out-as-an-open", "canonical_source": "https://promptcube3.com/en/threads/9581/", "published_at": "2026-09-23 16:12:03+00:00", "updated_at": "2026-09-23 16:31:00.938547+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents", "ai-products"], "entities": ["Xiaomi", "MiMo-V2.6-Pro", "MiMo-V2.6-Flash", "MiMo-V2.6-Pro-UltraSpeed", "ARVO", "mimoagent"], "alternates": {"html": "https://wpnews.pro/news/xiaomi-spent-3m-to-train-mimo-v2-6-pro-1t-a42b-and-it-is-now-out-as-an-open", "markdown": "https://wpnews.pro/news/xiaomi-spent-3m-to-train-mimo-v2-6-pro-1t-a42b-and-it-is-now-out-as-an-open.md", "text": "https://wpnews.pro/news/xiaomi-spent-3m-to-train-mimo-v2-6-pro-1t-a42b-and-it-is-now-out-as-an-open.txt", "jsonld": "https://wpnews.pro/news/xiaomi-spent-3m-to-train-mimo-v2-6-pro-1t-a42b-and-it-is-now-out-as-an-open.jsonld"}}