Xiaomi livestreams MiMo-V2.6 reinforcement-learning runs Xiaomi's MiMo team is publicly livestreaming the reinforcement-learning training of two unreleased models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, via a dashboard at mimo.xiaomi.com/rl that shows reward curves, rollout counts, step timing, running compute costs, and mid-training coding evaluations. MiMo head Luo Fuli said the team spent nearly six months studying how far reinforcement learning could scale after the release of MiMo-V2.5, and the dashboard states the MiMo-V2.6 series is "coming soon." Xiaomi has not released either model for public use or announced final specifications, pricing, or an API. Xiaomi’s MiMo team is livestreaming the reinforcement-learning training of two unreleased models, MiMo-V2.6-Pro and MiMo-V2.6-Flash, through a public dashboard. The page displays training metrics directly from the trainer logs, including reward curves, rollout counts, step timing and running compute costs. The dashboard also tracks the models’ mid-training coding performance and other evaluation results as the runs continue. Xiaomi has not yet released either model for public use or announced final specifications, pricing or an API. MiMo head Luo Fuli said the team had spent nearly six months studying how far reinforcement learning could scale after the release of MiMo-V2.5. The dashboard says the MiMo-V2.6 series is “coming soon.” Xiaomi https://mimo.xiaomi.com/rl/