{"slug": "multi-teacher-on-policy-distillation-for-capability-integration-in-post-training", "title": "Multi-Teacher On-Policy Distillation for Capability Integration in Post-Training", "summary": "Researchers propose Multi-teacher On-Policy Distillation (MOPD), a post-training paradigm that combines the capabilities of multiple domain-specific reinforcement learning teachers into a single large language model by distilling them on the student's own rollouts. On Qwen3-30B-A3B, MOPD outperforms Mix-RL, Cascade RL, Off-Policy Finetune, and Param-Merge baselines, inheriting nearly all of each teacher's capability, and has been deployed in the post-training of MiMo-V2-Flash, an industrial-scale frontier model.", "body_md": "# MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training\n\n[Download PDF](/papers/mopd.pdf)\n\n[Next →](/paper/arl-tangram)\n\nModern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard. Existing methods, such as Off-Policy Finetune and Mix-RL, are either inefficient or lose performance. In this work, we propose Multi-teacher On-Policy Distillation (MOPD), a post-training paradigm for combining the capabilities of multiple domain RL teachers: we first run per-domain specialised RL to obtain a set of domain teachers, then distill these teachers into the student on its own rollouts. This eliminates exposure bias and provides a dense optimization signal. On Qwen3-30B-A3B, MOPD outperforms Mix-RL, Cascade RL, Off-Policy Finetune, and Param-Merge baselines, inheriting nearly all of each teacher’s capability. MOPD also enables parallel, independent development of domain teachers, removing the cross-domain coupling typical of multi-domain post-training. MOPD has been deployed in the post-training of MiMo-V2-Flash, an industrial-scale frontier model, demonstrating its practical value for capability integration in frontier-scale LLMs.", "url": "https://wpnews.pro/news/multi-teacher-on-policy-distillation-for-capability-integration-in-post-training", "canonical_source": "https://mimo.xiaomi.com/paper/mopd", "published_at": "2026-08-27 10:49:58+00:00", "updated_at": "2026-08-27 11:19:55.129885+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["Qwen3-30B-A3B", "MiMo-V2-Flash", "MOPD"], "alternates": {"html": "https://wpnews.pro/news/multi-teacher-on-policy-distillation-for-capability-integration-in-post-training", "markdown": "https://wpnews.pro/news/multi-teacher-on-policy-distillation-for-capability-integration-in-post-training.md", "text": "https://wpnews.pro/news/multi-teacher-on-policy-distillation-for-capability-integration-in-post-training.txt", "jsonld": "https://wpnews.pro/news/multi-teacher-on-policy-distillation-for-capability-integration-in-post-training.jsonld"}}