cd /news/artificial-intelligence/multi-teacher-on-policy-distillation… · home topics artificial-intelligence article
[ARTICLE · art-112932] src=mimo.xiaomi.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Multi-Teacher On-Policy Distillation for Capability Integration in Post-Training

Researchers propose Multi-teacher On-Policy Distillation (MOPD), a post-training paradigm that combines the capabilities of multiple domain-specific reinforcement learning teachers into a single large language model by distilling them on the student's own rollouts. On Qwen3-30B-A3B, MOPD outperforms Mix-RL, Cascade RL, Off-Policy Finetune, and Param-Merge baselines, inheriting nearly all of each teacher's capability, and has been deployed in the post-training of MiMo-V2-Flash, an industrial-scale frontier model.

read1 min views1 publishedAug 27, 2026
[Download PDF](/papers/mopd.pdf)

[Next →](/paper/arl-tangram)

Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard. Existing methods, such as Off-Policy Finetune and Mix-RL, are either inefficient or lose performance. In this work, we propose Multi-teacher On-Policy Distillation (MOPD), a post-training paradigm for combining the capabilities of multiple domain RL teachers: we first run per-domain specialised RL to obtain a set of domain teachers, then distill these teachers into the student on its own rollouts. This eliminates exposure bias and provides a dense optimization signal. On Qwen3-30B-A3B, MOPD outperforms Mix-RL, Cascade RL, Off-Policy Finetune, and Param-Merge baselines, inheriting nearly all of each teacher’s capability. MOPD also enables parallel, independent development of domain teachers, removing the cross-domain coupling typical of multi-domain post-training. MOPD has been deployed in the post-training of MiMo-V2-Flash, an industrial-scale frontier model, demonstrating its practical value for capability integration in frontier-scale LLMs.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @qwen3-30b-a3b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/multi-teacher-on-pol…] indexed:0 read:1min 2026-08-27 ·