{"slug": "lexicographic-multi-objective-on-policy-distillation", "title": "Lexicographic Multi-Objective On-Policy Distillation", "summary": "Researchers introduced Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method that integrates reward-specialized policies under explicit priority orders, according to an arXiv paper (arXiv:2610.02359v1). Evaluated on 30B-A3B mixture-of-experts transformer models across three math benchmarks, LMOPD with two experts fully retained accuracy and reasoning-quality gains while acquiring 46.9% of the conciseness gain, and with four experts retained roughly 90% of both the accuracy and reasoning-correctness gains versus about 57% for the next best baseline. The authors report that lexicographic routing outperforms random routing and that projection further strengthens top-priority capabilities.", "body_md": "arXiv:2610.02359v1 Announce Type: cross \nAbstract: Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\\%$ of the conciseness gain. With four experts, it retains $\\approx90\\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\\approx57\\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.", "url": "https://wpnews.pro/news/lexicographic-multi-objective-on-policy-distillation", "canonical_source": "https://www.machinebrief.com/news/lexicographic-multi-objective-on-policy-distillation-2sqc", "published_at": "2026-10-05 04:00:00+00:00", "updated_at": "2026-10-05 05:13:07.567838+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["Lexicographic Multi-Objective On-Policy Distillation", "LMOPD", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/lexicographic-multi-objective-on-policy-distillation", "markdown": "https://wpnews.pro/news/lexicographic-multi-objective-on-policy-distillation.md", "text": "https://wpnews.pro/news/lexicographic-multi-objective-on-policy-distillation.txt", "jsonld": "https://wpnews.pro/news/lexicographic-multi-objective-on-policy-distillation.jsonld"}}