cd /news/machine-learning/language-specialized-multi-teacher-o… · home topics machine-learning article
[ARTICLE · art-87269] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

Researchers propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), a method that uses language-specialized teachers optimized via reinforcement learning and integrates their expertise into a generalist multilingual student through language routing and token-level distillation. In experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English, LS-MOPD outperformed RL baselines and surpassed the performance of the best-performing RL teachers, suggesting it can generalize beyond all teachers in multilingual ASR.

read1 min views1 publishedAug 5, 2026

arXiv:2608.03610v1 Announce Type: new Abstract: Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross-lingual knowledge to achieve competitive performance across multilingual benchmarks. However, joint modeling of languages with heterogeneous acoustic, phonological, and lexical characteristics inevitably introduces optimization conflicts, undermining language-wise specialization. To address this challenge, we propose Language-Specialized Multi-Teacher On-Policy Distillation (LS-MOPD), which decouples language-specific knowledge acquisition from multilingual capability integration: language-specialized teachers are independently optimized via reinforcement learning (RL), after which their expertise is integrated into a generalist multilingual student through language routing and token-level multi-teacher distillation, thereby reducing direct cross-lingual optimization conflicts. We further explore two acoustic-prefix configurations, static and dynamic, to examine how teacher--student prefix consistency influences the efficacy of on-policy distillation. Experiments on benchmarks covering Mandarin, Mandarin subdialects, Cantonese, and English demonstrate that LS-MOPD substantially outperforms RL baselines and consistently surpasses the empirical performance envelope defined by best-performing RL teachers, revealing its potential to generalize beyond all teachers in multilingual ASR.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/language-specialized…] indexed:0 read:1min 2026-08-05 ·