{"slug": "calibrating-teacher-student-discrepancy-for-on-policy-distillation", "title": "Calibrating Teacher--Student Discrepancy for On-Policy Distillation", "summary": "A new research paper addresses on-policy distillation (OPD), a method for improving reasoning models by learning the token-level discrepancy between a stronger teacher model and an on-policy student model. The work argues that this discrepancy does not purely reflect the capability gap between teacher and student, since it also contains other deviations, and proposes calibrating it accordingly.", "body_md": "On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the t", "url": "https://wpnews.pro/news/calibrating-teacher-student-discrepancy-for-on-policy-distillation", "canonical_source": "https://aiflash.com/news/123521/", "published_at": "2026-09-21 04:30:16+00:00", "updated_at": "2026-09-21 04:54:03.832369+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "large-language-models", "natural-language-processing"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/calibrating-teacher-student-discrepancy-for-on-policy-distillation", "markdown": "https://wpnews.pro/news/calibrating-teacher-student-discrepancy-for-on-policy-distillation.md", "text": "https://wpnews.pro/news/calibrating-teacher-student-discrepancy-for-on-policy-distillation.txt", "jsonld": "https://wpnews.pro/news/calibrating-teacher-student-discrepancy-for-on-policy-distillation.jsonld"}}