# Calibrating Teacher--Student Discrepancy for On-Policy Distillation

> Source: <https://aiflash.com/news/123521/>
> Published: 2026-09-21 04:30:16+00:00

On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the t
