Rethinking Privileged Information in On-Policy Self-Distillation
A new arXiv study (2608.18271v1) finds that performance gains from on-policy self-distillation (OPSD) in Qwen3 models (1.7B to 8B) do not consistently show that the student learned privileged reference information, as im…