Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Researchers propose trace-based on-policy distillation (TOPD), a teacher-supervised framework that transfers reasoning ability to a diffusion large language model (dLLM) without reward estimation. TOP…