04:00
2026-07-21
arxiv.org
artificial-intelligence
Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Researchers propose trace-based on-policy distillation (TOPD), a teacher-supervised framework that transfers reasoning ability to a diffusion large language model (dLLM) without reward estimation. TOPβ¦