Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents Researchers identified a failure mode they call Decision-Timestamp Mismatch in on-policy self-distillation (OPSD), a technique used to add dense privileged feedback to reinforcement learning with verifiable rewards (RLVR) for long-horizon agents. The finding addresses how RLVR's sparse outcome rewards provide only coarse supervision, while OPSD's dense feedback can misalign decisions with their timestamps. Reinforcement learning with verifiable rewards RLVR often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation OPSD complements this signal with dense privileged feedback. However, we identify Decision--Timestamp Mismatch: privileged