RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning A new method called RetireOPD proposes self-retiring on-policy distillation to give multi-turn reinforcement learning agents dense token-level supervision from a self-teacher with privileged task skills, so a skill-free student can internalize them. The approach targets the sparse single scalar reward per trajectory that multi-turn RL agents receive. The source provides no further results, figures, or named organizations. Multi-turn agents trained with reinforcement learning RL receive a single scalar reward per trajectory, which motivates self on-policy distillation OPD to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This rec