On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics A systematic study of distillation dynamics finds that existing comparisons between supervised fine-tuning and reinforcement learning vary too many factors at once, making it impossible to isolate the contribution of rollout policy to on-policy learning's claimed benefits of reduced catastrophic forgetting, sparser parameter updates, and improved generalisation. The research examines on-policy versus off-policy learning to disentangle those factors. On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy dif