# On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

> Source: <https://aiflash.com/news/129930/>
> Published: 2026-10-02 16:30:01+00:00

On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy dif
