arXiv:2610.10878v1 Announce Type: new Abstract: On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
Stochastic Teacher Intervention for Agentic On-Policy Distillation
Researchers introduced STI-OPD, a stochastic teacher intervention framework for multi-turn agentic on-policy distillation (OPD) that replaces a student model's proposed action with a teacher-generated one based on teacher-student policy discrepancy, according to the arXiv paper 2610.10878v1. STI-OPD estimates policy discrepancy using KL divergence and maps it to an intervention probability, sampling whether to intervene to balance teacher control with student exploration, and adds an Importance-Weighted Reverse KL objective to correct token sampling mismatch. The framework outperformed the strongest prior OPD baseline on every evaluated benchmark and student size across tool-integrated reasoning and long-horizon interaction, with ablations showing both discrepancy-guided intervention and importance weighting contribute to the gains.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.