{"slug": "stochastic-teacher-intervention-for-agentic-on-policy-distillation", "title": "Stochastic Teacher Intervention for Agentic On-Policy Distillation", "summary": "Researchers introduced STI-OPD, a stochastic teacher intervention framework for multi-turn agentic on-policy distillation (OPD) that replaces a student model's proposed action with a teacher-generated one based on teacher-student policy discrepancy, according to the arXiv paper 2610.10878v1. STI-OPD estimates policy discrepancy using KL divergence and maps it to an intervention probability, sampling whether to intervene to balance teacher control with student exploration, and adds an Importance-Weighted Reverse KL objective to correct token sampling mismatch. The framework outperformed the strongest prior OPD baseline on every evaluated benchmark and student size across tool-integrated reasoning and long-horizon interaction, with ablations showing both discrepancy-guided intervention and importance weighting contribute to the gains.", "body_md": "arXiv:2610.10878v1 Announce Type: new \nAbstract: On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.", "url": "https://wpnews.pro/news/stochastic-teacher-intervention-for-agentic-on-policy-distillation", "canonical_source": "https://arxiv.org/abs/2610.10878", "published_at": "2026-10-09 04:00:00+00:00", "updated_at": "2026-10-09 04:17:56.163341+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-agents", "ai-research"], "entities": ["STI-OPD", "arXiv", "KL divergence", "Importance-Weighted Reverse KL"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stochastic-teacher-intervention-for-agentic-on-policy-distillation", "markdown": "https://wpnews.pro/news/stochastic-teacher-intervention-for-agentic-on-policy-distillation.md", "text": "https://wpnews.pro/news/stochastic-teacher-intervention-for-agentic-on-policy-distillation.txt", "jsonld": "https://wpnews.pro/news/stochastic-teacher-intervention-for-agentic-on-policy-distillation.jsonld"}}