cd /news/artificial-intelligence/on-the-off-policy-teacher-in-on-poli… · home › topics › artificial-intelligence › article
[ARTICLE · art-143643] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

On the Off-Policy Teacher in On-Policy Distillation

A new arXiv paper (2609.38360v1) proposes Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher model to student-generated prefixes in on-policy distillation (OPD). The authors report that a teacher's continuation performance degrades as student-generated prefixes grow longer, since those prefixes are off-policy for the teacher, and that SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards. Across multiple teacher-student configurations, model scales, and reasoning domains, SCOUT consistently improved the effectiveness of on-policy distillation, the authors report.

by read1 min views2 publishedOct 2, 2026

arXiv:2609.38360v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @scout 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/on-the-off-policy-te…] indexed:0 read:1min 2026-10-02 · —