arXiv:2609.25048v1 Announce Type: new Abstract: How many prompts does on-policy distillation (OPD) need, and how does the answer depend on the student policies that generate its training responses? We study these two controls jointly: prompt breadth and rollout refresh. A 3x3 mathematical-reasoning experiment fixes 14,080 trajectories and 110 optimizer updates while varying the prompt bank and the number of response-generating policy snapshots. With ten snapshots, eight prompts reach 24.09% average accuracy, close to 24.51% for 14,080 distinct prompts. With responses frozen at the initial policy, however, increasing breadth lowers accuracy from 21.16% to 19.05%; under per-update refresh, it raises accuracy from 23.61% to 25.57%. The resulting interaction is 4.07 percentage points, with a 95% question-paired interval of [2.00, 6.28]. Matched comparisons under two teachers reveal a second reversal: the periodic models have higher short-budget accuracy and answer completion, but frozen-response models overtake in average accuracy at a 32K output limit, using 1.7-1.8x as many response tokens. These results show that prompt efficiency in OPD can depend on both refresh and inference budget.
Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation
A 3x3 mathematical-reasoning experiment on on-policy distillation (OPD) found that prompt breadth and rollout refresh interact, producing a 4.07 percentage-point effect with a 95% question-paired interval of [2.00, 6.28]. Holding 14,080 trajectories and 110 optimizer updates fixed, eight prompts with ten policy snapshots reached 24.09% average accuracy versus 24.51% for 14,080 distinct prompts, while frozen responses lowered accuracy from 21.16% to 19.05% as breadth increased and per-update refresh raised it from 23.61% to 25.57%. Matched comparisons under two teachers showed periodic models led on short-budget accuracy and answer completion, but frozen-response models overtook in average accuracy at a 32K output limit using 1.7-1.8x as many response tokens, indicating prompt efficiency in OPD depends on both refresh and inference budget.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.