cd /news/large-language-models/one-student-many-teachers-multi-task… · home topics large-language-models article
[ARTICLE · art-68025] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context

A new method called Soft-Prompt Privileged Context (SPPC) for on-policy self-distillation in large language models uses a teacher that differs from the student only by a learnable soft prompt, avoiding post-hoc rationalization and weight drift. On Qwen3-1.7B-Base and Phi-4-mini-instruct across four tasks, the single-task variant matches or exceeds full fine-tuning while training far fewer parameters, and the multi-task variant achieves the best overall average (56.2 on Qwen3-1.7B-Base) while preserving general-capability benchmarks.

read1 min publishedJul 22, 2026

arXiv:2607.18293v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts. Existing teachers either inject privileged context at the input -- inducing post-hoc rationalization -- or fine-tune weights, accumulating drift and forgetting across tasks. We propose \method, whose teacher differs from the student only by a learnable soft prompt: trained on $(x, y_\text{gold})$ pairs with the backbone frozen, the prompt yields a task-specific teacher that preserves the student's exact representational geometry. \method\ extends naturally to multi-task settings by routing each example in a merged corpus to its corresponding soft-prompt teacher, allowing a single student to absorb knowledge from $K$ teachers in parallel; at inference, all prompts are discarded. On Qwen3-1.7B-Base and Phi-4-mini-instruct across four tasks (Science, Tool Use, Biology, Math), the single-task variant (OPD with a PT teacher) matches or exceeds full fine-tuning while training orders of magnitude fewer parameters, and the multi-task variant achieves the best overall average ($56.2$ on Qwen3-1.7B-Base) while preserving general-capability benchmarks -- in contrast to sequential SFT, which degrades both.

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen3-1.7b-base 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/one-student-many-tea…] indexed:0 read:1min 2026-07-22 ·