04:00
2026-07-22
arxiv.org
large-language-models
One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context
A new method called Soft-Prompt Privileged Context (SPPC) for on-policy self-distillation in large language models uses a teacher that differs from the student only by a learnable soft prompt, avoidinβ¦