{"slug": "one-symptom-three-levers-a-critical-review-of-on-policy-self-distillation", "title": "One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation", "summary": "A critical review of on-policy self-distillation for language models finds that the method, which trains a model on its own generations with token-level teacher scores, requires a second, larger teacher model, limiting its practicality. The review highlights the trade-offs between dense supervision and on-policy sampling inherent in the approach.", "body_md": "On-policy distillation trains a language model on its own generations while a teacher scores them token by token. It combines the dense supervision of imitation learning with the on-policy sampling of reinforcement learning. But it requires a second, larger model to act as teacher. On-Policy Self-Di", "url": "https://wpnews.pro/news/one-symptom-three-levers-a-critical-review-of-on-policy-self-distillation", "canonical_source": "https://aiflash.com/news/115419/", "published_at": "2026-09-08 07:30:10+00:00", "updated_at": "2026-09-08 08:01:53.444455+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/one-symptom-three-levers-a-critical-review-of-on-policy-self-distillation", "markdown": "https://wpnews.pro/news/one-symptom-three-levers-a-critical-review-of-on-policy-self-distillation.md", "text": "https://wpnews.pro/news/one-symptom-three-levers-a-critical-review-of-on-policy-self-distillation.txt", "jsonld": "https://wpnews.pro/news/one-symptom-three-levers-a-critical-review-of-on-policy-self-distillation.jsonld"}}