{"slug": "pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few", "title": "PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation", "summary": "Researchers proposed PIVOT (Perplexity-Informed Transition Optimization), a dynamic framework that routes samples between On-Policy Distillation and GRPO reinforcement learning based on teacher-evaluated sequence perplexity, according to arXiv paper 2610.11167v1. On the Banking77 and HWU64 benchmarks, PIVOT outperformed continued OPD and globally synchronized OPD-to-GRPO baselines under the same number of post-warm-up student optimization steps, yielding stronger downstream performance and more stable training dynamics. The method addresses vertical-domain few-shot classification for small language models by moving low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition.", "body_md": "arXiv:2610.11167v1 Announce Type: new \nAbstract: Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.", "url": "https://wpnews.pro/news/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few", "canonical_source": "https://www.machinebrief.com/news/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-x7s9", "published_at": "2026-10-09 04:00:00+00:00", "updated_at": "2026-10-09 05:46:59.563640+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["PIVOT", "Perplexity-Informed Transition Optimization", "On-Policy Distillation", "GRPO", "Banking77", "HWU64", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few", "markdown": "https://wpnews.pro/news/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few.md", "text": "https://wpnews.pro/news/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few.txt", "jsonld": "https://wpnews.pro/news/pivot-perplexity-informed-kd-to-rl-transition-scheduling-for-vertical-domain-few.jsonld"}}