{"slug": "strategically-diverse-sampling-for-self-training", "title": "Strategically Diverse Sampling for Self-Training", "summary": "A new arXiv paper (2609.31571v1) reports that self-training on strategically diverse data outperforms IID sampling, with self-training on strategically diverse but incorrect traces from Qwen3-4B beating IID distillation from a 235B teacher. The authors introduce GROOT, which builds a hierarchical tree of approaches and samples distinct paths, and adapt Verbalized Sampling to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction, models trained on strategically sampled data outperformed IID-trained counterparts on difficult tasks and provided strong initializations for RL and test-time scaling.", "body_md": "arXiv:2609.31571v1 Announce Type: new \nAbstract: Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.", "url": "https://wpnews.pro/news/strategically-diverse-sampling-for-self-training", "canonical_source": "https://www.machinebrief.com/news/strategically-diverse-sampling-for-self-training-mm4r", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 05:48:18.451027+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["GROOT", "Verbalized Sampling", "Qwen3-4B", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/strategically-diverse-sampling-for-self-training", "markdown": "https://wpnews.pro/news/strategically-diverse-sampling-for-self-training.md", "text": "https://wpnews.pro/news/strategically-diverse-sampling-for-self-training.txt", "jsonld": "https://wpnews.pro/news/strategically-diverse-sampling-for-self-training.jsonld"}}