cd /news/large-language-models/strategically-diverse-sampling-for-s… · home › topics › large-language-models › article
[ARTICLE · art-140800] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Strategically Diverse Sampling for Self-Training

A new arXiv paper (2609.31571v1) reports that self-training on strategically diverse data outperforms IID sampling, with self-training on strategically diverse but incorrect traces from Qwen3-4B beating IID distillation from a 235B teacher. The authors introduce GROOT, which builds a hierarchical tree of approaches and samples distinct paths, and adapt Verbalized Sampling to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction, models trained on strategically sampled data outperformed IID-trained counterparts on difficult tasks and provided strong initializations for RL and test-time scaling.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.31571v1 Announce Type: new Abstract: Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.

── more in #large-language-models 4 stories · sorted by recency
── more on @groot 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/strategically-divers…] indexed:0 read:1min 2026-09-28 · —