cd /news/artificial-intelligence/it-takes-8-tokens-weak-to-strong-off… · home topics artificial-intelligence article
[ARTICLE · art-66378] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches

Researchers propose W2SPO, an off-policy reinforcement learning method that uses short auxiliary segments from a weaker model to improve reasoning in large language models, achieving a Pass@1 improvement from 62.3% to 64.2% and a 3.55 times training speedup over vanilla GRPO on mathematical reasoning benchmarks.

read1 min views2 publishedJul 21, 2026

arXiv:2607.16205v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts. However, we identify a critical support limited bottleneck in this paradigm: on challenging reasoning tasks, the target model's samples often exhibit semantic redundancy, converging into the same erroneous "reasoning basins" that offer negligible reward contrast for policy updates. In this paper, we propose to overcome this limitation through a weak to strong learning paradigm, where a policy's exploration is informed by a weaker but computationally efficient auxiliary model. We introduce W2SPO, an off policy RL method that injects short auxiliary segments often as brief as 8 tokens into intermediate target model trajectories and the target model then completes the reasoning path from these diverted states. Policy updates are restricted to these short inserted segments based on final verifiable rewards. Empirically, W2SPO achieves superior performance among evaluated 4B scale models on mathematical reasoning benchmarks, outperforming evaluated post trained baselines. Compared with vanilla GRPO under the same sampling budget, W2SPO improves Pass@1 from 62.3% to 64.2% while achieving a 3.55 times training speedup. These results suggest that weak auxiliary branches can induce stronger target reasoning policies by expanding local exploration support.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @w2spo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/it-takes-8-tokens-we…] indexed:0 read:1min 2026-07-21 ·