04:00
2026-07-21
arxiv.org
artificial-intelligence
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Researchers propose W2SPO, an off-policy reinforcement learning method that uses short auxiliary segments from a weaker model to improve reasoning in large language models, achieving a Pass@1 improvemβ¦