13:32
2026-08-16
arxiv.org
artificial-intelligence
RL for LLM Reasoning Is Sparse Policy Selection, Not Capability Learning
A new arXiv preprint (2605.06241v2) finds that reinforcement learning (RL) for large language model reasoning acts as sparse policy selection rather than capability learning, with only 1β3% of token pβ¦