arXiv:2608.02867v1 Announce Type: new Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reasoning tasks, there is significant debate as to whether RLVR expands the reasoning capability boundary, or just improves sampling efficiency. In this paper, we investigate the nature of test-time exploration in RLVR-trained LLMs by employing controlled maze-solving experiments and extracting a tree structure from mathematical reasoning traces (BODHI-Trees) based on semantic equivalence. This helps us delineate between entropy arising from stylistic variations and genuine inferential branching. Our findings demonstrate that the policy entropy collapse observed in RLVR models is not merely syntactic, and is accompanied by a significant reduction in semantic branching entropy. While RLVR improves adherence to environmental constraints and backtracking capabilities, it constricts the space of continuations; we provide evidence suggesting that this might be responsible for the sample efficiency gains of RLVR, albeit at the cost of genuine rollout diversity.
BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
A new arXiv preprint (2608.02867v1) from researchers studying reinforcement learning with verifiable rewards (RLVR) finds that RLVR-trained large language models (LLMs) exhibit reduced semantic branching entropy, not just stylistic variation, indicating a constriction of genuine inferential diversity. The study, which uses controlled maze-solving experiments and BODHI-Trees extracted from mathematical reasoning traces, suggests that RLVR's sample efficiency gains come at the cost of rollout diversity, despite improving adherence to constraints and backtracking.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.