Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR A new method called Difficulty-Adaptive Tree-Structured Policy Optimization (DAPO) targets the failure of Reinforcement Learning with Verifiable Rewards (RLVR) to expand a large reasoning model's intrinsic reasoning coverage, measured as pass@k, which the work attributes to limited exploration during training. The approach is proposed to improve reasoning coverage beyond single-sample accuracy gains achieved by RLVR. Reinforcement Learning with Verifiable Rewards RLVR has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage pass@k due to limited exploration during trai