{"slug": "difficulty-adaptive-tree-structured-policy-optimization-for-expanding-reasoning", "title": "Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR", "summary": "A new method called Difficulty-Adaptive Tree-Structured Policy Optimization (DAPO) targets the failure of Reinforcement Learning with Verifiable Rewards (RLVR) to expand a large reasoning model's intrinsic reasoning coverage, measured as pass@k, which the work attributes to limited exploration during training. The approach is proposed to improve reasoning coverage beyond single-sample accuracy gains achieved by RLVR.", "body_md": "Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during trai", "url": "https://wpnews.pro/news/difficulty-adaptive-tree-structured-policy-optimization-for-expanding-reasoning", "canonical_source": "https://aiflash.com/news/116759/", "published_at": "2026-09-10 02:30:39+00:00", "updated_at": "2026-09-10 02:48:50.884978+00:00", "lang": "en", "topics": ["ai-research", "machine-learning", "large-language-models", "artificial-intelligence"], "entities": ["Reinforcement Learning with Verifiable Rewards", "Difficulty-Adaptive Tree-Structured Policy Optimization", "Large Reasoning Models"], "alternates": {"html": "https://wpnews.pro/news/difficulty-adaptive-tree-structured-policy-optimization-for-expanding-reasoning", "markdown": "https://wpnews.pro/news/difficulty-adaptive-tree-structured-policy-optimization-for-expanding-reasoning.md", "text": "https://wpnews.pro/news/difficulty-adaptive-tree-structured-policy-optimization-for-expanding-reasoning.txt", "jsonld": "https://wpnews.pro/news/difficulty-adaptive-tree-structured-policy-optimization-for-expanding-reasoning.jsonld"}}