04:00
2026-09-24
arxiv.org
artificial-intelligence
Reinforcement Learning with Decomposed Subtasks
A new arXiv paper (arXiv:2609.27035v1) introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that replaces the scalar GRPO advantage with Subtask-Decomposed Advantage Estimation …