Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards
Researchers propose Gradient-Aligned Reward (GAR), a dense reward method for reinforcement learning from verifiable rewards (RLVR) that uses cosine similarity between rollout and expert-anchor gradien…