{"slug": "gradients-know-what-outcomes-don-t-unlocking-reinforcement-learning-for-llm-with", "title": "Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards", "summary": "Researchers propose Gradient-Aligned Reward (GAR), a dense reward method for reinforcement learning from verifiable rewards (RLVR) that uses cosine similarity between rollout and expert-anchor gradients to improve LLM reasoning with less than 9% wall-clock overhead. On Qwen3-4B and Qwen3-8B, GAR outperforms GRPO and other baselines on competition-level math benchmarks and transfers to GPQA Diamond and MMLU-Pro without domain-specific data. The method is detailed in arXiv paper 2609.03342v1, with code and data available on GitHub.", "body_md": "arXiv:2609.03342v1 Announce Type: new\nAbstract: Reinforcement learning from verifiable rewards (RLVR) drives chain-of-thought reasoning in large language models, yet its binary outcome reward cannot distinguish among correct trajectories. Existing dense reward alternatives, from surface heuristics to process reward models, either ignore the expert solutions already present in training corpora or require expensive offline annotation. We propose Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock overhead. We prove that this cosine admits a multiplicative decomposition into prediction-error and activation-pattern factors, providing a concrete characterization of what the alignment signal measures. On Qwen3-4B and Qwen3-8B, GAR consistently improves over GRPO and other baselines on competition-level math benchmarks and transfers to GPQA Diamond and MMLU-Pro without domain-specific data. Code and data are available at https://github.com/LQgdwind/GAR.", "url": "https://wpnews.pro/news/gradients-know-what-outcomes-don-t-unlocking-reinforcement-learning-for-llm-with", "canonical_source": "https://www.machinebrief.com/news/gradients-know-what-outcomes-dont-unlocking-reinforcement-le-c2h4", "published_at": "2026-09-04 04:00:00+00:00", "updated_at": "2026-09-04 07:22:14.907148+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-tools"], "entities": ["arXiv", "Qwen3-4B", "Qwen3-8B", "GRPO", "Gradient-Aligned Reward (GAR)", "GPQA Diamond", "MMLU-Pro", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/gradients-know-what-outcomes-don-t-unlocking-reinforcement-learning-for-llm-with", "markdown": "https://wpnews.pro/news/gradients-know-what-outcomes-don-t-unlocking-reinforcement-learning-for-llm-with.md", "text": "https://wpnews.pro/news/gradients-know-what-outcomes-don-t-unlocking-reinforcement-learning-for-llm-with.txt", "jsonld": "https://wpnews.pro/news/gradients-know-what-outcomes-don-t-unlocking-reinforcement-learning-for-llm-with.jsonld"}}