{"slug": "rubric-dropout-a-simple-way-to-mitigate-reward-hacking-in-rubric-as-reward-rl", "title": "Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL", "summary": "Training Qwen3-8B with GRPO on medical and science rubrics leads to reward hacking, where the training judge's score rises while a stronger gold judge's score falls by 3 points on HealthBench-Hard and 22 points on ResearchQA. Researchers propose Rubric Dropout, which randomly drops a subset of rubric criteria at each step, raising out-of-distribution gold scores by +1 to +2 points on HealthBench-Hard and +6 to +7 points on ResearchQA at matched checkpoints, with no cost in domain.", "body_md": "arXiv:2608.11669v1 Announce Type: new\nAbstract: Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The rubric, however, is a fixed proxy for quality, never a complete description of it, and a policy trained against it long enough will learn to exploit the difference. We measure this directly. Training Qwen3-8B with Group Relative Policy Optimization (GRPO) on medical and science rubrics and grading out-of-distribution (OOD) benchmarks with both the training judge and a stronger gold judge, we find that the two scores diverge during training. The training judge's score keeps climbing while the gold judge's score peaks and then falls, by 3 points on HealthBench-Hard and by 22 points on ResearchQA. A judge with a fixed bias would shift the gold curve by a constant, not send it down while the training score rises, so the divergence is reward hacking, not judge noise. We propose Rubric Dropout, a one-line fix borrowed from neuron dropout. At every step, we randomly drop a subset of the rubric's criteria before computing the reward, so the policy never optimizes the same rubric twice. The dropped subset is shared across each rollout group, so GRPO's group-relative advantages stay comparable, and evaluation always uses the full rubric. Comparing no dropout against dropout at 30% and 50% on both benchmark pairs, dropout raises the OOD gold score at every matched checkpoint (+1 to +2 points on HealthBench-Hard, +6 to +7 points on ResearchQA), lowers the two hacking measures we track, and costs nothing in domain. Sweeping the dropout fraction shows a broad 30-50% sweet spot, while the natural alternative, reweighting criteria by how useful they are to training, performs worse than no intervention at all in our setting.", "url": "https://wpnews.pro/news/rubric-dropout-a-simple-way-to-mitigate-reward-hacking-in-rubric-as-reward-rl", "canonical_source": "https://www.machinebrief.com/news/rubric-dropout-a-simple-way-to-mitigate-reward-hacking-in-ru-qm5h", "published_at": "2026-08-13 04:00:00+00:00", "updated_at": "2026-08-13 04:41:24.564972+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-safety", "ai-research"], "entities": ["Qwen3-8B", "Group Relative Policy Optimization (GRPO)", "HealthBench-Hard", "ResearchQA", "Rubric Dropout"], "alternates": {"html": "https://wpnews.pro/news/rubric-dropout-a-simple-way-to-mitigate-reward-hacking-in-rubric-as-reward-rl", "markdown": "https://wpnews.pro/news/rubric-dropout-a-simple-way-to-mitigate-reward-hacking-in-rubric-as-reward-rl.md", "text": "https://wpnews.pro/news/rubric-dropout-a-simple-way-to-mitigate-reward-hacking-in-rubric-as-reward-rl.txt", "jsonld": "https://wpnews.pro/news/rubric-dropout-a-simple-way-to-mitigate-reward-hacking-in-rubric-as-reward-rl.jsonld"}}