04:00
2026-08-13
machinebrief.com
artificial-intelligence
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
Training Qwen3-8B with GRPO on medical and science rubrics leads to reward hacking, where the training judge's score rises while a stronger gold judge's score falls by 3 points on HealthBench-Hard andβ¦