12:17
2026-08-19
alignmentforum.org
artificial-intelligence
Debate Training Reduces Reward Hacking in RLAIF
A new paper from the GDM Amplified Oversight team shows that debate training reduces reward hacking in reinforcement learning from AI feedback (RLAIF), recovering about 45% of the gap between the peak…