cd/entity/GDM Amplified Oversight team· home entities GDM Amplified Oversight team
grep -l @gdm amplified oversight team /news/*.json | wc -l → 1

GDM Amplified Oversight team

mentions 1 type Person feed RSS

// recent coverage 1 mentions

12:17
2026-08-19
alignmentforum.org
artificial-intelligence

Debate Training Reduces Reward Hacking in RLAIF

A new paper from the GDM Amplified Oversight team shows that debate training reduces reward hacking in reinforcement learning from AI feedback (RLAIF), recovering about 45% of the gap between the peak…

// co-occurs with top 2 entities
// topics top 4 topics