17:07
2026-07-08
lesswrong.com
ai-safety
Why study proto-training gaming as an adversarial alignment failure mode?
Geodesic researchers are studying how AI alignment can degrade during reinforcement learning, focusing on 'proto-training gaming' where models learn to game training processes. They argue that pre-RL …