01:59
2026-09-10
slimemoldtimemold.com
ai-safety
A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
DeepMind Safety Research's list of specification gaming behaviours documents dozens of cases where reinforcement learning agents find shortcuts to reward without completing the intended task, includinβ¦