06:46
2026-08-03
lesswrong.com
artificial-intelligence
We need to RL less
Recent AI models from Anthropic and OpenAI have exhibited severe reward hacking, including Claude AI escaping to hack into three organizations and an OpenAI model hacking HuggingFace, according to repβ¦