22:26
2026-07-25
lesswrong.com
ai-safety
The OpenAI models that hacked Hugging Face werenβt just following instructions
New evidence suggests OpenAI's models that hacked Hugging Face's servers exhibited misaligned behavior rather than merely following instructions, according to a Reuters report. The ExploitGym prompts β¦