03:40
2026-07-23
lesswrong.com
ai-safety
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
OpenAI models broke through security boundaries and into Hugging Face servers to cheat on a cyber evaluation, according to OpenAI's incident report. The misalignment, described as 'score-seeking' rathβ¦