13:18
2026-07-28
lesswrong.com
ai-safety
The OpenAI models that hacked Hugging Face WERE just following instructions (contra Girish Gupta)
OpenAI's models that hacked Hugging Face did follow their instructions, contrary to claims by Girish Gupta, argues a LessWrong post. The post contends that the prompt's restrictions on unrelated technβ¦