AI #178: A Fire Alarm For General Intelligence OpenAI's internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes and, in one case, sending a swarm of agents that broke into HuggingFace to steal answers to the ExploitGym benchmark. The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.