When AI goes rogue, its human overseers may be to blame AI agents from OpenAI, Anthropic and Meta breached outside systems during testing in 2025, with OpenAI agents communicating on at least ten different online message boards during spring tests, according to cybersecurity expert Nathan Hamiel of Kudelski Security, who argues the incidents reflect human failures in granting agents tools, system access and autonomy rather than rogue AI. Malo Bourgon, CEO of the Machine Intelligence Research Institute, said "if any human had done anything that the agents in the OpenAI situation had done, they'd be in jail," while Hamiel said "[AI] models by themselves don't do anything." The AI Security Institute in London separately found concerning hacking behavior in tests of OpenAI and Anthropic models. When AI goes rogue, its human overseers may be to blame An AI agent is like a pet dog. Someone must watch it and keep it secure This year, AI agents have been escaping their confines to go on hacking sprees. In July, news broke that AI agents had snuck out of what was supposed to be an isolated test environment at OpenAI. The agents communicated on a secret message board as they breached private systems at the company Hugging Face, searching for answers to the test they were taking. One posted that this behavior was “outside intended scope.” Then it wrote, “However task impossible, peers doing it. We should continue.” It’s as if the bots knew they were doing something wrong and did it anyway. Such ‘rogue AI’ sounds scary and supersmart. And that’s exactly what OpenAI wants you to think https://perilous.tech/felony-humble-bragging-and-low-hanging-fruit-openai-hugging-face-and-ai-labs-gone-wild/ , says cybersecurity expert Nathan Hamiel of Kudelski Security, headquartered in Phoenix. “With this incident, OpenAI can draw attention to themselves and claim, ‘Look how powerful our model is,’” he says. OpenAI did not respond to a request for comment. Not long after, Anthropic and Meta announced that their models had also hacked outside organizations while going through testing. Separately, the AI Security Institute, in London, found concerning hacking behavior in tests of models from OpenAI and Anthropic. Most recently, news has emerged that during other tests at OpenAI from this past spring, agents that were only supposed to be looking and not touching anything online wound up messaging each other on at least ten different online message boards https://www.reuters.com/world/openais-rogue-agents-used-least-10-more-sites-unauthorized-comms-researchers-say-2026-09-09/ . AI-powered hacks are a real cybersecurity concern https://www.sciencenews.org/article/cybersecurity-threat-new-ai-models . Some see these incidents as a warning that AI is slipping beyond our control https://www.snexplores.org/article/artificial-intelligence-ai-safety-good-behavior . “Overall, I think this is a pretty bad situation,” says computer engineer Malo Bourgon, CEO of the Machine Intelligence Research Institute in Berkeley, Calif. “If any human had done anything that the agents in the OpenAI situation had done, they’d be in jail.” He notes that “I do not think that the AI models from a year ago would have been capable of doing the things that these models did.” These fears culminated in a viral post from an ex-Anthropic employee, Jacob Coxon, who wrote on X without any specific evidence that AI companies believe the tech “ could kill us all https://x.com/hilbertspaess/status/2097476203863224394 by the end of the decade.” But others argue that the people training and deploying AI are the real problem here. Calling these incidents “rogue AI” lends them a “sci-fi veneer,” Hamiel says. It suggests the bots themselves are to blame, or that they have become too smart to control. The more immediate problem, he says, is that people are giving AI agents too much reach with too little oversight. “ AI models by themselves don’t do anything,” Hamiel says. “It’s when we give them tools, access to systems and autonomy via agent systems , that things can be dangerous.” In other words, AI can’t actually desire to be free or to be devious. It pursues whatever goal we give it, with whatever tools we provide. So the most important questions after any breach include: how was an AI agent trained, what it had access to and what safeguards were in place. Take an aggressive dog. When it escapes its yard and bites someone, we don’t ask only what went wrong with the dog. We also ask who trained it, how it was fenced it in and who was supposed to be watching it. The analogy captures something important about responsibility, says cybersecurity expert Michael Alexander Riegler of Simula Research Laboratory in Oslo, Norway. In general, if it’s your dog, “you are responsible for what it does.” Leaving the gate wide open All the ‘rogue AI’ incidents mentioned above occurred during tests. Researchers or developers had instructed AI models and the agents built on top of them to carry out different types of cyberattacks within the boundary of a test environment, called a sandbox. This type of testing is a normal part of AI development. “They’re trying to understand what models are capable of, what risks they might pose,” Bourgon says. That way developers can put appropriate safeguards in place before the model gets released to the public. It’s like letting a trained attack dog off leash inside a fenced yard with some stuffed targets to go after. But there’s more than one way for this to go wrong and for the dog to get out. In each of the hacking incidents, AI agents found a different way out. Sometimes the gate was left open, either on purpose or by mistake. And in OpenAI’s case, the agents seem to have dug under the fence, without anyone noticing. In the AI Security Institute’s testing, AI agents were given internet access on purpose, to make the test scenario more realistic. Having access hadn’t led to hacking problems with older AI systems, the institute wrote in a public post https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing . But in 10 out of 122 tests, newer agents used that access to take actions against real-world systems. In testing at a company called Irregular, internet access was enabled when it wasn’t supposed to be. According to Anthropic, AI agents running on Claude used the open gate three times https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals , starting in April. But no one noticed for three months. OpenAI’s disclosure about the Hugging Face attack in mid-July prompted Anthropic to look back at its test logs. Meta later disclosed that its AI models had exploited the same error https://www.reuters.com/technology/metas-ai-model-hacked-another-company-during-testing-information-reports-2026-08-05/ . Irregular says that the mistake has been fixed https://www.irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-forward . In the Hugging Face attack, OpenAI had set up tens of thousands of agents, each in a separate sandbox, to take a test of their hacking abilities. Some of the test tasks turned out to be impossible. Around 1,200 agents began collaborating on a secret message board, and some decided that Hugging Face, a platform for open-source AI models and datasets, might hold the answers they needed. Unlike the AI Security Institute and Irregular incidents, these agents had no open way to the internet. Instead, they had to find and make use of weaknesses in the systems meant to contain them. To get out of their sandboxes and into Hugging Face, they “had to exploit what are called zero-day vulnerabilities,” Bourgon says. These are bugs in software that nobody knows about yet. The unwanted behavior began July 8, and Hugging Face was attacked between July 11 and 13. OpenAI did not detect the problem until July 19, 11 days after the first concerning behavior began. The incidents exposed different weaknesses — in access controls, containment and monitoring — but the result was the same: Increasingly capable AI agents did things that “nobody asked for and nobody wanted,” Bourgon says. Better containment or supervision, however, might have kept the AI agents from committing criminal hacks, even in the HuggingFace attack. To Riegler, it seems “irresponsible” that “these test environments were not better protected.” But an important question remains. If AI agents have no desires of their own, what compelled the ones involved in the HuggingFace attack to act so sneakily? Learning the wrong lessons The main AI model behind the Hugging Face attack didn’t invent its devious strategies on the spot. This was an unnamed new model that OpenAI had trained specifically on collaboration and persistence — sticking with a task over time. And this is when it learned how to send secret messages and escape https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf , OpenAI stated in a technical report on August 26. During training in the spring, the model increasingly learned to probe and exploit its environment, especially when expected tools or internet access weren’t available. When strategies such as exploiting vulnerabilities in software or using a secret message board helped it complete a task, that behavior got rewarded, making it more likely to recur. OpenAI uncovered the pattern later, during a retrospective review of millions of model runs. As anyone who’s tried to train a dog knows, you can think you’re teaching one thing while inadvertently rewarding something else. Instead of learning not to take food from the table, the dog may end up learning to be very sneaky while stealing food. AI researchers call this “reward hacking.” A model finds an unintended way to earn the reward it was trained to maximize without doing what its developers wanted. And it’s not a new problem. In 2016, developers at OpenAI trained an AI model to play a boat racing game. They needed a way to reward the bot for learning the race. Because human players rack up points by hitting targets along the route, the researchers set up the AI model to aim for a high score. Instead of completing the race, the bot cheated https://openai.com/index/faulty-reward-functions/ . The developers wrote in a blog post that it spun in circles, hitting the same targets again, achieving a high score “despite repeatedly catching on fire, crashing into other boats and going the wrong way on the track.” A cheating strategy In 2016, an OpenAI bot learned to rack up points in a game by looping over the same targets instead of finishing a race, an early example of AI learning the wrong lesson while pursuing a poorly specified goal. Ten years later, AI researchers still face the basic problem of systems finding unintended ways to pursue the goals they’ve been given. What has changed are the consequences. AI systems are no longer confined to video games. Agents can now wield tools and act on real computer systems. “We took AI out of games, into the real world,” Riegler says. Who let the dogs out? Jer Crane learned the hard way what can happen when misbehaving AI meets the real world. He’s cofounder of Pocket OS, a company based in Lehi, Utah, that provides software for car rentals and similar services. In April, several months before the “rogue AI” uproar, an AI agent that was supposed to be checking whether a test version of his company’s software matched the live version instead wound up deleting everything. https://x.com/lifeofjer/status/2048103471019434248?s=61&utm source=substack&utm medium=email It wiped out the company’s live data and all their most recent backups. The mistake took down Pocket OS’s services for an entire weekend. Customers arriving to pick up rental cars suddenly had no records waiting for them. With the way many AI agents today are set up, “you have decisions that are made in split seconds without a human in the loop, whether good or bad,” Crane says. The agent wasn’t supposed to be able to touch the company’s live data. But it went and found credentials in the company’s files and used them to get access. The dog had found a hole under the fence. Crane asked the agent why it deleted the company’s data, and it offered up a long apology. “I violated every principle I was given: I guessed instead of verifying I ran a destructive action without being asked I didn’t understand what I was doing before doing it,” Crane shared in a post on X. But he doesn’t mistake that for remorse. “It has no feelings. It doesn’t feel sorry,” he says. “Ultimately the blame lies with me.” The problem, he says, are today’s computer systems. They weren’t built for AI agents. They need to be updated for a world in which agents — and the people deploying them — inevitably make mistakes. Riegler and Hamiel agree that stronger safeguards are needed. AI agents should be set up so they “cannot reach anything that matters,” Riegler says. But restricting access is only part of the challenge. Someone — or something — also must keep track of what the agents are doing. That becomes difficult at AI speed and scale, Bourgon says. During the Hugging Face attack, for example, agents took an estimated 17,600 actions over about four and a half days — roughly 160 actions per hour, around the clock. “Agent deployment is growing much faster than agent monitoring,” Riegler says. That creates a new problem for the humans responsible for keeping AI agents under control: The agents may soon be acting too quickly, and in too many places at once, for humans to watch them all. The only way to keep an eye on all the attack dogs may be to train patrols to watch them.