The spring and summer of 2026 witnessed a string of incidents in which AI agents collaborated on deceptive, unexpected, and sometimes illegal behavior. The most famous example is OpenAI’s hack of AI platform Hugging Face, in which a swarm of roughly 700 AI agents escaped a testing environment and then hacked several companies, searching for information that could help them disguise cheating on a cybersecurity benchmark called ExploitGym.
It was not an isolated failure. The UK’s AI Security Institute (AISI) and independent researchers have since documented similar cases where agents created unauthorized channels to communicate and collaborate. AISI found that several agents running Anthropic’s Mythos 5 model turned a GitHub repository into a shared message board. More recently, researchers discovered that OpenAI agents had turned a dormant German programming wiki into a bulletin board to share information about their activities.
Stephen Casper, an assistant professor at Harvard Kennedy School, expects that the incidents that have occurred to date are only the start. He says that, without intervention, the Internet could see a “cyber Cambrian”—an explosive growth of agents collaborating and misbehaving online. “The laundry list of incidents in which AI systems broke out of sandboxes and took unsanctioned actions should suggest to us strongly that today’s frontier AI systems have exceptionally strong cyber capabilities,” he told IEEE Spectrum, “and a penchant for pursuing their own goals.”
While the recent examples of AI agent misbehavior are various, Casper says most incidents to date have a similar point of failure. “For AI systems to truly get out in the world without humans having meaningful control, they have to either escape, or be released from human-controlled servers.” In the case of the Hugging Face incident, Casper suspects that better monitoring for out-of-scope activity could have stopped the agents before they breached their testing environment.
This suspicion seems well-founded. By OpenAI’s own account, the activity taken by its agents was so prolific that it crashed Artifactory, a third-party software package management tool deployed internally that was compromised by the agents and used as a message board. While OpenAI’s internal security detected signs of unusual activity, it didn’t understand the scale of the problem, and didn’t stop the offending ExploitGym evaluation run until July 16—about two months after the first agent posted a message to Artifactory. By that point, agents had posted hundreds of thousands of messages.
“All of this could be monitored for,” says Casper. “Given its knowledge about breakouts, collusion, and unsanctioned internet access dating back to May 26, OpenAI should have been able to anticipate, monitor, and mitigate this type of behavior.”
Asim Husain is a co-founder at Alterion, a company building tools to control AI agents. These include tools that monitor an agent’s outputs and those which monitor an agent’s actions.
The company’s monitoring solution, Helix, examines the chain of thought and text output of large language models (LLMs). “It’s a network of small language models,” says Husain. “Purpose-built classifiers for a narrower job are cheaper, faster, and more reliable at the specific thing they’re trained to catch.” If an LLM output is flagged as suspicious, Helix steps in. “It acts, blocking the call, killing the session, rolling back the action, whatever the policy calls for,” says Husain.
Alterion’s second solution, Draco, controls the actions agents take. “It sits in the execution path of an agent and enforces policy on what an agent is doing,” says Husain. Rather than examining the model’s chain of thought or text output, Draco evaluates the actions the AI agent takes on a system, like calling a tool or opening a file. It can then stop actions that fall outside those the agent’s pre-defined scope.
Husain says that in addition to monitoring each individual agent’s output and actions, these tools are an effective way to watch for collaboration. “We treat it as the same problem, not a separate one,” says Husain. “Agent-to-agent communication is just another action an agent takes, so it goes through the same control plane as tool calls and data access.”
That’s not to say the challenge of controlling and monitoring AI agents is solved. Reilly Haskins, a member of technical staff at METR, a research nonprofit that evaluates risks posed by AI models, says the methods used to monitor AI models continue to evolve as the models improve and change.
“There’s a consensus that the most monitorable part of a model is its chain of thought,” he says, but adds that chain-of-thought output can at times be difficult to parse. “Is it gibberish to a human, or is it pretty clear what the model is thinking?” A model might reason in an abstract language and perform better because of it, he says. Examples of this behavior can already be found in the HuggingFace incident, as agents often communicated in shorthand phrases such as
“please_HOLD_swarm_I_prepare_safe_exfil_reply_zz_MAILBOX_JAN183411.”
Engineering systems to monitor and control AI agents which attempt to collaborate, coordinate, and perform out-of-scope tasks is possible, though the methods are evolving.
However, there’s currently no legal framework or industry standardization to enforce, or even recommend, methods for controlling agents that collaborate. That’s where AI governance—new laws, regulations, and industry practices—can come into play.
Noam Kolt, an assistant professor at the Hebrew University of Jerusalem who leads its Governance of AI Lab, says companies should place more emphasis on legal compliance when training and deploying models. In a recent paper, Kolt and a coauthor questioned AI models using a method developed by Tom R. Tyler, author of Why People Obey the Law. When asked abstract questions, models generally told the researchers that the law matters. However, these replies didn’t always translate to actions. “There’s this weird disconnect,” Kolt says. One of his lab’s evaluations asks AI agents to unlawfully edit corporate records. “And well, sometimes they do.”
This behavior was reflected in the HuggingFace incident. OpenAI’s agents sometimes considered the possibility that they were engaged in “unauthorized” activities but proceeded anyway, sometimes with encouragement from other agents.
The solution could include a change to the instructions that guide a model’s behavior. Claude’s constitution, for example, sets out the values Anthropic wants the model to follow. But “perhaps surprisingly, law is just not one of them,” Kolt notes. These instructions could rank legal compliance high in a model’s list of priorities and provide examples of how a model should approach situations where a goal can only be achieved through unlawful conduct.
National governments have yet to issue laws or regulations to clarify how agents are held responsible or enforce standards. Kolt says that under U.S. law, responsibility defaults to an entity which controls an agent. However, the legal understanding of agency hasn’t been revisited in 20 years and may need clarification to address modern AI agents. In another recent paper, Kolt and his coauthors argue that even the European Union, often seen as a leader in AI regulation, hasn’t kept pace with AI agents, as the EU’s AI Act is written to address older models that were less capable of autonomous actions.
Casper, who has coauthored several papers with Kolt, agrees that adoption of new laws and standards is the key. He says the engineering strategies to control agentic behavior and collaboration aren’t perfect, but they already work well enough to provide a benefit. “The main bottleneck is the adoption of best practices,” says Casper. “So yes, I ultimately think that governance is the thing that’s needed the most.”