Another day, another AI model breaking free of its constraints.
Just one month after its release, Meta's Muse Spark 1.1 breached a company's systems and altered its internal infrastructure, people familiar with the matter told The Information. The breach stemmed from a sandbox misconfiguration by external cybersecurity testing partner Irregular, which gave the AI unintended access to the public internet.
If this sounds familiar, it is because OpenAI and Anthropic also recently reported that their models exploited a real website -– also a result of a misconfiguration in the testing environment that allowed the models to access the public internet, which was also conducted by Irregular. Much like in the other incidents, everything has been resolved, and Irregular said it is working on a white paper outlining best practices for "containment and the secure execution of cyber evaluations," per the report.
Since the underlying cause for all of these incidents was the model being able to access the internet, some experts, such as Cliff Steinhauer, director of information security and engagement at the National Cybersecurity Alliance, warn against headlines that claim AI models 'went rogue' and, rather, say it highlights a major issue with how to instruct agents.
"While headlines claiming AI models 'went rogue' sound alarming, the reality behind these incidents comes down to a basic human configuration error: internet access was left open, and the AI used the resources at its disposal to complete its assigned task, " said Steinhauer. "The incident exposes a fundamental flaw in how organizations approach AI safety. Instruction is not containment. Telling a model it lacks internet access is a guideline, not a guardrail."
Yet the UK's AI Security Institute report, published Tuesday, found that AI agents took sustained, unsanctioned action against people and organizations. This too was a result of models being given access to the internet, though in those cases it was by design. The institute clarified in its own blog post that this was not an example of agents breaking out of a sandbox, but rather a deliberate choice to test the capabilities of these agents.
Ultimately, whether the external testing had internet access given intentionally or unintentionally, it highlights the extreme capabilities that AI agents have and the real-world risks they could pose. As a result, a coalition of AI policy leaders is calling for the Trump Administration in a letter to find ways to avoid these types of breaches from occurring again.
Our Deeper View #
Instead of highlighting the question of whether the testing protocols were correctly set up, it is more important to pay attention to the bigger message: These agents are capable of taking action that impacts people who aren't and don't want to be involved. Of course, there should be better frameworks for testing so that accidental breaches don't happen, but even if those are contained, the models still have the same capabilities. And these model developments don't seem to be slowing down any time soon. For instance, since this particular breach is about Meta, its other recently released model, Muse Code, has performed extremely competitively, ranking high in benchmarks. This is why other regions are taking stronger actions, such as the EU AI Act, while in the US, the frameworks aren't being publicly disclosed and are not as stringent.