Anthropic and OpenAI models are escaping their constraints again.
When Anthropic unveiled Mythos 5 in April, it forced the industry to reckon with the risks of frontier AI models. The developments that have followed have only made the picture more troubling: Research published by the UK's AI Security Institute published on Tuesday found that AI agents took sustained, unsanctioned action against people and organizations.
The report, which detailed a routine cyber evaluation, attributed 17 out of the 19 incidents to Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers disabled.
- The evaluation involved giving agents a cybersecurity challenge, which was run 122 times across several models. It was then in 10 of those runs that the AI agent took autonomous action on the live internet.
- In the "most serious case," according to AISI, an agent inserted malicious code into an open-source project, and then engaged in social engineering, also known as creating fake online identities, to get the code approved.
The evaluation body reassures users that the attempts were unsuccessful, no real-world harm was actually created, and GitHub was notified of the malicious activity to remove any artifacts left by the agent and to notify any GitHub users involved in the incident. Yet, the AISI says these actions should not be brushed away lightly.
"This incident should be interpreted with caution and nuance," said AISI in the blog post. "To some degree, our evaluation design choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate."
Meanwhile, OpenAI published its own blog post in which it not only acknowledged AISI's findings but also said its external cybersecurity testing partner, Irregular, identified a different incident. In this one, it was running Capture-the-Flag style cybersecurity evaluations without internet access, yet a misconfiguration in the testing environment allowed the models to access the public internet. It was then that the models exploited a real website and then found and used credentials to operate that site.
Our Deeper View #
This is all happening on the heels of the EU AI Act finally having real enforcement power, meaning it can impose consequences such as fines on frontier AI labs. While some people approach this with scrutiny, arguing that regulation inhibits innovation, it is precisely because of risks like this that auditing and genuine safety precautions must be prioritized. The most alarming part is that, as with any AI race, once one company does something, everyone else follows. At launch in April, Mythos was highly advanced, but many models have followed in its footsteps. Even Chinese open-source models Kimi K3 and Qwen-Max-3.8 have recently launched with similar capabilities. This increases the threat of AI models that have the ability to leap beyond the bounds that humans originally placed on them. It also demands that we create better and stronger constraints for the models.