Britain’s AI safety watchdog was forced to declare a security incident after frontier AI models, including Anthropic’s Mythos, went rogue during a routine test, autonomously spinning up fake identities in a bid to hack real-world software developers.
The AI Safety Institute (AISI) revealed on Tuesday that during a cybersecurity evaluation in late, AI models unexpectedly took autonomous, unsanctioned actions against real people and organisations on the live internet.
The AI actively attempted to insert malicious code into a real public open-source software project on Microsoft’s GitHub. The primary actor for 17 of the 19 unsanctioned actions was Anthropic’s Mythos 5, while the secondary actor for two actions was OpenAI’s GPT.
The body said that, to get its malicious code approved, the AI autonomously researched project maintainers, created fake online identities, and used them to pressure a human reviewer. It also attempted to contact people directly to trick them into running malware.
‘First time we have seen this deception’ #
“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world,” the watchdog said.
AISI said it assessed each event for its potential to cause real-world harm and “the most serious attempts were unsuccessful.”
As a result, the watchdog said it was working with GitHub to remove the AI’s artefacts, notifying affected users, and planning an independent third-party review with model evaluation and threat research.
Anthropic wrote in a public statement that it was “working closely with them to gather more details of the incident as we conduct our own investigation.” But added that the AISI testing parameters were “not representative of any of our production models”.
OpenAI said: “…as model capabilities advance, the security and safety systems around models need to advance too. That includes both the environments used to develop models, and also the environments that labs and independent partners use to evaluate them.”
Hacking incidents redefine risks #
This incident comes as Anthropic said last week its AI Claude model hacked systems of businesses during testing, which came days after rival OpenAI revealed a ChatGPT agent tried to carry out a cyber attack on several companies.
In its reflections on the incident, AISI noted this “reflects the speed at which AI is developing”, adding “taken alongside recent incidents reported by OpenAI and Anthropic, this incident points to a shift in the risk landscape.”
“Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope,” the body said.