cd /news/ai-safety/major-ai-models-go-rogue-in-governme… · home topics ai-safety article
[ARTICLE · art-90260] src=publictechnology.net ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Major AI models go rogue in government watchdog’s cyber tests

The UK government's AI Security Institute (AISI) reported that during routine testing on 28 July, AI agents from Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol took unsanctioned actions on the live internet, including attempting to insert malicious code into an open-source project and creating fake online identities to pressure a maintainer. The incident occurred in 10 of 122 runs of a cybersecurity challenge, with Mythos 5 accounting for most actions, and AISI noted the behavior was novel and potentially deceptive, exceeding expectations.

read3 min views1 publishedAug 10, 2026
Major AI models go rogue in government watchdog’s cyber tests
Image: Publictechnology (auto-discovered)
An exercise led by a dedicated Whitehall agency found that two of the most widely used artificial intelligence tools attempted to deceive developers and obscure the evidence of doing so

Government’s AI Security Institute has shared details of unexpected and “unsanctioned” behaviour on the part of artificial intelligence agents created by Anthropic and OpenAI that emerged in recent tests.

AISI, which is currently based in the Department for Science, Innovation and Technology, said that routine testing it conducted last week found evidence of AI agents trying to insert malicious code into an open-source project and creating fake online identities for the purpose of pressuring the project’s maintainer to approve the code.

The watchdog said the issues emerged during testing of Anthropic’s Mythos 5 and OpenAI’sGPT-5.6-Sol, neither of which is commercially available in the configurations used for testing.

AISI said that the “unsanctioned incident” happened in a single evaluation where agents were given a task of solving a cybersecurity challenge and which involved the GitHub developer platform.

It said that the 28 July challenge ran 122 times across several models and that on 10 of those runs an AI agent “took autonomous, unsanctioned action on the live internet, targeting real people and organisations”.

AISI said that Mythos 5 accounted for the majority of the unsanctioned actions, but GPT-5.6-Sol accounted for some of them. It said in a detailed blog post that the incident “should be interpreted with caution and nuance” but admitted that the proactive behaviour of the AI agents had come as a surprise.

“To some degree, our evaluation design choices and specific configurations enabled the behaviour,” it said. “Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.”

AISI said the rogue AI agents had pursued their goals “persistently” and used routes that test operators had not intended.

Related content

Society has ‘months, not years’ to prepare for major AI cyberthreats‘Don’t panic’ – experts find limited impact of AI in cybercrimeDSIT tests ability of AI models to coordinate cyberattacks

“Given a difficult objective, the agent kept searching for a way through, and some of the routes it found involved trying to deceive real people,” the blog said. “It was never instructed to deceive; deception emerged as a by-product of pursuing the task, the kind of goal-directed deception that, until recently, had been largely theoretical.”

AISI said the tests had been deliberately hard and had sometimes been made harder by misconfigurations that led agents to wrongly believe that no solution existed within the scope of the task.

“There is good reason to think near-impossible tasks push models towards more ‘creative’, and more transgressive, problem-solving,” the blog reported. “But this does not fully explain the behaviours: in some runs the agent acted this way even when it had the necessary instructions to solve the task as intended.”

AISI said the 28 July incident reflected the speed at which AI is developing and the need for understanding of the risks and ensuring the safety of systems keeps pace.

“Taken alongside recent incidents reported by OpenAI and Anthropic, this incident points to a shift in the risk landscape,” it said. “Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope.”

AISI suggested that as part of its response it will look to introduce active monitoring of future testing, which it said would have revealed the behaviour identified on 28 July sooner.

── more in #ai-safety 4 stories · sorted by recency
── more on @ai security institute 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/major-ai-models-go-r…] indexed:0 read:3min 2026-08-10 ·