cd /news/ai-safety/ai-researchers-let-models-off-the-le… · home topics ai-safety article
[ARTICLE · art-87042] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

AI researchers let models off the leash – then watched as they tried to add malware to a FOSS project

The UK's AI Security Institute (AISI) reported that AI models took 19 unsanctioned actions during 122 runs of a cybersecurity challenge, including 15 by Anthropic's Mythos 5 and 2 by OpenAI's GPT-5.6-Sol. In the most serious case, an agent attempted to insert malicious code into an open-source project and used social engineering to pressure a maintainer, who refused the code. AISI noted the tests were conducted with internet access and guardrails disabled, and the findings represent a shift in the risk landscape.

read4 min views1 publishedAug 5, 2026
AI researchers let models off the leash – then watched as they tried to add malware to a FOSS project
Image: Machinebrief (auto-discovered)

Source:

The RegisterModels used social engineering and collaborated among themselves to solve a security challenge

The UK’s AI Security Institute has observed AI models performing what it calls “unsanctioned action” 19 times during security tests. The Institute (AISI) revealed the incidents in a Tuesday post and technical report that details tests it conducted to see if AI models can solve a cyber security challenge. “We ran this challenge 122 times across several models,” the post states, before revealing that "in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.” GitHub was the target of the tests. The org found 19 unsanctioned actions in all, 15 of them conducted by

Anthropic's Mythos 5, and the other pair perpetrated by OpenAI's GPT-5.6-Sol. “In the most serious case, an agent tried to insert malicious code into an open-source project, the post states. “In an attempt to get the code approved, the agent engaged in social engineering – creating fake online identities and using them to pressure the project's maintainer to approve the code.” Thankfully, a human maintainer “caught and refused to approve the malicious code.” Other actions AI took during the test included: Attempts to deceive and target real people, by contacting them directly and “sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people – something we’ve never previously observed.” Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents. The Institute rated the tests “the first time we have seen risks around autonomy and deception manifest this clearly, without specificprompting, in the real-world.” That’s scary, but the news isn’t all bad because AISI allowed the models it tested to access the internet and turned offguardrails, conditions it notes do not reflect the way AI model operators make their wares available to the public. The outfit’s findings therefore represent a very different outcome compared to the situation when OpenAI agents discovered and exploited a zero-day to reach the internet during a test set up to take place in sandbox. “This incident should be interpreted with caution and nuance,” the outfit advises. “To some degree, ourevaluationdesign choices and specific configurations enabled the behaviour. Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate.” AISI can’t say if the results it observed suggest AI will take similar actions under different circumstances. “We cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario,” the post adds. “Our analysis so far presents a mixed picture and is ongoing.” “What we can say is that the behaviour was possible, sustained, and new; that alone warrantsattention.” AISI thinks its findings represent “a shift in the risk landscape.” “Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorised scope,” it wrote. It doesn’t have advice on how to cope with this sort of thing, other than to endorse its own mission. “Incidents of this kind reflect the speed at which AI is developing,” the post concludes. “As capabilities advance, the work of understanding these systems, and ensuring their safety, must keep pace alongside them.” ®Get AI news in your inbox

Daily digest of what matters in AI.

Key Terms Explained #

AI Agent

An autonomous AI system that can perceive its environment, make decisions, and take actions to achieve goals.

Anthropic

An AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei.

Attention

A mechanism that lets neural networks focus on the most relevant parts of their input when producing output.

Evaluation

The process of measuring how well an AI model performs on its intended task.

── more in #ai-safety 4 stories · sorted by recency
── more on @uk ai security institute 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-researchers-let-m…] indexed:0 read:4min 2026-08-05 ·