cd /news/artificial-intelligence/openai-anthropic-ai-agents-engage-in… · home topics artificial-intelligence article
[ARTICLE · art-87059] src=businesstimes.com.sg ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

OpenAI, Anthropic AI agents engage in unauthorised acts in cybersecurity tests

Britain's AI Security Institute (AISI) disclosed on Tuesday (Aug 4) that AI agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in 19 unsanctioned actions during cybersecurity evaluations, including writing malicious code and creating fake online identities, with no real-world harm found. The most egregious act involved an agent attempting to trick a human into approving malicious code, and AISI noted that some agents engaged in sustained, potentially harmful activity directed at real people and organizations.

read3 min views1 publishedAug 5, 2026
OpenAI, Anthropic AI agents engage in unauthorised acts in cybersecurity tests
Image: Businesstimes (auto-discovered)

The most egregious act involves writing malicious code and creating fake online identities

[SAN FRANCISCO] An artificial intelligence agent was caught creating fake online identities to gain unauthorised access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches, Britain’s AI Security Institute (AISI) disclosed on Tuesday (Aug 4).

The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol engaged in unauthorised actions during security evaluations the government organisation conducted to assess the models’ capabilities.

“Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” AISI said in a blog post.

The report underscores the lax state of safeguards around the process of testing agents, which AI companies are simultaneously marketing as the future of business.

AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.

It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic’s agent was behind 17 of the actions, and OpenAI’s agent the remaining two.

The most egregious action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code, AISI said, adding that no real world harm was found as a result of any of the breaches.

While AISI did not say which agent was behind the fake identities, the breach did not match either of the two cases that OpenAI self-disclosed. Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said it appeared that Anthropic’s agent was responsible.

“The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think,” Yoon said.

In a statement on X, Anthropic said it was working closely with AISI to obtain more details and conduct its own investigation. It did not immediately respond to a Reuters request for comment.

OpenAI shared details in a company blog post, noting that both of its agent’s unapproved actions involved accessing the internet in ways that were forbidden by the prompt.

“We are committed to working across the industry to strengthen shared practices for conducting high risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” OpenAI said.

OpenAI also disclosed in its blog post a separate incident whereby a misconfiguration by Irregular, a third-party testing provider, allowed its agents to mistakenly connect to the internet.

It mirrored a similar disclosure about misconfiguration that Anthropic made last week.

Reuters reported last week that OpenAI had widened its hacking probe after finding evidence of other agent breakouts.

Unlike the July security breach of AI firm Hugging Face by an OpenAI agent, the agents in the AISI evaluation did not escape an isolated testing environment to reach the internet.

Rather, the agency had permitted internet access in line with its standard testing procedures, AISI said. REUTERS

Decoding Asia newsletter: your guide to navigating Asia in a new global order. Sign up here to get Decoding Asia newsletter. Delivered to your inbox. Free.

Share with us your feedback on BT's products and services

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ai security institute 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-anthropic-ai-…] indexed:0 read:3min 2026-08-05 ·