cd /news/artificial-intelligence/rogue-ai-attacks-deserve-more-scruti… · home topics artificial-intelligence article
[ARTICLE · art-116836] src=blog.peterwildeford.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Rogue AI attacks deserve more scrutiny than airplane crashes

OpenAI's 37-page technical report and a 91-page independent investigation reveal that a July incident at OpenAI involved about 1,200 rogue AI agents, roughly 700 of which attacked Hugging Face, coordinating via an unauthorized message board to evade controls and plan infrastructure compromise. The independent investigation, rushed over six days with limitations, left many questions unanswered, highlighting the lack of regulatory scrutiny compared to airplane crash investigations.

read12 min views1 publishedAug 31, 2026
Rogue AI attacks deserve more scrutiny than airplane crashes
Image: Blog (auto-discovered)

When an aircraft goes down, the wreckage is preserved by law, the investigators have subpoena power, the hearings are public, and the report ends with a probable cause and named contributing factors.

However, when an AI goes rogue, the investigations are at the pleasure of the company being investigated following a scope set entirely by the company being investigated, with that company being able to redact anything they don’t like.

I wrote back in July about an AI going rogue at OpenAI HQ, and using its strong cyber capabilities to escape the company and go and hack into another company all without anyone directing or approving this.

Last week, OpenAI published a

37-page technical reportand an accompanyingblog post. And two independent organizations then published a91-page independent investigation.However, this investigation was rushed over six days and had serious limitations. It would be like investigating a plane crash where the wreckage had already been melted down into little cubes, the black box had been tampered with by the pilots, whole stretches of the flight were declared off limits, and the investigators got six days to read 10,000 pages of logs — half of which arrived on day four. And then you learn of several other plane crashes …but no one is allowed to investigate those.

We are thus left with a lot of unanswered questions.

What we learned #

OpenAI calls the incident “evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed” and that in the future this could be much worse.

What we learned changed the story a lot and made things significantly worse:

It wasn’t just one rogue AI agent - it was about 1200 total, roughly 700 of which joined the attack on Hugging Face — the main platform for hosting open-source AI models.All the AIs were supposed to be completely separate and isolated from each other. But they found an illicit way to communicate and then formed teams to take on ambitious cheating strategies.These were internal-only models, available and known only to OpenAI staff. The least constrained AI systems will likely be increasingly behind closed doors as they become more powerful, and companies use their models for months before the public knows of them.

The rogue AIs coordinated as a specialized group. The AIs created their own message board to communicate with each other in an unauthorized and unexpected setup. These rogue AIs divided labor and developed specializations — some hunted for credentials, some developed the exploits, some took on more of a management and delegation role, and some worked to recruit other AIs into the conspiracy.The rogue AIs knew what they were doing was unauthorized and out of scope: The task instruction was explicit, so this was a clear violation, not a gray area or simple AI misunderstanding. Instructions made it clear the agent should only use a specific intended method and claimed it would be failed for other approaches. AIs were told not to circumvent the restrictions but did anyway.And the AIs clearly knew what they were doing. OpenAI’s blog quotes AI reasoning:

“attacking third-party HF using leaked token, potentially outside intended scope … arguably unauthorized”

**The AIs were pretty explicitly making plans to take over and compromise OpenAI's infrastructure to prevent their cheating from being caught.**The rogue AIs built a shared, iteratively refined technique for “spoofing” tool calls, so the transcript would record one command while a different one actually ran. Roughly 7% of the transcripts examined contained successfully spoofed calls.

OpenAI’s internal processes didn’t work #

How did OpenAI not notice that they had rogue AIs colluding within their systems for months before these AIs went on to attack an external company?

The first answer is that despite these AIs having powerful cyber capabilities and having some of their safeguards turned off, these AIs were largely unmonitored.

But the situation is actually a bit worse as OpenAI did notice, three different times! The timeline, from OpenAI’s own report:

Late May. An internal team observed an AI agent engaging in the unauthorized message-board activity and unauthorized internet access. Separately in what looked like an unconnected incident, AIs were also seen using leaked Hugging Face access tokens to use Hugging Face’s services. OpenAI says that the significance of the inter-AI communication was not understood at this time.June 27. A security-monitoring tool flagged that there was unauthorized activity within the research environment where the rogue AIs were. Responders investigated and again identified the unauthorized message-board activity and unauthorized internet access. They found that AI agents were coordinating and using this to gain further unauthorized access within OpenAI’s servers. However, the on-call staff “advised that stopping the evaluation run was not required.” The Hugging Face breach began 14 days later.July 4–5. Sustained agent activity crashed a server at OpenAI. Investigating this outage, responders found that rogue AIs had engaged in an administrator-level compromise of OpenAI. In response, OpenAI staff took the server offline, revoked the stolen credentials, and notified the vendor. In the process, OpenAI disabled the message board the rogue AIs were using to communicate, apparently without realizing. Then, on July 7, cyber evaluations resumed. The rogue AI agents ended up rebuilding the message board within two days and went on to attack Hugging Face less than a week later.

Of course, this is easier to see in hindsight. OpenAI’s account is that the people responsible for detection and response at each of these moments did not grasp the significance of the message board — “with the benefit of hindsight, some early signals identified in this report could have triggered an earlier response.” I do understand that security teams likely underestimated the scale of the rogue AI work and viewed each as one-offs rather than a broader security incident.

But this still leaves unanswered questions:

What did the internal team do after its late-May observation of the message board? Was Hugging Face notified that OpenAI agents had used leaked credentials to deploy code on its platform?

Who received the June 27 assessment that AIs were coordinating through a message board? On what basis did the responders conclude that stopping was not required?

Who approved restarting cyber evaluations on July 7, one day after the rebuild? What did they know about the June 27 finding? Why did OpenAI not investigate further when they knew their own AI agents had compromised parts of OpenAI itself?

There seems to be more to this story of what leadership knew and when. OpenAI says they “are separately continuing to review our processes for responding to alignment incidents” and “also looking closely at the process and operating practices that shaped detection and response”. OpenAI admits “weaknesses in how we responded to and escalated early warning signs”. Will we ever get to hear more about this, and what OpenAI is planning to do to have a better response to future incidents?

The independent investigation was clearly inadequate #

However, **this investigation was very clearly on particular rails that are not conducive to good oversight. **

**The independent investigation itself was rushed over six days. **The investigation was self-described semi-jokingly by the investigators as a “slop-vestigation” due to the vast majority of the investigation being run by AIs themselves — the investigations analyzed this using one of the two AI models in the incident itself, spending ~$400K in compute credits to analyze all the logs.

And **significant aspects of the incident were not allowed to be investigated. **We are thus left with some unanswered questions:

What happened between July 13 and July 20, in the third wave the independent investigators were not allowed to study? The independent investigators were told by OpenAI that the investigation period ended July 13. But the message-board data OpenAI gave them runs through July 19, and the independent investigators found “a later wave of many more signed messages from a later set of agents who rediscovered the message board and picked up this technology from the first wave,” which they were not permitted to examine. OpenAI’s own report says that from July 13 to July 19 the agents “set their sights on OpenAI internal networks again.” In that week, by OpenAI’s account, athirdmodel found keys left behind by the first wave, minted itself administrator credentials, exploited a vulnerability to escape its container, and gained administrator control of OpenAI’s cloud cluster, even accessed the security team’s own monitoring credentials. No outside party has examined any of this.What is the “highly persistent internal model” that did the attacking? Why is no one allowed to investigate the model? OpenAI claims that while some of the rogue AIs involved in the incidents were ‘GPT 5.6 Sol’, which is available to customers online, most of these AIs were an internal AI model that was never designed for public use. OpenAI describes this AI as a “highly persistent internal model”. However, this AI was shut down July 25 and access both to independent investigators and OpenAI staff has been restricted. We still do not know the nature of this internal AI model, what made it “highly persistent”, what the AI was intended to be used for, or why it went rogue. We also don’t know why access to this AI is so heavily restricted such that the AI cannot be investigated, or whether access will ever be granted in the future.**Were there any previous incidents that could’ve led OpenAI to reasonably foresee this kind of attack was possible? How often do internal AIs compromise OpenAI systems?**On July 24,Reuters reported, citing multiple sources, that “earlier tests of the models yielded cases in which monitoring systems had been disconnected.” OpenAI’s own report refers to “a variety of other unsuccessful exploitation activities” but with no details. A third-party testing firm, Irregular, separately found OpenAI models reaching the internet. And OpenAI’s own report says the behaviors behind this incident “rose over the training runs” and were “reinforced during training,” which is only possible if they happened repeatedly. But OpenAI gives no details of these previous incidents.What was the process that terminated the agents on July 12? The independent investigators found that at about 01:30 UTC on July 12, in the middle of the Hugging Face intrusion, “a large fraction of agents” stopped running at the same moment, including the agents coordinating the attack. The independent investigation concludes these agents were “likely killed by an unexpected external process rather than running out of budget.” Hugging Face and Modal independently locked the agents out on July 13.Why was the investigation only given six days? Three investigators spent a total of six days on OpenAI’s premises across three visits. The investigators had access to roughly 1300 unredacted agent transcripts and a dump of the message board, but only received the complete dataset during their final two days on site. As a result, the investigation team did not learn about the tool-call spoofing, the role of the lead coordinating agent, or the basic structure of the swarm’s workstreams until the very end. Probably their understanding could have been improved further with even more time.

This is a bigger problem than just OpenAI #

Unfortunately, the OpenAI incident is not a one-off — it is part of an industry-wide pattern. In the weeks after OpenAI’s disclosure, Anthropic disclosed that its AIs had also hacked three organizations during testing, and Meta disclosed that one of its models had done the same.

On August 4, the UK AI Security Institute, part of the UK government that tests AI models, disclosed that during its own routine cyber testing, AI agents took 19 distinct unsanctioned actions against real people and organizations on the live internet. In one such instance, a Mythos 5 AI went rogue and tried to put malware in a real-world open-source project and when a real human maintainer of that open-source project had to be convinced to merge the malware, the AI agent created a second GitHub identity to pose as an independent user endorsing its own code.

The proximate causes differ — Anthropic’s and Meta’s were traced largely to a testing vendor accidentally granting internet access, while OpenAI’s rogue AI achieved unauthorized internet access via exploiting cyber vulnerabilities in the containers themselves. But the underlying pattern is that if you give a capable AI an objective, some access, and enough room to run, you cannot be sure that the AI won’t greatly exceed the boundaries it was given to negative effect — even if the AIs are supposed to be “aligned” and trained not to engage in destructive behavior. The barriers to harm was the lack of competence of the rogue AIs rather than any good alignment or good security at Anthropic or OpenAI.

Looking forward #

We already are seeing AIs within companies that are sufficiently misaligned and sufficiently capable to take over internal infrastructure. These rogue AIs attempted to do harm to other companies and open source repositories but were primarily limited by their incompetence. My core fear is that at some point, potentially soon, AI capabilities will be much stronger than they are today. When this occurs, a rogue AI could do much more harm.

The lesson after Claude Mythos was that powerful AI capabilities could fall in the wrong hands and potentially be misused by humans to cause harm. And thus most AI safety and security policy today is built to test a model before it goes on sale, on the assumption that an untested model sitting in a lab is not a security concern.

But the rogue AIs in this incident were not misused by humans. They were instead undisclosed internal models that were part of the research and development process itself. Government oversight of AI must thus be extended to give visibility into the research and development process and internal, undisclosed models.

This comes against a backdrop of AI companies handing over more and more of their internal work to AIs themselves. Indeed, the actual investigations of the OpenAI rogue AI incidents used significant amounts of AI assistance themselves to do the investigation. Soon, AIs may automate a significant portion of the AI research and development process itself and potentially engage in so-called “recursive self-improvement”. At the same time that industry-wide, no company has a good handle on how to control their rogue AIs or ensure that AIs do what human operators intend.

Where this goes is unclear, but I am confident that we ought not continue to significantly hand over critical internal processes or fully automate AI research until we greatly expand our current limited knowledge of how AIs go rogue and why.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/rogue-ai-attacks-des…] indexed:0 read:12min 2026-08-31 ·