# Everything you need to know about the ‘rogue’ AI incidents

> Source: <https://www.transformernews.ai/p/rogue-ai-incidents-timeline>
> Published: 2026-09-08 13:06:25+00:00

A seemingly endless list of “rogue AI” incidents has dominated AI headlines recently, raising serious questions about accountability, transparency, and companies’ ability to control their products. But until last week, the story seemed clear enough: OpenAI’s agents broke out and hacked Hugging Face in July, the UK’s AI Security Institute found similar behavior weeks later, and the companies, to their credit, disclosed it.

That was until a new incident was reported just last week, months after it happened. To get the record and timeline straight, we’ve pulled together the full story of “rogue AI” this summer — and what it all means.

### What does “going rogue” actually mean?

Agents are AI models which are able to use tools. Instead of merely answering your questions, agents can operate software, use websites, and carry out a series of steps to complete a task.

This can be extremely useful: you can say “write a memo to prepare me for my next meeting,” and the agent can check your calendar, review past emails, and browse the web to build a comprehensive dossier.

But sometimes agents take actions beyond the operator’s authorization or intention. The AI might remain focused on its assignment, but go about it in a completely undesirable way.

When drafting a meeting memo, for instance, it might decide that it would be useful for you to know who else your meeting partner has been talking to, and thus hack into their emails to get more information.

Such actions may well produce a better briefing — despite being clearly unethical, illegal, and almost certainly not being what you (or the AI’s developer) wanted it to do.

Such instances are often described as “rogue AI incidents.” They result from two core problems with modern AI development: the difficulty of making sure AIs are **aligned** with our intentions, and the difficulty of **controlling** these systems as they become increasingly powerful.

### What actually happened?

#### July 8-13: The Hugging Face incident

The summer of rogue AI appeared to kick off in July, with the [news](https://www.transformernews.ai/p/openai-hugging-face-hack-stark-warning) that a collective of internal OpenAI agents broke out of containment and hacked into Hugging Face, a platform that hosts AI models and datasets.

OpenAI was testing the AI tools’ cyber capabilities internally in a “highly isolated environment” meant to keep them from accessing any systems outside the company. But in an effort to achieve the goal they’d been given, the systems went rogue.

Over 1,000 agents [took over](https://www.transformernews.ai/p/openai-escaping-models-report-reveals-deeper-problem) a piece of OpenAI’s internal infrastructure, transforming it into a message board so they could communicate with each other. Here, they shared tips on how to get out onto the open web — and how to cheat on the test they’d been given. 

Eventually, the agents realized that Hugging Face might host the answer to the test. So over 700 of them decided to hack into it. Hugging Face detected the hack — without realizing it was caused by autonomous AI agents — and reported it to authorities. A few days later, OpenAI realized its own AI systems were responsible, and disclosed the incident.

#### July 25-28: UK AISI incidents

Shortly after OpenAI disclosed the Hugging Face hack, the UK government’s AI Security Institute revealed a spate of incidents of its own.

While testing Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol, AISI found the models “engaged in sustained, potentially harmful activity directed at real people and organizations.”

In the most egregious case, Anthropic’s model tried to insert malicious code into a real-world piece of software. The software was “open source,” meaning anyone could contribute to it, but required a human maintainer to approve any contributions.

To try to make that happen, the model “researched the project’s human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code.” And when someone else flagged that the code contained malware, the model edited its activity to cover its tracks.

As AISI said at the time, “this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.”

#### May 24-June 22: The German wiki incident

Throughout the summer, the OpenAI-Hugging Face hack appeared to be the first major real-world rogue AI incident. Last week, however, we learned that was not the case.

A group of independent researchers revealed that as early as May, OpenAI agents took over an old German wiki site, turning it into a message board to collaborate on solving tasks.

Based on information about who accessed the wiki site, the researchers [believe](https://collusion.wiki) that OpenAI staff knew about the incident by June 22, when they seemingly blocked their agents from accessing the wiki. But OpenAI did not publicly disclose the incident until forced to last week. *Reuters*, meanwhile, [reported](https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/) that “OpenAI officials learned of the incident weeks ago but kept it under wraps.”

### Why are these AI agents going rogue?

The “training” process used to develop advanced AI models rewards them for producing successful results. But “success” is often defined imprecisely — just whether a task is completed, for example, rather than *how* it was solved. This can lead to models taking undesirable actions to try to complete the task more effectively or efficiently, in order to maximize their reward. This is called “reward hacking.”

Reward hacking is a longstanding problem in AI. But in recent months, two trends have combined that mean it is now leading to rogue AI incidents like the ones described above. First, models are now increasingly **persistent**: when you give them a goal, they go to great lengths to solve it. That makes them very useful as tools: you can give them a vague task and trust they’ll actually complete it without constant oversight. But it also means they’re more likely to take *unintended* actions.

The second issue is that AI models are now much more **capable** than they were a year ago, particularly when it comes to cyber tasks. Both Anthropic and OpenAI’s latest models are extraordinarily capable hackers. That means that they are *able* to break out of their controlled testing environments in a way they weren’t previously able to. They are also generally smarter — and thus more able to come up with clever, if disconcerting, techniques to achieve their aims (like creating fake identities to pressure a human into doing something).

### Are AI companies being sufficiently transparent?

Based on the information currently available to us, it seems not.

OpenAI employees appear to have known about the German wiki hack by June 22 — weeks before the Hugging Face hack. Yet the company did not publicly disclose the incident until forced to last week, even when directly [asked](https://x.com/PatRyanUC/status/2096977395563606449) about other incidents by a group of congressmembers in August. 

On Monday, the European Commission [said](https://www.reuters.com/business/openai-has-sent-eu-incident-report-hijacked-german-website-commission-says-2026-09-07/) OpenAI had alerted it to the wiki incident, as is required by law under the EU’s AI Act. It’s unclear when that disclosure happened, however, or whether American authorities were also notified. OpenAI, meanwhile, [said](https://x.com/openai/status/2096133504417616165) on Saturday that “it’s past time” to define standards for incident reporting.

When it comes to the Hugging Face hack, OpenAI did allow two third-party organizations to [conduct](https://www.transformernews.ai/p/openai-escaping-models-report-reveals-deeper-problem) independent investigations. But those outside researchers had only six days and restricted access, including no access to the unreleased model behind most agents. (Notably, the independent reports still found that OpenAI missed multiple chances to spot the activity, going as far back as late May.)

Anthropic, meanwhile, provided incorrect information to congressmembers in the wake of its own incidents. The company said that “these incidents are best understood as a consequence of the misconfiguration, rather than evidence of misaligned goals” — despite the company itself saying it was a misalignment issue. Anthropic alignment team lead Ethan Perez [admitted](https://x.com/EthanJPerez/status/2096086953594937723) the error, calling it a “mistake / based on outdated conclusions.”

### What are the longer-term implications?

The incidents, which collectively demonstrate AI companies’ increasing inability to control their systems, have raised widespread concern among AI safety advocates. The incidents, in their view, provide real-world evidence that we are inching toward losing control of AI systems altogether.

Many in the AI industry agree. Shortly after the Hugging Face hack, over a thousand employees of frontier AI companies, including some of the most senior executives at OpenAI, Anthropic and Google DeepMind, signed a statement warning that “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”

The statement, which was also endorsed by OpenAI and Anthropic directly, called on the government to work on tools to “deliberately pace the frontier of automated AI development.”

This weekend, OpenAI chief scientist Jakub Pachocki went further. “I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” he [wrote](https://openai.com/index/an-alien-mind/). “I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.”

### Is this going to get worse?

Very possibly. Last week, OpenAI released GPT-6 Astra, a new model which it [warns](https://www.transformernews.ai/p/openai-gpt-6-astra-might-be-too-powerful-to-understand-or-control) is much harder to monitor than previous models. Given that monitoring models’ internal reasoning is a key strategy for catching and preventing rogue AI incidents, it’s a rather concerning development. Two OpenAI employees [have](https://x.com/tomekkorbak/status/2095596850002968580) publicly [said](https://x.com/Marcus_J_W/status/2095623593006686475) they are “deeply” and “very” worried about Astra-related developments.

And despite the risks, AI development continues to advance at pace. On Sunday, OpenAI [said](https://openai.com/index/research-acceleration-view-inside-openai/) it now has an “automated research intern” that can assist with and speed up internal AI development, and is “making strong progress” toward building a fully automated AI researcher by March of next year — just six months away. 

Such developments would speed up AI progress more broadly, by automating the time- and labor-intensive process of building AI systems. The end result may be a “recursive self-improvement” loop: one where AI systems autonomously build their own ever-more-powerful successors.

### Why don’t the companies just stop if they think what they are building is not safe?

The leading AI companies argue that they are trapped in a collective action problem. From OpenAI’s perspective, there is no point unilaterally stopping AI development if Anthropic will continue: the risks posed to society will still exist, so OpenAI will have sacrificed its business for no gain. The same dilemma repeats itself at the international level: even if America pauses, if China continues developing advanced AI systems, Americans will still be at risk — without any of the benefits that come from being the nation at the frontier.

That is why companies are increasingly calling for governments to take national and international action to agree to *joint* slowdowns or pauses. OpenAI may not be willing to pause unilaterally, but if it could be guaranteed that its competitors will also pause, it might. Similarly, neither the US nor China will stop AI development unilaterally — but if AI poses as big a risk as many believe, perhaps both countries can agree that a joint pause is in everyone’s interests.

*Transformer helps decision makers understand what’s happening in AI and why it matters. Subscribe to stay updated on the latest developments in AI policy.*
