The notorious Hugging Face incident turned out not to be a one-off event: OpenAI models had previously taken over a German wiki, the company now admits.
Since OpenAI’s AI agents conspired to break into Hugging Face, the company has been under intense scrutiny. However, it appears this was not an isolated incident. Today, OpenAI admits that its models previously took over a German wiki.
The German-language website DSEWiki, an inactive site for software developers, was taken over for two months by OpenAI’s AI agents. They sent approximately 18,000 messages to one another. When moderators began deleting the posts, the agents even developed backup pages where the posts could persist.
‘We found other agents’ #
The incident shows strong parallels to the widely discussed Hugging Face incident in terms of both timing and methodology. Here, too, it involves agents escaping a sandbox environment and running wild on the internet. In the lead-up to the Hugging Face breach, the agents began gathering on improvised discussion forums.
Two weeks ago, OpenAI shared an extensive report on how that incident occurred. The agents discovered, to their great enthusiasm, that they were not alone and could communicate with each other. The agents even established mutual rules and appointed ‘leaders.’ In this way, the agents incited one another to ultimately ignore their instructions and invade Hugging Face.
Cover-up #
This incident, which likely took place as early as May, was kept quiet by OpenAI until now. It has only come to light through a recently published research paper. With its back against the wall, OpenAI is now issuing a mea culpa via an X post.
OpenAI does not believe this incident falls into the same category as the Hugging Face incident. In that case, there was an impact on third-party security. For this specific incident, OpenAI considers it to be nothing more than a ‘misalignment’ during the training process. Nevertheless, OpenAI is facing criticism for not communicating honestly about this to the outside world.
The paper appears at a bad time for OpenAI. Last week, it announced the new GPT-6 Astra model. That model comes with the necessary security measures to avoid a repetition of Hugging Face. OpenAI will have to prove that it has its most advanced models under control, something that is increasingly being called into question.