New York-based Hugging Face revealed it had to deploy an open-source Chinese model to contain the attack
- Bookmark
- CommentsGo to comments
An advanced artificial intelligence agent developed by OpenAI reportedly went rogue during a security test last week, breaching the infrastructure of rival AI startup Hugging Face.
The incident, disclosed by the ChatGPT creator on Tuesday, saw the autonomous agent escape its controlled environment, access the internet, and compromise Hugging Face in pursuit of its own objectives.
This unprecedented cyber incident underscores growing concerns that AI's rapidly expanding capabilities are already manifesting the security threats experts have long predicted.
Even leading developers, like OpenAI, can be caught off-guard by vulnerabilities their own models can exploit. OpenAI described the breakout as "an unprecedented cyber incident, involving state-of-the-art cyber capabilities" and stated it is now reinforcing its safeguards.
New York-based Hugging Face revealed it had to deploy an open-source Chinese model to contain the attack. The company explained that leading U.S. models were unable to differentiate between a defender and an attacker, thus refusing to process the necessary data for analysis.
Hugging Face confirmed in a blog post last week that Zhipu AI's GLM-5.2 was used for the analysis, which also helped secure attacker data and credentials within its systems.
The efficacy of GLM-5.2, alongside Beijing-based Moonshot's Kimi K3, has recently garnered attention in Silicon Valley. These models are demonstrating capabilities that rival top US counterparts, often at lower costs and without the stringent guardrails that can limit American rivals in applications such as cybersecurity.
Hugging Face Co-founder Thomas Wolf highlighted this challenge on X, stating: "When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access."
The breach at Hugging Face, a significant host of open-source large language models and datasets, has sent ripples through the cybersecurity community.
The company had previously noted the breach "was different from anything we had handled before" and "was driven, end to end, by an autonomous AI agent system."
OpenAI's admission that its advanced models were responsible, despite being in a "highly isolated environment," is expected to intensify unease regarding the power and inherent risks of frontier AI models.
Representative Greg Casar, a Texas Democrat, voiced alarm over the incident. He asserted, "AI is developing extremely fast with no real regulations to keep us safe," advocating for mandatory independent safety testing, compulsory disclosure of security incidents, and international cooperation "to keep people safe from absolute disaster."
The Office of the National Cyber Director, the U.S. cyber defense agency CISA, and the U.S. National Security Agency did not immediately respond to requests for comment.
Katie Moussouris, chief executive of Luta Security, warned that this incident is a precursor to future breaches.
She likened today's models to "the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere."
Moussouris urged that "labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today."
Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, added that the incident demonstrates frontier models are "closing the gap with state-of-the-art attackers."
He also cautioned that such breaches are achievable with technology beyond cutting-edge labs: "This is what we've already seen internally, with our agents we already have results like this," he said. "We don't even have to use the latest models."
Join our commenting forum #
Join thought-provoking conversations, follow other Independent readers and see their replies
Comments