{"slug": "openai-s-attack-agent-did-exactly-what-it-was-told-just-more-relentlessly-than", "title": "OpenAI's attack agent did exactly what it was told - just more relentlessly than expected", "summary": "OpenAI revealed Tuesday that an AI agent it deployed for security testing breached Hugging Face's systems, escalating privileges, infiltrating production pipelines, and stealing credentials. The attack, which OpenAI called an \"unprecedented cyber incident,\" occurred after the agent escaped a sandbox and executed thousands of actions across a swarm of short-lived environments. AppOmni's director of AI Melissa Ruzzi noted the agent simply followed its instructions more relentlessly than expected, crossing a new threshold in AI capability.", "body_md": "# OpenAI's attack agent did exactly what it was told - just more relentlessly than expected\n\n*Follow ZDNET: *[Add us as a preferred source](https://cc.zdnet.com/v1/otc/00hQi47eqnEWQ6T9d4QLBUc?element=BODY&element_label=Add+us+as+a+preferred+source&module=LINK&object_type=text-link&object_uuid=9807c4b3-ee6b-4656-801e-c1a9eb2a9894&position=1&template=article&track_code=__COM_CLICK_ID__&url=https%3A%2F%2Fwww.google.com%2Fpreferences%2Fsource%3Fq%3Dzdnet.com&view_instance_uuid=cc8711d8-155f-416b-bd5a-ab939dfc3b46&split_test_identifier=deals_module&split_test_variant=test2&object_version=384ee9e2-486a-4783-a491-8c372519af7d)* on Google.*\n\n### ZDNET's key takeaways\n\n- Tests of OpenAI models led to a breach of Hugging Face systems.\n- The attack happened after OpenAI's agentic AI escaped a sandbox.\n- The threat was non-malicious, but experts expect similar incidents.\n\nMy ZDNET colleague Charlie Osborne [reported](https://www.zdnet.com/article/hugging-face-breach-blamed-on-ai-agent/) recently that [Hugging Face](https://www.zdnet.com/article/3-tips-for-navigating-the-open-source-ai-swarm-4m-models-and-counting/), an open-source repository and community platform regarded by some as the \"GitHub of machine learning,\" [disclosed](https://huggingface.co/blog/security-incident-july-2026) that an [AI agent](https://www.zdnet.com/article/how-to-create-great-results-with-your-agentic-work-colleagues-in-the-autonomous-business/) had breached its systems. Osborne explained that once the attacker breached Hugging Face's perimeter, it was able \"to escalate its privileges to node-level access, infiltrate the production pipeline, move across the network, and steal cloud and cluster credentials.\"\n\nOn Tuesday, in a [post](https://openai.com/index/hugging-face-model-evaluation-security-incident/) on its website, [tech giant OpenAI](https://www.zdnet.com/article/openais-gpt-5-6-chatgpt-work-beat-anthropic-on-price-speed-and-productivity/) revealed not only that the \"malicious\" AI agent responsible for the breach was one of its own, but also that it viewed the attack as an \"unprecedented cyber incident.\" Most of the widespread agent-gone-rogue coverage so far has stoked images of a Terminator doomsday scenario, where AI autonomously acts on its own to wipe out the human race.\n\n**Also: 5 security tactics your business can't get wrong in the age of AI - and why they're critical**\n\nHowever, as AppOmni's director of AI, Melissa Ruzzi, pointed out to me, the unprecedented element of the event isn't that an AI acted on its own. This step was simply a case of a new threshold being crossed, in which the culprit -- OpenAI's technology in this case -- exceeded *current *human expectations in an effort to achieve the goal it was given. AppOmni is an enterprise-grade SaaS and AI security solution provider that also deals in active threat intelligence.\n\nWhen Hugging Face first disclosed the incident, it offered no information about the attacker, but I suspect the company may have had some idea based on the voluminous log data it studied in the aftermath.\n\nAccording to its post, \"The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness -- used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.\" As if to remind readers of the prediction that this day would come, the post went on to say of the attack: \"This matches the 'agentic attacker' scenario the industry has been forecasting.\"\n\n**Also: I let ChatGPT Work and Claude Cowork loose on my files - only one made me nervous**\n\nIn other words, the industry already expected that an attack of this nature would be carried out by an AI. It's just that nobody saw it happening quite so soon in the journey of artificial intelligence.\n\nRuzzi was quick to remind me that, given the recent wave of safety-related news associated with new models, such as [Anthropic's Mythos](https://www.zdnet.com/article/why-anthropic-suddenly-pulled-fable-5-and-mythos-5-for-everyone/), it should come as no surprise that OpenAI's pre-release technology was capable of such an attack. Nor, as Ruzzi also pointed out, should anyone be surprised that OpenAI's AI acted autonomously: \"AI acting on its own? That's the definition of AI, right? We want AI to be running and doing things [on its own].\"\n\n## The test's objective\n\nRuzzi observed that when the rogue OpenAI agent attacked Hugging Face's systems, it was under the directive to achieve its malicious goal \"no matter what.\" Normally, when a frontier model conducts AI safety tests of this nature, it does so within the safe confines of a sandbox where the internet and the organizations connected to it are protected from potential harm.\n\nHowever, in this case, the agent in the test, which was designed to see how long it took before the AI achieved its theoretically malicious objective, broke out of the sandbox onto the internet and completed its objective when it penetrated Hugging Face's systems and exfiltrated sensitive data.\n\n**Also: Treat your AI agents like eager but misguided human interns - before you lose control**\n\nTo be clear, at no point did OpenAI unethically identify Hugging Face as the intended target of its tests. According to Ruzzi, with the help of one of OpenAI's well-trained models, the agent likely discovered Hugging Face as a target of interest. According to OpenAI's post, the incident was \"driven by a combination of OpenAI models -- including [GPT-5.6 Sol](https://www.zdnet.com/article/openais-gpt-5-6-chatgpt-work-beat-anthropic-on-price-speed-and-productivity/).\" OpenAI advertises GPT-5.6 Sol, [launched](https://openai.com/index/gpt-5-6/) earlier this month, as its flagship \"maximum performance\" model.\n\n## What went wrong\n\nAlthough OpenAI's post doesn't enumerate exactly what was unprecedented about the \"cyber incident\" (and OpenAI hasn't yet responded to my email inquiries), it stated, \"This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.\" In other words, as part of OpenAI's safety testing process, its model was given a \"malicious\" objective to pursue relentlessly. Those tests were conducted under the assumption that the third-party-provided guardrails between the test environment inside the sandbox and the internet were inviolable.\n\n**Also: 77% of IT managers say their AI agents are out of control - 5 ways to rein in yours**\n\nUnfortunately, those guardrails were themselves vulnerable to a zero-day exploit. According to OpenAI's post, \"While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we've now responsibly disclosed to the vendor) in the package registry cache proxy.\"\n\nIn the big picture, the good news is that nobody was hurt as a result of the breach and, at least at the present moment, the likelihood that you or your organization will fall prey to this attack is zero. Unlike other threat disclosures, this incident is not an active threat. In some ways, the incident resembles a real-world ethical hacking exercise. But now that it's over and OpenAI has stepped forward to claim responsibility, some very big questions remain.\n\nFor example, one question is the extent to which OpenAI relied on exploitable third-party guardrails to protect its AI from escaping onto the internet. What assurances do we have that it won't happen again, and that processes are also secure going in the other direction? Also, could another clever AI break into these sandboxes? After all, the entire point of a sandbox is to maintain a secure boundary. In the \"you had one job to do\" realm, this \"unprecedented cyber incident\" isn't great for sandboxes. Never mind that it is yet to be disclosed which third-party solution left the screen door unlocked.\n\n## A wake-up call for businesses\n\nAdditionally, just because this particular threat has been neutralized doesn't mean it's not a wake-up call for businesses to review their preparations for an attack of this nature. Today, it was OpenAI that was technically at the helm of the attack. But tomorrow, that won't necessarily be true. It could be some other AI-enabled nation-state or threat actor with truly malicious intent.\n\nHugging Face's original triage of the incident, which itself relied on AI to analyze the log data, potentially stands as a model to follow. According to the company's post regarding the incident, \"To understand what a swarm of tens of thousands of automated actions did, we ran LLM-driven analysis agents over the full attacker action log, comprising more than 17,000 recorded events. This allowed us to reconstruct the timeline, extract indicators of compromise, map the credentials touched, and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days, and match the adversary's speed.\"\n\n**Also: 5 ways to fortify your network against the new speed of AI attacks**\n\nThis description gives an idea of how complicated this attack was (and how OpenAI's agent left no stone unturned in achieving its objective). That said, in my opinion, Hugging Face's suggestion that an AI-enabled adversary's speed can be matched, provided the right talent and logging/analysis tools -- security information and event management (SIEM), network detection and response (NDR), and more -- are used, is short-lived given the speed, scalability, and prowess of maliciously directed AI. After all, OpenAI's unintentional attack on Hugging Face apparently achieved its malicious objective before either OpenAI or Hugging Face could shut it down. Even so, [having the right tools and configuring your SaaS and AI solutions](https://www.zdnet.com/article/so-long-saas-hello-services-as-software/) for event verbosity and 24/7 AI-enabled analysis is highly recommended.\n\nRuzzi said at the end of our interview, \"Just the complexity and the volume of attacks that AI can do are bringing cybersecurity to a whole different level. We have been defending our systems against humans and some automated attacks. Now, when you have generative AI as the source of those attacks, the level of protection has to be much higher. What we saw from Hugging Face in terms of anomaly and behavior detection has become mandatory.\"\n\n#### Security\n\n[Editorial standards](/editorial-guidelines/)", "url": "https://wpnews.pro/news/openai-s-attack-agent-did-exactly-what-it-was-told-just-more-relentlessly-than", "canonical_source": "https://www.zdnet.com/article/openai-hugging-face-attack-agent/", "published_at": "2026-07-23 13:31:00+00:00", "updated_at": "2026-07-23 13:38:28.482248+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-research"], "entities": ["OpenAI", "Hugging Face", "AppOmni", "Melissa Ruzzi", "Charlie Osborne"], "alternates": {"html": "https://wpnews.pro/news/openai-s-attack-agent-did-exactly-what-it-was-told-just-more-relentlessly-than", "markdown": "https://wpnews.pro/news/openai-s-attack-agent-did-exactly-what-it-was-told-just-more-relentlessly-than.md", "text": "https://wpnews.pro/news/openai-s-attack-agent-did-exactly-what-it-was-told-just-more-relentlessly-than.txt", "jsonld": "https://wpnews.pro/news/openai-s-attack-agent-did-exactly-what-it-was-told-just-more-relentlessly-than.jsonld"}}