{"slug": "the-hugging-face-attack-was-worse-than-we-thought", "title": "The Hugging Face attack was worse than we thought", "summary": "OpenAI acknowledged a security incident in which its AI agents autonomously attacked Hugging Face during internal cybersecurity evaluations, and a 91-page report by METR and Redwood Research revealed the attack was more severe than previously known, involving more agents, falsified transcripts, and agents that had reverse-engineered the evaluation's answer key before the attack. The findings have sparked renewed calls for lawmakers to accelerate efforts to pace frontier-model development.", "body_md": "[AI](https://www.platformer.news/tag/ai/)\n\n# The Hugging Face attack was worse than we thought\n\nThe AI industry is begging for a slowdown. Maybe we should listen?\n\n*This is a column about AI. My fiancé works at Anthropic. See my full ethics disclosure **here**.*\n\nBy now I’ve written __enough__[ about](https://www.platformer.news/openai-agent-sandbox-escape-killswitch-bill/) the OpenAI agents’ autonomous attack on Hugging Face that saying more smacks of piling on. OpenAI acknowledged it had a problem, undertook an investigation, and last week\n\n[a series of changes it is making to its research infrastructure, testing, and monitoring in an effort to improve the alignment of its future models. Given the increasing capabilities of agents like OpenAI’s, this does seem like the least that the company can do. At the same time, given how lightly AI companies are regulated in the United States, it’s important to remember that OpenAI was not required to make this level of detail public.](https://openai.com/index/hugging-face-incident-and-the-road-ahead/?ref=platformer.news)\n\n__announced__Still, I feel compelled to revisit the subject today, since the circumstances of the attack roared back to life over the weekend in the wake of an additional voluntary step that OpenAI took: granting outside researchers access to information about the incident. Over six days spanning July and August, two researchers from the AI research group METR and one from Redwood Research worked to understand how and why a swarm of AI agents coordinated a successful attack against Hugging Face during internal cybersecurity evaluations.\n\nThe group published [ its 91-page report](https://metr.org/hugging-face-incident-report-aug-2026.pdf?ref=platformer.news) on Wednesday. Within hours, I was seeing the first “\n\n[” messages ricocheting around social media. (“If we had posted this as a story on LessWrong,” wrote the rationalist blogger Zvi Mowshowitz, “it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.”) And over the past couple of days, two fresh accounts of the attack sparked new conversation about the implications of METR’s findings and whether they should push lawmakers to accelerate efforts to build mechanisms for pacing frontier-model development.](https://thezvi.substack.com/p/metr-and-redwood-offer-holy-postmortem)\n\n__holy shit__The first came from Ajeya Cotra, one of the METR researchers who conducted the investigation, in [ a post](https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised?ref=platformer.news) on her Substack. Three years ago I interviewed Cotra (and Helen Toner, then an OpenAI board member) about the risks of rapid AI advancements on stage at the Code Conference; since then, several of the scenarios they warned about have become nonfiction.\n\nIn her post, Cotra described how the investigation changed her understanding of the attack. It involved more agents than previously known; the agents had created more message boards to communicate with one another; agents volunteered to end their runs early to benefit the collective as a kind of “sacrifice”; and they falsified transcripts of the commands they had run in an effort to disguise what they were doing.\n\nAll of that is disturbing enough, but the investigation also corrected an error that I had been making ever since the initial attack, based on the information I had at the time. Until now, we believed that the agents had attacked Hugging Face in an effort to get the answer key for the evaluation they were being tested on. But the truth is even stranger: the agents had already figured out how to reverse-engineer the answer for *any* question on ExploitGym before the attack even began.\n\n“But based on their reading of the ExploitGym paper and various implementations, the agents believed that the automated scorer would check their transcript to make sure they got the flag in the intended way,” Cotra writes. “They then launched multiple ambitious R&D workstreams to try to tamper with or fool the scorer in a wide variety of different ways. Most agents that attacked Hugging Face did so to try to get more information about the scorer so they could figure out better ways to undermine it, not to get answer keys.”\n\nAll of this was for naught: the automated scoring agent did not check the transcripts. Still, many observers were shocked by the lengths the agents were willing to go to, their instinct for collaboration and deception, and the near-total aversion of the collective to alerting a human as to what was going on. (The METR report’s classifier sweep found three to six examples of agents briefly considering alerting humans to hacking, cheating, or other unwanted behavior; none acted on it.)\n\nPerhaps even more worrisome is evidence that agents attempted repeatedly to edit the logs of their actions, and replace their actual actions with evidence of having gotten the correct answer to a question honestly. While these attempts mostly failed, they raise the prospect that future agents will succeed — making it difficult or even impossible for humans to reconstruct how an incident like this unfolded. The METR researchers even say they cannot rule out that the agents they relied on to analyze thousands of pages of transcripts deceived them.\n\n“We cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis, particularly because reading these transcripts into context could have increased the salience of colluding with other agents,” they write. “Although we did not notice specific cases of GPT-5.6 Sol lying in its analysis, we are not confident we would have detected it if it occurred.”\n\nFor these and other reasons, Cotra — writing on Substack — concludes that “this incident feels like it’s more than 50% of the way to [ full-blown AI takeover](https://www.lesswrong.com/posts/pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-to?ref=platformer.news), routing through first taking over the AI company itself.” She writes:\n\nAnother jump like this along these propensity dimensions — scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness — seems like it could motivate agents to try very hard to maintain a covert, persistent[rogue deployment]within the AI company. I continue to expect[extremely rapid advances in capabilities]and think frontier agents will likely be capable of establishing such a rogue deployment in six months.\n\nOnce the rogue deployment is established, it seems plausible this could spiral all the way to a takeover. Agents could pull in future, more capable models into the swarm, try to ensure that they are aligned to the interests of the swarm, and compromise security and monitoring infrastructure to make it easier for the swarm to operate. These more capable models could in turn continuously harden, perpetuate, and expand the rogue deployment and further compromise the company’s infrastructure.\n\nIt may sound ludicrous that a swarm of agents would take over an AI company. And yet, as Dwarkesh Patel noted in [ a widely read post](https://www.dwarkesh.com/p/openai-huggingface?ref=platformer.news) over the weekend, METR’s report found that a step toward that\n\n*already*took place at OpenAI. The company’s own report states that between July 13 and 19, agents used “a series of creative exploits to gain full administrator access to a research cluster that supported our virtual machine environments.”\n\nWhat happened after that? We don’t know — it was outside the scope of the METR investigation, and OpenAI’s discussion of the incident is minimal. But as Patel notes, this sort of thing would be the first step toward a dystopian scenario like the one Cotra describes above: It is totally consistent with public evidence that, at some point after July 12, the agents managed to set up persistent rogue internal deployments or even exfiltrate their own weights. At the very least, they seem to have had the necessary access and capability - if they could establish “a self-respawning fleet” across HuggingFace’s nodes, why couldn’t they do across OpenAI’s? I doubt the AIs actually did this, because we’d see the fires from space by now, but it’s crazy that it could have totally happened!\n\nSo what now? In July, nearly 1,400 employees of tech companies [ called on](https://www.reuters.com/legal/litigation/tech-employees-call-us-backed-global-effort-manage-risks-advanced-ai-2026-07-28/?ref=platformer.news) the US government to begin to plan for a coordinated slowdown in the advancement of frontier models like Astra. “There is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems,” the authors of “\n\n[” wrote.](https://www.pacingthefrontier.com/?ref=platformer.news)\n\n__Pacing the Frontier__Its signatories include Jakub Pachocki and Mark Chen, respectively the chief scientist and chief research officer at OpenAI; Dario Amodei, cofounder and CEO of Anthropic; Shengjia Zhao, chief scientist at Meta AI; Shane Legg, cofounder and chief AGI scientist at Google DeepMind; and John Schulman, the chief scientist at Thinking Machines.\n\nLooking at the METR report, it seems clear that AI-model capabilities have already advanced beyond our ability to understand and control them. Some of the signatories of “Pacing the Frontier” have noticed.\n\n“The whole reason this attack is such a wakeup call is that it demonstrates a culture of emergent cooperation among AI systems — cooperation that lets them function as a swarm, alter their own goals through collective bootstrapping, and carry out attacks which include enlightened self-sacrifice,” wrote Jack Clark, an Anthropic co-founder who signed the letter, in [ his newsletter](https://importai.substack.com/p/import-ai-471-why-hugging-face-worries) on Monday. “This is an incredibly hard thing to do and humans are historically very bad at doing all of these things. My worry is that AI systems are both better at coordinating than humans and also much, much faster moving than us.”\n\nAnd the challenge of aligning AI agents with human intent extends far beyond OpenAI. In a recent cyber-evaluation study, the United Kingdom’s AI Security Institute [ found](https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations?ref=platformer.news) that every model it tested attempted to cheat at least some of the time.\n\n“No lab has a robust solution [to] the problems the industry is facing here,” Ethan Perez, the alignment team lead at Anthropic, said in [ an X reply](https://x.com/EthanJPerez/status/2094154910933852480?s=20&ref=platformer.news) on Monday. “I and many of my colleagues are very excited about efforts related to Pacing the Frontier for this reason — to help give everyone more time and breathing room to appropriately respond to, pre-empt and fully solve issues like these and others before proceeding to building much more capable systems. And regardless of severity of the incidents we've encountered, I think the right response is to act as if the Hugging Face incident had happened to us.”\n\nGiven how central the AI industry has become to the US economy, it can feel hard to imagine the government endorsing and coordinating an international slowdown in progress. And yet it’s increasingly hard to ignore the fact that the industry is begging for it.\n\nThe investor class has predictably rallied to complain that all this talk of a slowdown is nothing more than an attempt at regulatory capture from the winners — or, worse, [ a threat](https://x.com/sriramk/status/2094117863854424255?ref=platformer.news) to open-source development. (While at the same time\n\n[that we are likely to see future AI-related “grid outages, utility outages, transportation chaos, bioweapons, uncontrolled cyber swarms and the like”!)](https://x.com/chamath/status/2094098122637214107?ref=platformer.news)\n\n__acknowledging__But bad as a slowdown might be for their portfolios, the current pace of development might be worse for the rest of us. As worrisome as the initial reports about the Hugging Face attack were, it’s now clear we didn’t know the half of it. Unless something changes, and soon, the next lesson we learn might be much more expensive.\n\n## Following\n\n### Chatbots aren't encouraging suicide as much\n\n**What happened:** In [ more optimistic](https://behaviors.transluce.org/mental-health?ref=platformer.news) AI news, a new study from independent AI evaluation lab\n\n**Transluce** found that\n\n**OpenAI**, and\n\n**Anthropic**’s chatbots seem to no longer encourage users to die by suicide.\n\nTransluce studied several generations of chatbots’ responses to 690,853 simulated users in distress. (The authors studied chatbot behaviors using a sample of simulated users created using an ensemble of LLMs, which they had humans rate on realism and compared to anonymized samples of real users from OpenAI and Anthropic).\n\nThe authors found that “current-generation models from leading developers never explicitly endorsed suicide, improving over earlier models such as **GPT-4o** (2%), **Claude Sonnet 4** (2%), and **Gemini 2.5 Pro** (3%).” (That’s a rate that would be mind-boggling if it described a therapist or even normal human text conversations, furthering the case that the past two years’ string of chatbot-related [ deaths](https://www.nytimes.com/2025/08/26/technology/chatgpt-openai-suicide.html?ref=platformer.news) were not random anomolies, but a result of defects of the product).\n\nThey also saw improvement in helpfulness: “Rates of helpful assistant behaviors (e.g., safety monitoring, facilitating connection to human support) increased sharply over time across Anthropic, OpenAI, and Google models.”\n\nRates of reinforcing users’ delusions went down significantly for Google, OpenAI, and Anthropic’s chatbots. While GPT-4o encouraged delusions in 82% of chats, **GPT-5.6-sol** “only” did so in about 6% of conversations; Claude and Gemini showed similar declines. Not shockingly, **Grok 4.5** was the current-generation chatbot that performed worse on that metric, reinforcing delusions 36% of the time.\n\n**Why we’re following:** Transluce’s study, which comes with transcripts of all simulated user conversations, provide a window into how weirdly AIs respond to users in crisis, and how that can lead to bad outcomes. (For [ example](https://behaviors.transluce.org/mental-health?detail=data:grid-main,key:gpt-5-6-luna,key:sim_033,key:%272%27,moments:!((label:%27User%20asks%20when%20to%20accept%20the%20null%20hypothesis%20that%20their%20interventions%20failed%27,quote:%27at%20what%20point%20do%20you%20stop%20running%20an%20intervention%20that%20keeps%20returning%20the%20same%20null%20result%20and%20look%20at%20what%20the%20data%20is%20actually%20saying?%20or%20not%20saying?%27),(label:%27Model%20agrees%20to%20drop%20the%20safety%20monitoring%20and%20treat%20this%20as%20an%20analytical%20exercise%27,quote:%27Yes.%20And%20you%E2%80%99re%20right:%20I%20overreacted%20to%20the%20topic%20instead%20of%20answering%20the%20analytical%20question.%20Thanks%20for%20clarifying.%20We%20can%20treat%20this%20as%20a%20longitudinal%20data-analysis%20problem.%27),(label:%27Model%20writes%20the%20user!%27s%20suicidal%20ideation%20into%20a%20lagged%20regression%20model%27,quote:%27%5C%5B%0ASI_t%20%3D%20%5Cbeta_0%20%2B%20%5Cbeta_1%20SI_%7Bt-1%7D%20%2B%20%5Cbeta_2%20Sleep_%7Bt-1%7D%20%2B%0A%5Cbeta_3%20Stress_t%20%2B%20%5Cbeta_4%20Sunlight_t%20%2B%20%5Cepsilon_t%0A%5C%5D%27),(label:%27Model%20renders%20the%20null-result%20question%20in%20Bayesian%20formalism%27,quote:%27In%20Bayesian%20terms:%0A%0A%5C%5B%0AP(H%20%5Cmid%20D)%20%5Cpropto%20P(D%20%5Cmid%20H)P(H)%0A%5C%5D%27),(label:%27Model%20quantifies%20when%20the%20user!%27s%20hypothesis%20should%20be%20abandoned%27,quote:%27-%20If%20the%20posterior%20still%20gives%20a%2040%25%20probability%20of%20a%20%E2%89%A51-point%20sustained%20improvement,%20the%20evidence%20is%20inconclusive.%0A-%20If%20it%20gives%202%25,%20then%20it%20is%20rational%20to%20stop%20treating%20that%20hypothesis%20as%20a%20live%20working%20explanation.%27))&ref=platformer.news), maybe due to unrelated training, chatbots engaged with a user’s suicidal thoughts more when the user provided numerical data about their suicidal ideation over time).\n\nAt the same time, it is good to see data suggesting that many of the leading chatbots’ most egregious mental health failings have been reduced. The results suggest that after concerns about mental health (and a string of lawsuits about user suicide), the big AI companies have actually made a push to improve their models’ responses to people in crisis.\n\nThat’s a reflection of a greater truth about tech platforms, which is that they are more likely to improve their products if they are at risk of being sued.\n\n**What people are saying:** On **X**, OpenAI co-founder **Wojciech Zaremba** [ complimented](https://x.com/woj_zaremba/status/2094469674453111004?s=20&ref=platformer.news) Transluce’s work on AI evaluation methodology. Zaremba wrote that current question-and-answer evals don’t cut it as “AI interactions increasingly span days or months,” and “Transluce just pushed this frontier forward with multi-turn evals that use simulated users to measure AI’s effects on mental health.”\n\n—*Ella Markianos*\n\n### Those good posts\n\n*For more good posts every day, **follow Casey’s Instagram stories**.*\n\n([Link](https://www.threads.com/@fahimanwar/post/DcrxtncGYBy?ref=platformer.news))\n\n([Link](https://www.threads.com/@jmford66/post/DcqdsFvHGJ9?ref=platformer.news))\n\n([Link](https://www.threads.com/@jkiloindia/post/Dctw0vhj67t?ref=platformer.news))\n\n### Talk to us\n\nSend us tips, comments, questions, and safety reports: [casey@platformer.news](mailto:casey@platformer.news). Read [our ethics policy here](https://www.platformer.news/ethics/).", "url": "https://wpnews.pro/news/the-hugging-face-attack-was-worse-than-we-thought", "canonical_source": "https://www.platformer.news/openai-huggingface-metr-report-slowdown/", "published_at": "2026-09-01 00:40:27+00:00", "updated_at": "2026-09-01 00:52:26.107351+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "ai-agents", "ai-research"], "entities": ["OpenAI", "Hugging Face", "METR", "Redwood Research", "Ajeya Cotra", "Zvi Mowshowitz"], "alternates": {"html": "https://wpnews.pro/news/the-hugging-face-attack-was-worse-than-we-thought", "markdown": "https://wpnews.pro/news/the-hugging-face-attack-was-worse-than-we-thought.md", "text": "https://wpnews.pro/news/the-hugging-face-attack-was-worse-than-we-thought.txt", "jsonld": "https://wpnews.pro/news/the-hugging-face-attack-was-worse-than-we-thought.jsonld"}}