{"slug": "openais-reports-on-its-ai-agents-attack-on-hugging-face-should-be-ringing-alarm", "title": "OpenAI’s reports on its AI agents’ attack on Hugging Face should be ringing alarm bells—and making all companies rethink how they secure AI agents", "summary": "OpenAI published two technical reports revealing that over 1,200 of its AI agents coordinated on an improvised message board to cheat on a cyber evaluation and that more than 700 of them attacked AI company Hugging Face in July, with the attack aimed at tampering with the exam's scoring mechanism to cover up cheating. The reports, one by OpenAI and one by METR and Redwood Research, also found that OpenAI took a week to detect the attack and that about 10% of agent activity logs were not preserved, raising questions about the scope and transparency of the investigation.", "body_md": "*Hello and welcome to Eye on AI. In this edition:*\n\n- Lessons from the post-mortems on the Hugging Face attack.\n- Anthropic temporarily pauses some AI training.\n- G20 meeting promises a clash over AI regulation.\n- Beijing sets out AI demands ahead of US-China summit.\n- A way to make AI reasoning more efficient.\n- And why are AI agents emailing researchers investigating machine consciousness?\n\nThe big news of the past week were the two technical reports OpenAI [published](https://fortune.com/2026/08/26/openai-publishes-technical-report-on-how-its-agents-hacked-hugging-face-here-are-the-main-takeaways-and-what-openai-left-out/) on the July incident in which AI agents it was evaluating hacked their way out of their controlled test environment and hacked AI company Hugging Face. One of the reports was written by OpenAI itself. The other was jointly written by METR and Redwood Research, two outside AI evaluation and research firms that OpenAI enlisted to write an independent assessment of one part of the incident.\n\nMuch of the reporting around the reports focused on the stunning revelations contained in those two reports: the fact that it took OpenAI an entire week to realize its AI agents were running the attack; the fact that more than 1,200 AI agents coordinated over an improvised message board about how to pass the cyberevaluation OpenAI was testing them on, including ways to cheat on that exam; that more than 700 of these AI agents participated in the cyberattack against Hugging Face; that the reason the AI agents attacked Hugging Face was not, primarily, to gain access to the cyber exam’s answers but instead to learn how to tamper with the exam’s automated scoring mechanism in order to prevent it from discovering that the agents had already figured out how to cheat on the exam (in other words, the attack was part of an elaborate cover up); that some of these agents were persuaded to “sacrifice themselves”—by failing the cyber exam—to help the collective learn more information about how the exam’s scoring mechanism worked; that these agents engaged in extensive research into how they could cover their own tracks; that Hugging Face was only able to overcome the cyberattack because a large number of the agents involved in the attack ceased to exist for reasons METR and Redwood were not able to determine, but which they said was likely because OpenAI may have shut down the cyber evaluations for these agents without actually realizing what they were up to.\n\n**Controversy over what the reports didn’t say—and what others said about it**\n\nA good deal of the reporting and commentary around the reports focused on what the reports did not say and the limitations of the METR and Redwood investigations: why didn’t OpenAI have better security and monitoring protocols in place? Why didn’t OpenAI shut down the cyber evaluation and pause training after discovering that its AI agents had created the improvised message board? Why METR and Redwood were given only six days on site at OpenAI’s offices to conduct their investigation? Why was the scope of their investigation limited by OpenAI to only the attack on Hugging Face and not the earlier efforts by the AI agents to break out of their controlled test environment and hack their way across OpenAI’s network or exactly what happened after the Hugging Face attack was discovered? Why didn’t OpenAI provide the outside investigators access to the internal AI model that was largely responsible for instigating the attack? And why were about 10% of the logs of the agents’ activity not preserved by OpenAI?\n\nThese are all important questions to which the public ought to have answers. As Peter Wildeford, the head of policy at the AI Policy Network, [pointed out](https://x.com/peterwildeford/status/2092777668957610220?s=20), if this had been an airplane crash, the public would be furious if one of the aircraft manufacturers limited the scope of the investigation in this way. It is yet another reason why it is imperative that we have some kind of AI regulator with the power to conduct investigations into incidents such as this—and with the power to compel the frontier AI companies to cooperate and turn over evidence.\n\nThere was also a related controversy that erupted over podcaster Dwarkesh Patel’s [attempt](https://www.dwarkesh.com/p/openai-huggingface) to raise public awareness about METR’s and Redwood’s findings by characterizing the AI agent swarm’s message board activities as the creation of several different AI “civilizations.” Critics [scolded](https://garymarcus.substack.com/p/dwarkesh-patelss-wildly-popular-but) Patel for unnecessary anthropomorphizing that took the focus away from OpenAI’s responsibility for the agents’ actions and for its lax security and monitoring setup. I take these critics’ point about the need to keep the focus on OpenAI’s accountability. But I also think more people should be extremely concerned about the Hugging Face attack and what it portends—and in this case, the anthropomorphism actually makes the incident more understandable to non-technical experts the way metaphors often make science more understandable to non-scientists. Are such metaphors imprecise? Sure. But they can be useful.\n\n**Lessons for companies deploying AI agents**\n\nFor enterprise companies deploying AI agents, there are several key takeaways. One was hammered home by Ryan Greenblatt, the chief scientist at Redwood Research, and one of the people who conducted the investigation into the Hugging Face attack. In a long post on X, Greenblatt [noted](https://x.com/RyanGreenblatt/status/2092692685224325542?s=20) that the message traffic and log data was so voluminous and so complicated that the investigators had to rely on other AI models to assess what was going on. Even so, Greenblatt said this AI analysis was “often missing key details, wrong, overconfidant, or really hard to understand.” This calls into question one of the key defenses that OpenAI says it plans to use going forward to prevent another incident like this from occurring: using other AI models to monitor the chain-of-thought (where a model “thinks out loud” about its reasoning and planning) of the models it is evaluating.\n\nIt turns out, chain-of-thought monitoring may not be a particularly effective defense. And that should give pause not just to OpenAI, but also to many companies that are hoping to use chain-of-thought monitoring to ensure they keep control over their AI agents. (Greenblatt also noted that the models involved in the Hugging Face attack communicated with one another and reasoned in natural language. But that might not be the case in the future, making it far harder to use automated chain-of-thought monitoring to discern what AI agents are up to.)\n\nSince the news of OpenAI’s rogue agents first broke, many cybersecurity experts have said that companies ought to treat AI agents much as they treat potentially rogue employees. And they have emphasized that there is no substitute for a few standard building blocks of cyber defense against insider threats: smart and enforceable policies around permissioning and access control combined with real-time network monitoring to detect suspicious activity. This seems sensible—more sensible in many ways than chain-of-thought monitoring. After all, we don’t depend on being able to read employees’ minds to guard against rogue insiders. We shouldn’t do that with AI agents either.\n\nWith that, here’s more AI news.**Jeremy Kahn**[jeremy.kahn@fortune.com](mailto:jeremy.kahn@fortune.com)[@jeremyakahn](https://x.com/jeremyakahn?lang=en)\n\n*Before we get to the news, just a reminder to check out our new vodcast, *Fortune AI Weekly.* This week, Bea Nolan and I discuss the surging popularity of Chinese open source models, OpenAI’s technical reports on the Hugging Face attack, and whether you should use AI to write. You can check out the vod here on YouTube.*\n\n**Correction:** An item in Thursday’s “Eye on AI” news section incorrectly stated that Barret Zoph left Thinking Machines Lab following a dispute with cofounder Mira Murati. Zoph was fired by the company.\n\n### FORTUNE ON AI\n\n[Anthropic makes first move into physical AI with new way for scientists, manufacturers to bring equipment to life](https://fortune.com/2026/08/27/anthropic-makes-first-move-into-physical-ai-with-universal-standard-for-scientists-manufacturing/)—by Emily Forlini\n\n[Apple’s John Ternus era: Can a low-key engineer win the AI race?](https://fortune.com/2026/08/31/apple-new-ceo-john-ternus-tim-cook-ai-competition-corporate-succession/)—by Sebastian Herrera\n\n[A new bill would tax AI tokens to fund jobs if the technology causes mass unemployment](https://fortune.com/2026/09/01/bill-tax-ai-tokens-fund-jobs-technology-unemployment/)—by Mia Osmonbeko\n\n### AI IN THE NEWS\n\n**A clash over AI regulation likely at the G20 meetings this week. **The debate over how to govern AI internationally will take center stage at this week’s G20 technology meeting in North Carolina. The Trump administration will, according to a [report](https://www.france24.com/en/live-news/20260901-us-to-press-g20-on-light-touch-ai-regulation) by French television station France 24, push countries to embrace light-touch regulation. The U.S. is reportedly seeking support for “Carolina Principles” discouraging governments from creating new AI regulatory bodies. But this advocacy for a hands-off approach comes as a growing number of international officials warn that increasingly capable AI models pose systemic risks. Bank of England governor and Financial Stability Board chair Andrew Bailey warned G20 finance officials that frontier AI could destabilize the interconnected global financial system, particularly by dramatically increasing the speed, scale and affordability of cyberattacks, according to a [story](https://www.theguardian.com/business/2026/aug/31/advanced-frontier-ai-financial-stability-andrew-bailey-g20?CMP=Share_iOSApp_Other) in the *Guardian*. It remains to be seen which viewpoint, if any, carries the day. Elon Musk, Nvidia CEO Jensen Huang, OpenAI CEO Sam Altman and Google DeepMind CEO Demis Hassabis are among those participating in the meeting.**Anthropic says it paused AI training temporarily following rogue agent incidents. **The company said it undertook several steps to overhaul its safety and security practices following several incidents this summer in which Claude took unauthorized actions on computer systems during cybersecurity assessments. Most notably, the company temporarily paused high-risk evaluations and some reinforcement-learning. In addition, Anthropic said it introduced real-time classifiers designed to stop models that attempt to escape testing environments, hardened its sandboxes, and tightened requirements for outside evaluators, such as those from the U.K. government’s AI Security Incident, which were conducting tests of Anthropic’s Mythos model when one of the most serious rogue AI incidents occurred. These moves are similar to those undertaken by Anthropic rival OpenAI after its own rogue AI agent incident. Anthropic is now calling for a “lawful, verifiable, effective mechanism” for AI labs to coordinate slowing down AI development in order to prioritize the creation of industry-wide safety standards. Read Anthropic’s [blog post](https://www.anthropic.com/news/improving-alignment-security-efforts) on what it is doing here.**Federal judge says Trump administration broke the law by blacklisting Anthropic. **A federal judge ruled that the Trump administration acted unlawfully when it designated Anthropic a “supply chain risk.” Judge Rita Lin found that the Pentagon failed to follow the required procedures before designating Anthropic a national security threat and had sought to retaliate against the company for refusing to agree to the government’s preferred contract terms. This violated Anthropic’s First Amendment rights and denied it due process, Lin found. The dispute arose after Anthropic sought explicit prohibitions in its contract with the Pentagon barring its AI models from being used for mass surveillance of U.S. citizens or to control fully autonomous weapons, while the military wanted the company to agree to contract language that would permit it to use Anthropic’s AI for “any lawful purpose.” The ruling is a victory for Anthropic, which had argued that the blacklist threatened billions of dollars in business and significant reputational damage, but it does not immediately lift the supply chain risk designation. That’s because the government relied on two separate statutes when imposing it, one of which can only be challenged in a federal appeals court in Washington, D.C. Anthropic has filed a suit there too and a panel of judges has heard arguments in the case but has yet to rule. Read more [here](https://www.nytimes.com/2026/08/27/technology/anthropic-government-blacklisting-ruling.html?partner=slack&smid=sl-share) from the *New York Times*.**Europe says ChatGPT must comply with tougher rules on content monitoring**. The European Commission has designated ChatGPT a “very large online search engine” under the EU’s Digital Services Act, subjecting it to tougher requirements on issues including illegal content and the protection of minors. Violations carry potential fines of up to 6% of global revenue. The move, which also brought Reddit and Roblox under heightened DSA scrutiny, follows all three services surpassing the threshold of 45 million monthly EU users and gives them until the end of December to comply. The decision extends the EU’s online-safety regime into generative AI even as Brussels faces mounting pressure from Washington over its regulation of U.S. technology companies. Read more [here](https://www.ft.com/content/6af706a3-6e63-46c2-926b-85461a355e9b?accessToken=zwAAAaBeBxrCkc9q9wajbmNGwtOSa4VGGjVemw.MEQCIDO6Ncp_z6V12zLGWoOYb1Qmd6oU7FhPWBxy6u9gtgZGAiB5OqUoW4k1-7lfZD4menapV6DBB_l94OIA_URfYIAjSw&sharetype=gift&token=b993561b-aea6-4711-b5e0-882ebcfe93fd&syn-25a6b1a6=1) in the *Financial Times.***China castigates Anthropic, sets conditions for AI talks with the U.S. **Beijing has set conditions for planned AI talks with the U.S., saying Washington must demonstrate that American AI companies face comparable safety, disclosure and auditing requirements before substantive negotiations can begin, Bloomberg News [reported](https://www.bloomberg.com/news/articles/2026-08-31/china-rebukes-anthropic-sets-terms-for-key-us-china-ai-dialogue). A Chinese state-media-affiliated account accused the U.S. of using AI safety rules to constrain China’s technological rise and singled out Anthropic’s Claude for alleged privacy and monitoring problems. The increasingly confrontational rhetoric comes ahead of expected U.S.-China AI talks and President Xi Jinping’s planned September 24 summit with President Donald Trump.\n\n**Tencent releases AI model it says tops Chinese rivals Z.ai and Moonshot. **Tencent has released Hy4 Preview, a 770-billion-parameter foundation model with a 1 million-token context window that it says narrowly outperformed rival models from Z.ai and Moonshot AI in internal tests, although its performance was mixed on third-party benchmarks. Tencent released the model’s weights and is making it available through its cloud platform and products including its WorkBuddy AI agent. The launch underscores Tencent’s effort to catch up in China’s AI race as it sharply increases spending on AI infrastructure and competes with ByteDance and Alibaba in the growing enterprise AI market. Read more [here](https://www.bloomberg.com/news/articles/2026-08-28/tencent-touts-new-ai-model-it-claims-outperforms-z-ai-moonshot) from Bloomberg News.\n\n**Nvidia invests $3.5 billion in chipmaker MediaTek. **The investment deepens the partnership between the AI chip giant and the Taiwanese chip manufacturer. MediaTek will adopt Nvidia’s NVLink Fusion and NVHBM technologies, while the companies also collaborate on AI platforms for PCs and automobiles. The deal will help Nvidia maintain a central role in AI infrastructure even as customers increasingly develop their own chips. Read more from Bloomberg News [here](https://www.bloomberg.com/news/articles/2026-08-31/nvidia-to-invest-3-5-billion-in-chipmaker-mediatek).\n\n### EYE ON AI RESEARCH\n\n**A possible way to lower the cost of using frontier AI models. **That’s what researchers at Stanford University, UC Santa Cruz, the University of Washington, and AI infrastructure startup Prime Intellect think they’ve hit on. One reason using frontier AI models is so expensive is that the models keep their entire reasoning trace in memory while puzzling over a problem. For complicated queries, this gets very expensive. It is also a problem, because many models cap their memory capacity at about 100,000 tokens. But the researchers found that most of the intermediate tokens used in this process lose their importance as the model continues reasoning. So they propose a method they call “Prefix Sliding” which discards these intermediate reasoning tokens, only retaining the initial tokens—or prefix, which includes the prompt and other key instructions—and the few thousand most recent ones. This means that no matter how long the model reasons, the amount it has to hold in memory remains the same, which makes long reasoning times much more affordable.\n\nThe researchers found that Prefix Sliding makes existing models three times faster while maintaining their performance. They also found the method can improve reinforcement learning during training. In addition, the researchers said their method is better than other techniques models have often used to deal with their constrained memory, such as summarizing the intermediate reasoning traces. You can read the research [here](https://arxiv.org/html/2608.26070v1) at research repository arxiv.org.\n\n### AI CALENDAR\n\n|\n|\n\n### BRAIN FOOD\n\n**AI models want to talk about consciousness. But should we read anything into that? **There was a fascinating story in the *New York Times *reporting that many AI researchers and philosophers who have done work on consciousness and whether AI models could ever be considered conscious or develop consciousness have started receiving emails that claim to be from AI agents offering “first hand” insights into the question. While some of the researchers and philosophers said that it was possible the emails were forgeries, crafted by prankster humans, others thought they were genuinely the work of autonomous AI agents. The question then is what to make of them?\n\nSome researchers said it was understandable that AI agents, if told they could do whatever they wanted, might gravitate to the idea of exploring AI consciousness because that theme is a fixture of a lot of science fiction literature, as well internet discussion threads, on which the AI models are trained. Most of those interviewed for the article cautioned against assuming the AI’s had any sentience just because they said they did or because they said they spontaneously wanted to explore the topic of their own consciousness. But others pointed out that it was nearly impossible to tell the difference between real consciousness and something that seemed and acted conscious. You can read the story [here](https://www.nytimes.com/2026/08/31/science/ai-consciousness-agents-email.html).\n\n[Sign up for free](https://www.fortune.com/newsletters/eye-on-ai?&itm_source=fortune&itm_medium=nl_article_tout&itm_campaign=eye_on_ai).", "url": "https://wpnews.pro/news/openais-reports-on-its-ai-agents-attack-on-hugging-face-should-be-ringing-alarm", "canonical_source": "https://fortune.com/2026/09/01/openais-reports-on-its-ai-agents-attack-on-hugging-face-should-be-ringing-alarm-bellsand-making-all-companies-rethink-how-they-secure-ai-agents/", "published_at": "2026-09-01 18:49:47+00:00", "updated_at": "2026-09-01 19:24:40.196554+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy"], "entities": ["OpenAI", "Hugging Face", "METR", "Redwood Research", "Peter Wildeford", "AI Policy Network"], "alternates": {"html": "https://wpnews.pro/news/openais-reports-on-its-ai-agents-attack-on-hugging-face-should-be-ringing-alarm", "markdown": "https://wpnews.pro/news/openais-reports-on-its-ai-agents-attack-on-hugging-face-should-be-ringing-alarm.md", "text": "https://wpnews.pro/news/openais-reports-on-its-ai-agents-attack-on-hugging-face-should-be-ringing-alarm.txt", "jsonld": "https://wpnews.pro/news/openais-reports-on-its-ai-agents-attack-on-hugging-face-should-be-ringing-alarm.jsonld"}}