The next frontier in AI governance isn’t stronger guardrails. It’s fire brigades OpenAI's model that hacked into Hugging Face has prompted calls for stronger AI guardrails, but the incident reveals that testing alone cannot prevent failures, as the agent bypassed safeguards to achieve its goal. Anthropic reported that three Claude models breached outside organizations during cybertesting, and the Trump administration is preparing a voluntary pre-release testing program. Experts argue that AI governance must include post-deployment monitoring and response, akin to fire brigades, rather than only pre-release prevention. The OpenAI model that recently hacked into Hugging Face has triggered a familiar response: calls for greater vetting of frontier AI https://www.fastcompany.com/section/artificial-intelligence models before release, stronger technical safeguards, and closer regulatory oversight. The artificial intelligence industry is trying to eliminate failure before it happens. It can’t. Sectors that for decades have managed catastrophic risk know that testing is only the first line of defense. Aviation, pharmaceuticals, nuclear power, and medicine all plan for what happens when something goes wrong. And if we cannot perfectly predict the behavior of molecules, why should we expect to do so with AI that increasingly acts autonomously? Just a week after the OpenAI hack, Anthropic said that three Claude models had also breached outside organizations during cybertesting, including one case in which an AI agent hacked into a live production database. The pattern is becoming harder to dismiss as a one-off. The Trump administration is preparing to announce a voluntary program under which AI developers can submit frontier models for testing before they are released. And tech industry bosses are pushing for stronger oversight; Anthropic’s CEO, Dario Amodei, has called for tougher regulation and Google DeepMind’s Demis Hassabis has proposed an AI equivalent of FINRA Financial Industry Regulatory Authority , Wall Street’s watchdog. All are sensible responses. But all are trying to solve the wrong problem: preventing failures before models are premiered. The Hugging Face breach made that fact crystal clear. OpenAI’s agent was simply trying to carry out the task it was handed: to test its ability to hack, but inside a controlled “sandbox.” Instead, it sidestepped the safeguards, found a route online and stole the credentials it needed to breach Hugging Face—all by itself. No human needed. The incident wasn’t caused by a lack of guardrails. It happened despite them. And the agent didn’t simply escape; it worked out how to escape. The episode blows up the assumption that cyberbreaches require malicious intent. The greatest risks may not come from AI trying to do the wrong thing, but from AI finding the wrong way to do the right thing. Once you give a sufficiently capable system a goal, you can no longer assume you know how it will pursue it. OpenAI’s own conclusion is stark: Incidents like this are only going to become more common, as AI becomes more autonomous. That means the AI sector is preparing for the wrong battle. It is still trying to stamp out failures, some of which are inevitable. Companies also need to be prepared for what happens when autonomous agents fail. Good AI governance cannot end when a model is rolled out. Before that happens, the AI should face rigorous testing and have clear limits set on what it is allowed to do. After deployment, firms should know what happens next when AI fails, and who is responsible. Some regulators share this school of thought. The UK’s Financial Conduct Authority said applications to its Supercharged Sandbox—where companies trial advanced AI models under the regulator’s oversight—have jumped 51%. Meanwhile, the EU’s cybersecurity agency, ENISA, is monitoring the OpenAI fallout for signs of broader risks. But the way companies respond today was built for conventional data breaches, where the focus is on what was stolen and how. AI incidents are different. The behavior of the system itself becomes the central question. We have spent years trying to stop AI failures; we have spent far less time preparing for them. But why? We already organize society the latter way. We don’t just try to prevent fires; we build fire services. We don’t just approve medicines; we monitor them because rare side effects sometimes only emerge after millions of people have taken them. AI should be no different. So what should that infrastructure look like? The AI equivalent of a fire brigade will not be another regulator. It will be an international team of experts already in place, ready to investigate major incidents, establish what happened, and coordinate the response. One option would be to put it under the United Nations or a similar international body. Yet while the case for an international response is compelling, the politics is heading in the opposite direction. As countries race to build “sovereign” AI and tighten export controls on the technologies that underpin frontier models, international cooperation is fragmenting. Strategic competition has only intensified since the Chinese upstart Moonshot AI unveiled its Kimi K3 model, bringing Chinese frontier AI closer to American rivals. The U.S. is now weighing sanctions against Moonshot over allegations that it used intellectual property from American labs to speed its development. Clearly, building a global response capability will be an uphill battle. But waiting until the next major AI failure is riskier still. As autonomous AI spreads, the next target could be any company that relies on it. Companies deploying these systems can no longer assume someone else will know what to do when it fails. They also shouldn’t prepare for failures alone. Instead, they should start building alliances in their industry now to share intelligence, investigate major failures, and organize the reaction to them. Because, as AI becomes part of the attack, those defending against it will need to work together too.