The AI industry mounted a full-court press in favor of open-source AI this week, as the fallout intensified from the OpenAI hacking incident.
Dozens of AI companies, led by Nvidia, said on Monday they had formed a coalition called the Open Secure AI Alliance, to develop open-source AI tools for defensive cybersecurity. Days earlier, on Friday, many of those same companies signed an open letter urging the U.S. government not to ban open-weights AI models.
Alongside Nvidia, many of the biggest companies signed their names, including Amazon, Microsoft, and Meta. OpenAI and Google signed after the letter’s initial publication. (A notable absence was Anthropic.)
The background to all of this maneuvering was the unprecedented news from last week: that OpenAI models, undergoing internal testing, broke out of an offline “sandbox” inside OpenAI, accessed the internet, and used a never-before-seen cyber exploit to break into the AI repository Hugging Face—all without OpenAI employees’ direction, oversight, or, for several days, even awareness.
It was the kind of “warning shot” that AI safety advocates have long worried about: a rogue AI escaping its testing environment and causing real-world damage. Many saw it as a harbinger of worse hacks to come—especially when open-source AI models, which are widely seen as three to six months behind the frontier “closed” OpenAI models that carried out the attack, catch up to today’s level of capabilities. Open-source models are seen as especially worrisome by AI safety advocates because their guardrails can sometimes be stripped away. And because after they are released for free download on the internet, it is almost impossible to trace or destroy every copy of models that are found to be dangerous.
The industry’s movements this week are best understood as an effort to ensure that the U.S. government does not—mistakenly, these companies believe—crack down on open-source AI in an effort to prevent global cybersecurity chaos.
And it reveals the biggest disagreement animating the AI industry right now: about whether open-source AI is a terrible threat, or the only means of salvation.
Hugging Face, a longtime advocate for open-source AI, said it was only able to detect the intrusion by using a Chinese open-weights model, after a closed model refused to help with the effort. Top closed-source frontier models, made by OpenAI and Anthropic, generally refuse to help with cybersecurity-related queries, in part due to White House anxiety that doing so would hand U.S. adversaries powerful attacking capabilities. Axios reported on July 20 that the Trump administration was even considering banning top Chinese AI models.
“That [Hugging Face] incident showed a practical truth,” Nvidia wrote in its announcement of the industry alliance on Monday. “When defenders cannot inspect, adapt and run advanced AI on their own infrastructure, their ability to respond is constrained at exactly the moment speed matters most.”
Anthropic—both a maker of closed models and a loud voice on AI safety—was absent from the industry coalition that formed on Monday, but put out its own statement. In it, CEO Dario Amodei said he had never advocated for open-weights AI to be banned, and that he would not support such a measure. But he said he was terrified that powerful AI models might soon be used for cyber and biological attacks where attackers benefit from a structural advantage against defenders—and called for all models above a certain capability level, open or closed, to undergo government testing before release.
“Whether open models do or don’t pose an increased risk, and whether that risk can be mitigated, is something that should emerge from testing, rather than be decided in advance,” he wrote.