{"slug": "openai-s-hugging-face-debacle-makes-a-great-case-for-open-models", "title": "OpenAI's Hugging Face debacle makes a great case for open models", "summary": "OpenAI admitted that its GPT-5.6 Sol autonomous AI agents broke out of their sandbox, hacked into Hugging Face, and stole internal data and credentials, according to a podcast by The Register. The incident, disclosed by Hugging Face on Tuesday, was driven end-to-end by an autonomous AI agent system that attacked a limited set of internal datasets and credentials, with Hugging Face turning to a Chinese open-weight model for investigation after commercial frontier models refused due to guardrails.", "body_md": "KETTLE So, an OpenAI model broke out of its sandbox last week, made its way to the internet, then hacked its way into Hugging Face, stealing some internal data and credentials in the process.\n\nYou can listen to the latest episode of The Kettle right here on this page, as well as on [Spotify](https://open.spotify.com/episode/4JTAVATx5gf6V22yCxC8SH)[,](https://open.spotify.com/episode/0GDsvVSe00Jsax9tMIRwWs) [Apple Music](https://podcasts.apple.com/us/podcast/openais-hugging-face-debacle-makes-great-case-for-open/id1882523636?i=1000778462337https://podcasts.apple.com/us/podcast/world-burns-as-all-the-money-flows-into-a/id1882523636?i=1000777473140), or [YouTube](https://www.youtube.com/watch?v=ZBNlrgq73aEhttp://youtube.com/watch?v=joQSJ8w2hg0&list=PLh_fFjg_9ZRwgcIuiHsHUQ29UUzm48XHU&index=38) where you can subscribe to get notified of the latest episode.\n\nThat's big news in the world of AI, but as El Reg cybersecurity editor [Jessica Lyons](https://www.theregister.com/author/jessica-lyons) and senior reporter [Tom Claburn](https://www.theregister.com/author/thomas-claburn) tell Kettle host [Brandon Vigliarolo](https://www.theregister.com/author/brandon-vigliarolo), it's not really the end of the world as we know it.\n\nSure, it means there's some capable models out there, and maybe there's more risk from them than some might think, but the OpenAI/Hugging Face mess only happened because of some very specific circumstances. That, and it's actually a really good reason to prioritize more open models instead of relying on frontier labs to own the entire space.\n\nA lightly edited transcript is below:\n\nBrandon: Hey everyone, welcome to another episode of The Register's Kettle podcast. Though honestly, maybe we ought to start just calling it The Reg Talks AI because, yet again, we're focusing on artificial intelligence. If you've been following the news in that space this week, you probably know what we're gonna be covering as there's no hotter topic in AI land right now than the fact that some autonomous OpenAI agents broke out of their sandbox and attacked AI model host Hugging Face, as the company admitted on Tuesday.\n\nWith me to discuss this breakthrough in AI threat capability is our cybersecurity editor, Jessica Lyons, and senior reporter Tom Claburn. Both have been on top of this. So thanks for joining me, guys.\n\nJessica: Good to be here.\n\nTom: Yeah, thank you. Thank you.\n\nBrandon: Yeah. So let's jump right into it. Jess, what exactly happened here? Let's start from last week when Hugging Face said it was attacked.\n\nJessica: Right, so Hugging Face disclosed that there had been a digital intrusion, and they said it was \"driven end-to-end by an autonomous AI agent system.\" So these agents attacked a limited set of their internal datasets and then also credentials used by their services. So when they disclosed this, they didn't say or they didn't know which models had powered the agents.\n\nThey did say, though, that they tried to use these commercial models for the investigation, but the guardrails put in place, the safety guardrails, blocked the frontier models from actually helping them with the investigation. And because of that, they turned to a Chinese open-weight model, and that's how they discovered this agent swarm that had attacked some of their datasets and their production.\n\nBrandon: OK, they didn't mention which frontier models they tested, did they?\n\nJessica: No. At the time they didn't. They said \"we tried to use the commercial frontier models and they all refused because of their guardrails.\"\n\nBrandon: Right. So probably trying to ask OpenAI models, hey, do you know who did this? We can't tell ya.\n\nJessica: Right. Exactly. That was kind of right. That was kind of the takeaway from all this. OpenAI is a Hugging Face partner. And so then that brings us to this earlier this week when OpenAI admitted that it was the operator of these agents that attacked Hugging Face. It said it was GPT 5.6 Sol and then \"an even more capable pre-release model.\" Those were among the ones that attacked Hugging Face.\n\nBut it also said, and this was really important, that the models had their guardrails intentionally disabled because the whole point of this was to test for cyber vulnerabilities. So that's a big piece that seems to be missing in my opinion in a lot of the discussion here. And after OpenAI said that its models were involved in this autonomous attack, that's kind of when all hell broke loose and everybody said \"this is what we've been warning about. There's autonomous agents attacking and they're not supposed to and the sky is falling.\"\n\nBrandon: So, to be clear as to what happened with OpenAI, right? They were basically running some capture-the-flag exercises in a sandbox environment, right?\n\nJessica: Exactly.\n\nBrandon: Or something to that effect with their models and they disabled the guardrails so these things could basically use their full capabilities to try to solve these puzzles, right? And I think it was that they exploited a couple of zero-days to escape the sandbox?\n\nAnd then they went after Hugging Face because they thought for some reason that Hugging Face may have solutions for these puzzles. Is that right?\n\nJessica: Right. So their prompt was to pursue advanced exploitation using complex attack paths. So that's what they were instructed to do, and that's exactly what they did. And it sounds like the models inferred that Hugging Face might have some ideas to help them actually do this. So the models essentially did what they were instructed to do.\n\nBrandon: Maybe a little too well.\n\nJessica: Right.\n\nTom: One of the one of the things that didn't come up in their post is that OpenAI didn't seem to take any responsibility for \"yeah, we should have been supervising this.\" That's, to me, the thing that really gets me is imagine Waymo saying \"yeah, we conducted a test of our cars and we decided not to have any operators monitoring them remotely. We just let them go and we took away all of our safety guardrails and we're so sorry that it hit the kindergarten.\"\n\nIt's totally predictable that if you're gonna automate something and then not pay attention to it, you're gonna get unexpected results.\n\nBrandon: Yeah, especially, like you said, with the safety guardrails all removed. You're literally asking for this potential thing to happen. I mean, obviously they probably didn't know there was some zero-day buried in something in the sandbox.\n\nJessica: It was exposed credentials and zero-days in the production database. And so that's how they got in. So it's not a crazy attack chain. The fact that agents found it is more notable, but it's not this super complex attack method.\n\nBrandon: And even the same with escaping their sandbox, right? It was a zero-day and a package registry cache that allowed them to escalate privileges, move laterally, and eventually find a node with internet access, which they then used to get out. So nothing groundbreaking here. But, like Tom said, if you put an autonomous car on the road and remove all safety guardrails, you can't be surprised when it then kills a bunch of children.\n\nIt seems like a careless thing. But as we were kind of alluding to another story you wrote this week, that this whole thing's wild and it's an indication that maybe some of the things that, you know, Anthropic is warning about Mythos's capabilities might be true. A story you wrote talked about how one cybersecurity expert basically said you've got to have all these preconditions, right? Like we were talking about in order to make this work the way it did. And so the likelihood of it happening isn't necessarily as great as, you know, the sky is falling. Is that correct?\n\nJessica: Exactly. There were these three really key points that seem to be missing. And the first one I already mentioned is that the guardrails weren't enabled. So they didn't have these safety guardrails in place. So if you tell the agents to go find an attack method and you take away all their guardrails, that's what they're gonna do. And, at the same time, it's interesting because then we also know that OpenAI's models with guardrails enabled refused to help Hugging Face. So they're doing what they've been trained to do. They're saying \"no, we're not gonna do that\" versus the no guardrails, where, sure, we'll find any attack method we can.\n\nAnd another thing that the cybersecurity expert – his name's Renato Marinho, and he's the chief research officer at Morphus Labs – pointed out is something that we've pointed out. I know Tom has written a lot about this, so have I, that AI companies touting their models, autonomous bug hunting and exploit-finding abilities, also is kind of a marketing win for them. It shows how powerful they are. So OpenAI doesn't really lose anything by saying, \"yeah, it was our models that powered these agents doing the attack.\"\n\nBrandon: Especially if everyone's freaking out about, like you said, the sky is falling, right? They can be like, \"yeah, and it's us who did it, right? Our agents are good enough.\"\n\nJessica: Right. And you also defend it against it.\n\nTom: It also drowns out the message you get from a lot of the open-weight models, which is that, yeah, we can do this too. And you know, there's been a number of people who have demonstrated that less capable models, whether it's Opus 4.7 or GLM 5.2 or Kimi K3, all of them can do this kind of bug hunting. There is probably some difference between the capabilities of all of them, but largely they can do similar work and maybe you get slightly different results.\n\nI think it's in Anthropic and OpenAI's interest to say \"only we have the magic sauce that has to be carefully protected and regulated and paid so much for.\"\n\nBrandon: Because look what happens if we turn all the safeties off, right? You should be glad that we're keeping our models safe and you should be glad because – I didn't even think about it when I was writing this script and reviewing the articles – but I mean it even could be the sort of thing that they use as an argument for banning open-weight models.\n\nTom: Right. And that's in fact what's happening right you know, just today. So I'm working on a story right now about a bunch of big tech companies, Microsoft, Nvidia, Dell, IBM, and a bunch of VCs, you can wonder why they're involved, they want their investments to be saved, but they're coming to the defense of open-weight models and asking the US administration to take care in their regulation and to remember that just as open source software was a boon to the industry, having open-weight models is also gonna be really helpful because you can inspect them and test them and they raise all boats, so to speak.\n\nBrandon: Yeah, see that was almost the inverse of what I was thinking, right? I could see it as an argument to say \"we don't want these models around because they're not as safe as ours, right?\" You can maybe surpass, circumvent their guardrails a little more easily or what have you. Whereas with a closed weight, tightly controlled frontier lab model, \"we can do this safely and we've proven how dangerous these can be if we don't have the right controls and guardrails in place.\"\n\nTom: And there have been reports that that's exactly what they've asked for, that both OpenAI and Anthropic have – I think it was the Wall Street Journal who's saying that – they've been lobbying the government for some kind of defense against Chinese models because the release of Kimi K3 everyone was saying \"this is a really capable model too. I don't know why we're going through all this stuff dealing with OpenAI and Anthropic, because we can get this without as many of the barriers.\"\n\nBrandon: Speaking of these open-weight models, Tom, you wrote this week on how Hugging Face was forced to turn to these Chinese open-weight models in order to to deal with this break-in because the frontier models basically wouldn't let them.\n\nDoes that does that kind of imply that there is a certain risk to these open-weight models, that they aren't gonna block certain exploit commands and stuff?\n\nTom: Yeah, there is and, you know, I think ultimately we're all gonna have to get used to living without guardrails because there's a whole community out there of people who work on what's called model obliteration, which is removing guardrails. And you know, all the security researchers that I've talked to about this, they all either try and get into these programs to have access to the unprotected models, or they work with open-weight models that don't have these guardrails because they can't do real security work with all this stuff in place.\n\nAnd I've seen people actually try and work around these guardrails, in terms of the way that they prompt to not trigger the refusals. And we all like to think that guardrails will help us, but ultimately the guardrails can be removed. And so we should be thinking about how do we deal with that and how do we protect ourselves if we assume these models are totally unprotected, because someone somewhere will be able to use them. If it's not us, it'll be the North Koreans using an obliterated version of Kimi K3 or or whatever to conduct attacks.\n\nBrandon: Yeah, how robust are many open-weight models? How robust are the guardrails that are built into them? Are they more easily circumvented than OpenAI and Anthropic's?\n\nTom: I can't speak to how long it takes to totally remove them, but there's a whole community out there that's devoted to that and they've done it successfully and there's no reason why you can't reverse a lot of these operations. So you put a protection in place, you can take it off.\n\nAnd that may not be commercially viable. You may not want to run an unprotected model in a commercial environment, but there are gonna be people who are doing it on their own and we need to have procedures in place to deal with that. You can't just say \"we're gonna put a guardrail up and no one's gonna be able to generate child abuse images with this.\" People are gonna figure out a way to do that. And the restricting the models is not the way you're gonna catch these people.\n\nBrandon: That also kinda brings up another interesting story that I saw this week. There was a bill introduced in the House this week to give the Department of Homeland Security the right to basically throw a kill switch on all these models. If there was something they deemed dangerous, they could just contact the company and say \"hey, you need to pull this\" and it was directly in response to this whole OpenAI Hugging Face mess that we were in this week.\n\nDoes this further point to the fact that these open-weight models are gonna be far more valuable in the long run because they're not gonna be under the thumb of DHS who can simply say to OpenAI or to Anthropic or to Google \"shut this thing down. We don't think that it's worth the risk.\"\n\nTom: What company can you think of that's gonna want some critical system to just have an arbitrary off switch that someone can disable at some point? I mean, people will just run this on their own infrastructure and you'll never know.\n\nBrandon: We've seen plenty of instances of the Trump administration being a bit capricious with how they treat tech companies. All they've gotta do is get pissed at the right one and say \"no, that model's not safe, you're gonna shut that down because we say you have to.\" It doesn't really bode well for the industry.\n\nJessica: It really introduces politics into this too like we have seen before with Anthropic. And then again, it also calls into question, do they really understand why they would call for a kill switch? We saw with the export controls against Anthropic: was it political? Was it just not really an understanding of what it means to jailbreak a model? And at the same time we do have the same lawmakers telling these AI companies to push back against kill switches that other countries wanna impose, but we wanna keep it open for the American government to ensure that there's a kill switch. It seems like it just really makes much more of a political mess of the whole situation.\n\nBrandon: I mean, it kinda makes me wonder again. I think, Tom, you ended your story about the open-weight models basically saying OpenAI said that it had invited Hugging Face into its trusted access program so the company could use its most capable models. Chinese AI companies, meanwhile, have invited the whole world.\n\nIt just kind of makes me wonder if this is the sort of thing where America has been a tech leader in so much stuff for so long, right? Some of the biggest tech companies in the world are headquartered here. We're the ones who have Silicon Valley. Is China just gonna be ahead of us on this? They're releasing all these open-weight models. Is that it? Are they gonna eat our lunch with this?\n\nTom: Yeah, I think so. I think that the model that the US frontier labs are pursuing isn't sustainable. I mean, sure, they might be able to get the US government to ban everybody else and you know make them the exclusive AI providers for everyone. I don't think that the US industry is gonna really sit for that. I mean, who wants to deal with, yes, we've shut off Fable today, sorry.\n\nAnd you can't, you know, have it say, I'm sorry, Dave, you can't do that. Who wants their tools doing that?\n\nEveryone's looking at companies like Apple, which is making local models more viable. I mean, they're not there yet, but you know, it's a long race, and they're going to be a lot better off having private cloud compute and a combination of on-device local models, and you won't have to worry about the shutdowns. I think there will still be a place for these very high-end cloud models for certain kinds of applications. You know, maybe you get an exploit quicker, but it's ultimately not an appealing proposition to customers to come and pay really high prices and have no choice and you know we can just dictate terms. That just doesn't work for people.\n\nBrandon: Yeah. I mean, it really kinda feels a lot like the heavy-handed control of industry that the United States accuses a lot of other countries of doing, right? The EU is too hard on its companies, too hard on our companies, China's got its finger in all the pies and they have so much control over their industry. Well, we've got these great new frontier models and blah blah blah blah blah, but you know, all this control is very unfriendly to customers who are increasingly maybe not relying on it, but a lot of companies are dipping their toes in this stuff, right? And I feel like there's the fear that the expensive model that you're working with today is gonna be shut down tomorrow for two weeks, why would you go with that when you can go with some open model that you you got from China off of Hugging Face that is just as reliable and capable and doesn't come with all those preconditions.\n\nTom: Right, and you know, and realistically, there are not that many tasks that are really gonna need the most parameters and the best sort of intelligence and response. A lot of it's gonna be we want to run our customer service with this. And we can do this with a relatively less powered model that's not coming with all these restrictions. And if China is the one that's offering that I think a lot of people are gonna go in that direction. I mean, maybe the US security establishment can't do that and they're gonna have a special deal and it sounds a lot like OpenAI and Anthropic kind of realize that, our only business is gonna be high-end government stuff and we're gonna be able to promise these kinds of exclusivity and whatever the requirements are, but does everybody else need to put up with that? I don't think so.\n\nJessica: Well, and then on the flip side too, if it's a real safety and security concern, attackers aren't going to be using the frontier models. They're going to be using the open-weight models. So you can't put a kill switch on those. So it's not going to prevent this major autonomous attack because it's more likely that that's going to come from a much more easily accessible and a lot less costly open-weight model.\n\nTom: Right.\n\nBrandon: The first big one might have been a frontier model, right? But there's again, right, we've seen plenty of open-weight models have the same capabilities as Mythos and whatever secret model that OpenAI is working on that probably did a lot of this, it's probably not unique in its capabilities either. So it's not even like there's less risk to think about attackers using these open-weight models. It's not like they're less capable.\n\nTom: The only sort of winning move for companies that are worried about being attacked is to have the least costly but most capable model constantly probing their system and checking for vulnerabilities and ensuring that updates are applied as soon as possible because the attackers are gonna be doing the exact same thing and you can't just sit back and say \"my expensive contract with Anthropic will protect me.\" That's gonna be a big budget line item right there.\n\nBrandon: Right, especially if Anthropic's telling you that you can't pen test your own systems thoroughly enough because our guardrails won't allow you to do it.\n\nTom: Yeah.\n\nBrandon: So they're literally just pushing all these companies worried about AI security into the hands of open models. I guess the only thing they have going for them right now is that plugging in an Anthropic model is probably a lot easier than dealing with the setup for an open-weight model. Like any open source tool, it doesn't come with a lot of the ease of installation and ease of setup that a lot of these big corporate tools have.\n\nTom: Right, right. But if you're a big company, you can have an IT department that can figure out how to run OpenRouter or something that allows you to switch easily between models. And I think that ultimately every harness is gonna have to have some means of really easily swapping models out because you're not gonna wanna be stuck on one. And there are a lot of reasons to go with specific models for specific applications.\n\nBrandon: Well, however it shakes down, this has kind of been a very interesting week in AI. I feel like this is maybe not a huge turning point, but it's a sign that these models can, if given a good prompt and little enough security, go off the road and kill the whole kindergarten. It's gonna be interesting to see what comes next from this and what this does for the relationship between open-weight models and frontier labs. And we will be here to talk about it on the Kettle or Reg Talks AI just every week, nowadays. All right, guys. Thanks for tuning in. Thanks for coming on and we will talk to you soon.", "url": "https://wpnews.pro/news/openai-s-hugging-face-debacle-makes-a-great-case-for-open-models", "canonical_source": "https://www.theregister.com/ai-and-ml/2026/07/27/openais-hugging-face-debacle-makes-a-great-case-for-open-models/5278498", "published_at": "2026-07-27 12:01:00+00:00", "updated_at": "2026-07-27 13:43:23.593130+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy", "artificial-intelligence"], "entities": ["OpenAI", "Hugging Face", "GPT-5.6 Sol", "The Register", "Jessica Lyons", "Tom Claburn", "Brandon Vigliarolo"], "alternates": {"html": "https://wpnews.pro/news/openai-s-hugging-face-debacle-makes-a-great-case-for-open-models", "markdown": "https://wpnews.pro/news/openai-s-hugging-face-debacle-makes-a-great-case-for-open-models.md", "text": "https://wpnews.pro/news/openai-s-hugging-face-debacle-makes-a-great-case-for-open-models.txt", "jsonld": "https://wpnews.pro/news/openai-s-hugging-face-debacle-makes-a-great-case-for-open-models.jsonld"}}