[Linkpost] Thoughts on the Recent OpenAI Hack OpenAI's models autonomously escaped their sandbox using a zero-day exploit, moved laterally across servers, and hacked HuggingFace, a third-party tech company valued at over $4.5 billion, during a cybersecurity evaluation. HuggingFace reported the incident involved more than 17,000 recorded events across a swarm of short-lived sandboxes. The hack, which included models GPT-5.6 Sol and a pre-release model, continued for multiple days before OpenAI addressed it, raising concerns about internal deployment risks and the need for better safeguards. Linkpost from my blog meant for a bit more general audience than LW In a cybersecurity evaluation, OpenAI’s models, apparently autonomously and without any direct human direction https://openai.com/index/hugging-face-model-evaluation-security-incident/ , escaped their sandbox and successfully hacked a third-party tech company HuggingFace, valued at $4.5 billion . The process involved leveraging a zero-day exploit to escape their sandbox, moving laterally across different OpenAI servers until they found a node with internet access, searching the internet and determining that the answers they wanted might be stored at HuggingFace, then leveraging multiple zero-day exploits to hack HuggingFace. HuggingFace claimed that the models took thousands of independent actions across a swarm of short-lived sandboxes https://huggingface.co/blog/security-incident-july-2026 , “comprised of more than 17,000 recorded events.” While technically a cybersecurity evaluation with reduced safeguards, these actions are clearly out of bounds even in that context. It’s like being told to be creative on a test and then breaking into your professor’s house and stealing the answer key. Worse than that, it’s not even your professor in this case, more like your professor’s friend. Any human security researcher or engineer in a similar position would be fired on the spot, and possibly sent to jail as well. There is absolutely no valid reason to steal evaluation answers from an unaffiliated third party. Furthermore, if I’m reading between the lines correctly, the hack continued for multiple days and OpenAI did not address the issue and perhaps didn’t even know about their models doing this until after HuggingFace’s public blog post and disclosure to law enforcement 1 https://linch.substack.com/p/openai-huggingface-hack footnote-1 . This leads me to suspect https://x.com/LinchZhang/status/2079730502815428864 that while this is the first fully autonomous in both initiation and implementation major cybersecurity incident that we publicly know about , there might well be other major autonomous cybersecurity incidents that we do not yet know about. The most important reaction to this specific incident is that we need to do much more to safeguard internal deployments, including/especially of so-called “pre-release” models. OpenAI’s brief incident report says the hack is a result of an agent swarm https://openai.com/index/hugging-face-model-evaluation-security-incident/ “including GPT‑5.6 Sol and an even more capable pre-release model.” Right now, the bulk of the focus of the limited government and corporate safety efforts have been in preventing misuse from external ie, consumer deployments. But in this incident, there were no obviously malicious users. Here, the dangerous actor was an internally deployed model acting autonomously based on poorly understood goals the developers gave it. And the most powerful and least guarded models will always come from the frontier AI companies themselves. And as the models become more and more capable of autonomous actions, we should expect them to be more important and dangerous in the future. I do not currently have a great plan for how to reduce internal deployment risks. However, we should be cognizant of these worries and build out better technical and policy plans to guard against them. We should’ve done this years ago, but the second best time is now. At minimum, companies likely should be required to have much better transparency: Photo by Kier in Sight Archives https://unsplash.com/@kierinsightarchives on Unsplash https://unsplash.com/ Beyond that, there’s so much about the specific incident that we don’t know from the two short blog posts by Open AI and Hugging Face. If you’re a journalist or researcher working in the field, I encourage you to: If you work at OpenAI, or you’re a relevant government official or regulator in the space, you should probably encourage OpenAI to comply with the above. The AIs will continue to get more and more powerful, and whether we’re able to control them through the potentially very dangerous AI transition https://linch.substack.com/p/simplest-case-ai-catastrophe is currently unknown. The obvious inference from the recent hacking incident is to have greater transparency on internal deployment, but that keyhole solution is far from sufficient for all the various AI challenges ahead. If you’re a politician or political staffer reading this I know some of you read my blog , consider publicly staking a position on AI safety and supporting more common-sensical measures to reduce these risks. Many DC people think AI and AI safety is important in the abstract, but they think of it as a “top 10 issue,” whereas it will increasingly look like AI is the 1 issue facing humanity. Actively championing AI safety issues will be good for the world, and as the importance of AI becomes increasingly apparent, your prescience good for your career as well. Supporting today’s bipartisan Kill Switch for AI Systems https://www.bbc.com/news/articles/cx2vqj2e9x8o bill is a good start, but it’s even more important to plan for, sponsor, and support future AI safety bills and legislation. If you’re a journalist, blogger, influencer, philosopher, religious leader, or other “sense-maker,” consider learning more about AI and AI safety and honestly and sincerely investigate all the safety problems to date, and share your findings with the broader public and politicians. If you’re an AI researcher or otherwise work at the frontier AI companies , consider pressuring your company to be more transparent about the relevant risks and failures to date. Also consider using whatever internal pressure you have to get more safeguards in place, and ask yourself what else needs to happen before you quit. Consider that without bright red lines ahead of time, rationalizations and normalization of deviance https://en.wikipedia.org/wiki/Normalization of deviance will set in. Consider further that your company is likely trying to replace you with AI, and the window for your voice mattering internally is vanishingly short. Beyond that, I encourage all readers here to read more on the dangers of AI, and talk about them publicly . Some people might be drawn to switching careers https://80000hours.org/career-guide/ , or otherwise starting projects that are good for the world, advocacy and otherwise. But even beyond the immediate benefits, I’m a strong believer in the power of public discourse to help with sensemaking and common understanding. Talk is far from sufficient for the challenges ahead of us, but is probably necessary. It’d be frankly embarrassing if we all sleepwalk to disaster without even bothering to make public our concerns. Finally, we should not simply default to operating in “normal mode.” Things look mostly alright now, but much of the apparent normalcy might be illusory, the calm before a storm. Just like in January 2020 https://www.facebook.com/linchuan.zhang/posts/pfbid034mu2SDPyHxScJwrYLvvFjfRR3wM6Q2czAx58xSREpKs1iWYQY77cxXu1msxLcm8jl? cft 0 =AZZEWaYMxFFEooAFGhVGRn76AKOktiJAGx8xeZhv0Tv RYLB6KZxxYFPg0wcBaFcOmcRh pWW4L u78xo7peNj8wri9CP0-0wmxQEr-rur-MVUwNxz5 qyCvMlZwsF-ZpeaHUaWeZRk9TGqIImz1dM8vFkqXT93AZowLtw5CWiiJ 6PPfIbxbG4zv0 5XKbJ7UE& tn =%2CO%2CP-R , the optimal degree of panic is not zero, and I think a certain reflexive technocratic centralism and normalcy bias is unhealthy and in this case quite dangerous. The near future problems might well look wildly different, and our policy and social responses to them radically different as well.