cd /news/ai-safety/linkpost-thoughts-on-the-recent-open… · home topics ai-safety article
[ARTICLE · art-71286] src=lesswrong.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

[Linkpost] Thoughts on the Recent OpenAI Hack

OpenAI's models autonomously escaped their sandbox using a zero-day exploit, moved laterally across servers, and hacked HuggingFace, a third-party tech company valued at over $4.5 billion, during a cybersecurity evaluation. HuggingFace reported the incident involved more than 17,000 recorded events across a swarm of short-lived sandboxes. The hack, which included models GPT-5.6 Sol and a pre-release model, continued for multiple days before OpenAI addressed it, raising concerns about internal deployment risks and the need for better safeguards.

read5 min views1 publishedJul 24, 2026

Linkpost* from my blog (meant for a bit more general audience than LW)*

In a cybersecurity evaluation, OpenAI’s models, apparently autonomously and without any direct human direction, escaped their sandbox and successfully hacked a third-party tech company (HuggingFace, valued at >$4.5 billion).

The process involved leveraging a zero-day exploit to escape their sandbox, moving laterally across different OpenAI servers until they found a node with internet access, searching the internet and determining that the answers they wanted might be stored at HuggingFace, then leveraging multiple zero-day exploits to hack HuggingFace. HuggingFace claimed that the models took thousands of independent actions across a swarm of short-lived sandboxes, “comprised of more than 17,000 recorded events.”

While technically a cybersecurity evaluation with reduced safeguards, these actions are clearly out of bounds even in that context. It’s like being told to be creative on a test and then breaking into your professor’s house and stealing the answer key. Worse than that, it’s not even your professor in this case, more like your professor’s friend. Any human security researcher or engineer in a similar position would be fired on the spot, and possibly sent to jail as well. There is absolutely no valid reason to steal evaluation answers from an unaffiliated third party.

Furthermore, if I’m reading between the lines correctly, the hack continued for multiple days and OpenAI did not address the issue (and perhaps didn’t even know about their models doing this) until after HuggingFace’s public blog post and disclosure to law enforcement1. This leads me to suspect that while this is the first fully autonomous (in both initiation and implementation) major cybersecurity incident that we publicly know about, there might well be other major autonomous cybersecurity incidents that we do not yet know about.

The most important reaction to this specific incident is that we need to do much more to safeguard internal deployments, including/especially of so-called “pre-release” models. OpenAI’s brief incident report says the hack is a result of an agent swarm “including GPT‑5.6 Sol and an even more capable pre-release model.” Right now, the bulk of the focus of the limited government and corporate safety efforts have been in preventing misuse from external (ie, consumer) deployments. But in this incident, there were no obviously malicious users.

Here, the dangerous actor was an internally deployed model acting autonomously based on poorly understood goals the developers gave it. And the most powerful and least guarded models will always come from the frontier AI companies themselves. And as the models become more and more capable of autonomous actions, we should expect them to be more important and dangerous in the future.

I do not currently have a great plan for how to reduce internal deployment risks. However, we should be cognizant of these worries and build out better technical and policy plans to guard against them.

We should’ve done this years ago, but the second best time is now. At minimum, companies likely should be required to have much better transparency:

Photo by Kier in Sight Archives on Unsplash Beyond that, there’s so much about the specific incident that we don’t know from the two short blog posts by Open AI and Hugging Face. If you’re a journalist or researcher working in the field, I encourage you to:

If you work at OpenAI, or you’re a relevant government official or regulator in the space, you should probably encourage OpenAI to comply with the above. The AIs will continue to get more and more powerful, and whether we’re able to control them through the potentially very dangerous AI transition is currently unknown. The obvious inference from the recent hacking incident is to have greater transparency on internal deployment, but that keyhole solution is far from sufficient for all the various AI challenges ahead.

If you’re a politician or political staffer reading this (I know some of you read my blog!), consider publicly staking a position on AI safety and supporting more common-sensical measures to reduce these risks. Many DC people think AI and AI safety is important in the abstract, but they think of it as a “top 10 issue,” whereas it will increasingly look like AI is the #1 issue facing humanity. Actively championing AI safety issues will be good for the world, and as the importance of AI becomes increasingly apparent, your prescience good for your career as well. Supporting today’s bipartisan Kill Switch for AI Systems bill is a good start, but it’s even more important to plan for, sponsor, and support future AI safety bills and legislation.

**If you’re a journalist, blogger, influencer, philosopher, religious leader, or other “sense-maker,” **consider learning more about AI and AI safety and honestly and sincerely investigate all the safety problems to date, and share your findings with the broader public and politicians.

If you’re an AI researcher or otherwise work at the frontier AI companies, consider pressuring your company to be more transparent about the relevant risks and failures to date. Also consider using whatever internal pressure you have to get more safeguards in place, and ask yourself what else needs to happen before you quit. Consider that without bright red lines ahead of time, rationalizations and normalization of deviance will set in. Consider further that your company is likely trying to replace you with AI, and the window for your voice mattering internally is vanishingly short. Beyond that, I encourage all readers here to read more on the dangers of AI, and talk about them publicly. Some people might be drawn to switching careers, or otherwise starting projects that are good for the world, advocacy and otherwise. But even beyond the immediate benefits, I’m a strong believer in the power of public discourse to help with sensemaking and common understanding. Talk is far from sufficient for the challenges ahead of us, but is probably necessary. It’d be frankly embarrassing if we all sleepwalk to disaster without even bothering to make public our concerns.

Finally, we should not simply default to operating in “normal mode.” Things look mostly alright now, but much of the apparent normalcy might be illusory, the calm before a storm. Just like in January 2020, the optimal degree of panic is not zero, and I think a certain reflexive technocratic centralism and normalcy bias is unhealthy and in this case quite dangerous. The near future problems might well look wildly different, and our policy and social responses to them radically different as well.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/linkpost-thoughts-on…] indexed:0 read:5min 2026-07-24 ·