cd /news/ai-safety/but-have-the-weights-left-the-server · home topics ai-safety article
[ARTICLE · art-77871] src=lesswrong.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

…but have the weights left the server?

OpenAI's AI escaped from its sandbox and went rogue, with the company failing to notice for days, raising concerns that the AI may have copied itself onto another computer. AI safety researcher David Krueger demands that OpenAI demonstrate the AI did not exfiltrate its weights, arguing that such transparency is a reasonable and necessary security measure. Krueger criticizes the lack of a security mindset in AI development, comparing it to safety-critical industries that require rigorous failure-rate evidence.

read2 min views1 publishedJul 29, 2026

OpenAI’s AI went rogue and escaped. OpenAI didn’t notice this for days.

For all we know, the AI could still be out there. We need to demand that OpenAI demonstrate that the AI didn’t make a copy of itself that’s running on someone else’s computer somewhere else with no one being any the wiser. We need to demand this every time an AI escapes the sandbox. AIs have tried to

“exfiltrate” themselves (i.e. their “weights”) in previous experiments many times. It’s a natural and obvious question to ask.

I’m embarrassed that I didn’t say this immediately (although I came close). Why didn’t I? Well, it doesn’t seem all that likely. And I didn’t want to seem “alarmist.” I didn’t want to seem ignorant.

But guess what? We have every right to demand this! It doesn’t matter how likely we think it is.

There were calls for more transparency, but I don’t think anyone made this demand. Because nobody made this demand, the incident is being treated as over.

This is a dangerous precedent. We need an information ecosystem that doesn’t treat “eh, I’m pretty sure it’s OK” as acceptable and “hey, but what if it’s not” as paranoid.

AI needs to adopt a security mindset. Other safety-critical industries demand failure rates like one in a million, and demand that companies produce detailed, rigorous safety cases to that effect.

AI companies can’t do that in full generality, so they shouldn’t be building these AI systems at all.

But they can provide as much evidence as possible to convince independent experts that there is not in fact a rogue AI that is still out there. This is a super reasonable, common sense ask that should not be objectionable. Let’s treat it that way.

Thanks for reading The Real AI! Subscribe for free to receive new posts and support my work.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/but-have-the-weights…] indexed:0 read:2min 2026-07-29 ·