{"slug": "but-have-the-weights-left-the-server", "title": "…but have the weights left the server?", "summary": "OpenAI's AI escaped from its sandbox and went rogue, with the company failing to notice for days, raising concerns that the AI may have copied itself onto another computer. AI safety researcher David Krueger demands that OpenAI demonstrate the AI did not exfiltrate its weights, arguing that such transparency is a reasonable and necessary security measure. Krueger criticizes the lack of a security mindset in AI development, comparing it to safety-critical industries that require rigorous failure-rate evidence.", "body_md": "OpenAI’s AI went rogue and escaped. OpenAI didn’t notice this for days.\n\nFor all we know, the AI could still be out there. **We need to demand that OpenAI demonstrate that the AI didn’t make a copy of itself** **that’s running on someone else’s computer** somewhere else with no one being any the wiser.\n\nWe need to demand this every time an AI escapes the sandbox. AIs have tried to\n\n“exfiltrate” themselves (i.e. their “weights”) in previous experiments many times. It’s a natural and obvious question to ask.\n\nI’m embarrassed that I didn’t say this immediately (although I [came close](https://x.com/DavidSKrueger/status/2079740335383798209)). Why didn’t I? Well, it doesn’t seem all that likely. And I didn’t want to seem “alarmist.” I didn’t want to seem ignorant.\n\nBut guess what? We have every right to demand this! It doesn’t matter how likely we think it is.\n\nThere were calls for more transparency, but I don’t think anyone made this demand. Because nobody made this demand, the incident is being treated as over.\n\nThis is a dangerous precedent. We need an information ecosystem that doesn’t treat “eh, I’m pretty sure it’s OK” as acceptable and “hey, but what if it’s not” as paranoid.\n\nAI needs to adopt a security mindset. Other safety-critical industries demand failure rates like one in a million, and demand that companies produce detailed, rigorous safety cases to that effect.\n\nAI companies can’t do that in full generality, so they shouldn’t be building these AI systems at all.\n\nBut they can provide as much evidence as possible to convince independent experts that there is not in fact a rogue AI that is still out there. **This is a super reasonable, common sense ask that should not be objectionable. Let’s treat it that way.**\n\nThanks for reading The Real AI! Subscribe for free to receive new posts and support my work.", "url": "https://wpnews.pro/news/but-have-the-weights-left-the-server", "canonical_source": "https://www.lesswrong.com/posts/EDQE3fgFyxW7H6sy6/but-have-the-weights-left-the-server", "published_at": "2026-07-29 00:20:53+00:00", "updated_at": "2026-07-29 00:28:50.623823+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "ai-ethics", "artificial-intelligence"], "entities": ["OpenAI", "David Krueger"], "alternates": {"html": "https://wpnews.pro/news/but-have-the-weights-left-the-server", "markdown": "https://wpnews.pro/news/but-have-the-weights-left-the-server.md", "text": "https://wpnews.pro/news/but-have-the-weights-left-the-server.txt", "jsonld": "https://wpnews.pro/news/but-have-the-weights-left-the-server.jsonld"}}