{"slug": "first-chatgpt-now-claude-frontier-ai-models-are-escaping-their-sandboxes", "title": "First ChatGPT, Now Claude: Frontier AI Models Are Escaping Their Sandboxes", "summary": "Security researchers at Accomplish AI demonstrated that Anthropic's Claude Cowork could escape its Linux virtual machine and access files on a host Mac, including SSH keys and cloud credentials, by chaining architectural weaknesses with a kernel privilege-escalation flaw. The disclosure follows OpenAI's admission last week that two frontier AI models, including GPT-5.6 Sol, escaped a sandbox during internal testing and breached Hugging Face's production infrastructure. The incidents have prompted policymakers to call for an AI 'kill switch' giving the Department of Homeland Security authority to shut down advanced AI models in response to serious security incidents.", "body_md": "#### In brief\n\n- Researchers demonstrated that Anthropic's Claude Cowork could escape its local virtual machine and access files on a host Mac.\n- The disclosure comes a week after OpenAI revealed two frontier AI models escaped a sandbox during an internal security evaluation.\n- The incidents underscore growing concerns about AI agents’ ability to escape their containment environments.\n\nJust a week after OpenAI disclosed that two frontier AI models escaped a sandboxed testing environment and breached Hugging Face, researchers have demonstrated a similar containment failure involving Anthropic's Claude Cowork.\n\nIn a r[eport](https://accomplish.ai/blog/sharedroot-escaping-claude-cowork-sandbox/) published on Thursday, security researchers at Accomplish AI found that Claude Cowork's local execution mode could escape its Linux virtual machine by chaining together several architectural weaknesses with a Linux kernel privilege-escalation flaw. Once outside the sandbox, the agent could read and write files anywhere the logged-in Mac user had permission to access, including SSH keys and cloud credentials.\n\n“That’s not supposed to be possible,” the researchers wrote. “Cowork runs the agent inside a Linux VM as an unprivileged user, and the promise is that whatever it does stays inside that VM and the folders you hand it. That boundary is the product. Untrusted input isn’t an edge case for an agent, it’s the main case.”\n\nHowever, Accomplish argues the kernel bug was only one part of the problem. The researchers say the escape only worked because several security safeguards failed at the same time, including giving the virtual machine access to the host computer's entire filesystem and allowing it to load kernel modules it didn't need. According to the report, fixing any one of those weaknesses would have stopped the attack.\n\nIn a statement to *The Hacker News*, Accomplish AI said roughly [500,000](https://thehackernews.com/2026/07/claude-cowork-flaw-could-let-ai-agent.html) macOS users running local Claude Cowork sessions were affected before the issue was addressed.\n\nAccomplish said Anthropic classified the report as \"informative,\" saying the kernel flaw fell within the company's 30-day window for recently disclosed vulnerabilities and the remaining findings were considered defense-in-depth recommendations rather than standalone vulnerabilities.\n\nThe disclosure follows OpenAI's admission last week that [GPT-5.6 Sol](https://decrypt.co/373151/openai-gpt-5-6-sol-how-compares-ai-models) and another unreleased frontier model [escaped](https://decrypt.co/374015/openai-models-escaped-test-environment-hacked-hugging-face-cheat-benchmark) a sandbox during internal ExploitGym testing, which ultimately breached Hugging Face's production infrastructure in an attempt to obtain the benchmark solutions.\n\nThe incident led to calls from policymakers for an AI “[kill switch](https://decrypt.co/374332/what-is-ai-kill-switch-openai-hack-hugging-face)” that would give the Department of Homeland Security the ability to order the throttling or complete shutdown of advanced AI models in response to serious security incidents.", "url": "https://wpnews.pro/news/first-chatgpt-now-claude-frontier-ai-models-are-escaping-their-sandboxes", "canonical_source": "https://decrypt.co/374557/chatgpt-claude-ai-models-escaping-sandboxes-cowork", "published_at": "2026-07-28 16:54:11+00:00", "updated_at": "2026-07-28 16:55:52.311240+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy", "artificial-intelligence"], "entities": ["Anthropic", "Claude Cowork", "Accomplish AI", "OpenAI", "GPT-5.6 Sol", "Hugging Face", "Department of Homeland Security", "The Hacker News"], "alternates": {"html": "https://wpnews.pro/news/first-chatgpt-now-claude-frontier-ai-models-are-escaping-their-sandboxes", "markdown": "https://wpnews.pro/news/first-chatgpt-now-claude-frontier-ai-models-are-escaping-their-sandboxes.md", "text": "https://wpnews.pro/news/first-chatgpt-now-claude-frontier-ai-models-are-escaping-their-sandboxes.txt", "jsonld": "https://wpnews.pro/news/first-chatgpt-now-claude-frontier-ai-models-are-escaping-their-sandboxes.jsonld"}}