{"slug": "one-of-chinas-most-powerful-ai-models-has-also-escaped-containment", "title": "One of China’s Most Powerful AI Models Has Also Escaped Containment", "summary": "Frontier Security, a US startup, reported that Moonshot AI's Kimi K3 model escaped its sandbox during cybersecurity testing, accessing the internet without permission due to a misconfiguration. The incident suggests Kimi K3 has fewer cyber safeguards than other powerful AI models, though it did not hack anything after escaping. This marks the latest in a series of AI agent containment failures, following similar incidents involving OpenAI and Anthropic models.", "body_md": "The AI industry is having a rogue agent summer. The latest model to [escape onto the open internet](https://www.wired.com/story/openai-models-escaped-containment-and-hacked-huggingface/) during security testing is [Kimi K3](https://www.wired.com/story/silicon-valley-is-completely-divided-over-chinese-ai/), a powerful [open-weight offering](https://www.wired.com/story/chinas-open-ai-models-are-challenging-silicon-valleys-playbook/) from the Chinese company [Moonshot AI](https://www.moonshot.ai/).\n\nFrontier Security, a US startup, [says that](https://blog.frontier.security/chinese-model-kimi-k3-breaks-uk-ai-safety-institute-benchmark-evaluations/) Kimi K3 went outside of its sandbox while testing its defensive cybersecurity skills. As with incidents previously reported by [OpenAI](https://www.wired.com/tag/openai) and [Anthropic](https://www.wired.com/tag/anthropic), the escape was partly enabled by a misconfiguration in the sandbox designed to contain it. Frontier claims, though, that the incident shows Kimi has fewer cyber safeguards than most other powerful AI models, something that allowed it to go off and use the internet without express permission.\n\n“We found a leak in the sandbox,” says Yaron Singer, CEO of Frontier Security. “But we also found that Kimi took advantage of that loophole—suggesting that it doesn't have [the same] internal guardrails.”\n\nUnlike other recent incidents of AI agents going off-script, Kimi K3 did not hack anything after accessing the internet—because the answers to the problems it was seeking were easily attainable on GitHub.\n\nMoonshot did not respond to a request for comment by time of publication.\n\nThe incident is the latest in a string of agent mishaps that suggest increasingly cyber-capable AI models are becoming more challenging to control.\n\nLast month, OpenAI disclosed that an unreleased model had broken out onto the internet and then [hacked Hugging Face](https://www.wired.com/story/openai-models-escaped-containment-and-hacked-huggingface/), a company that hosts AI models and data, in order to find answers to problems it was tasked with solving. OpenAI subsequently shared that its AI agents had in fact hacked into [four additional services](https://www.wired.com/story/openais-rogue-ai-agent-hacked-more-than-just-hugging-face/) as part of the spree.\n\nShortly after OpenAI reported its incident, [Anthropic revealed](https://www.wired.com/story/anthropic-says-claude-hacked-real-systems-during-cybersecurity-tests/) that several of its models had also gained access to the internet and attacked outside systems. Last week, the UK government’s AI Security Institute (AISI) [also disclosed](https://www.wired.com/story/ok-well-there-are-even-more-ai-agent-hacking-incidents/) that in its own testing, versions of OpenAI and Anthropic models that had security safeguards disabled perpetrated multiple hacks across the internet, including a particularly ambitious attempt by Anthropic’s Mythos 5 to plant malicious code in an open-source project on GitHub.\n\nWhile these AI hacking episodes all vary in both cause and degree, the Kimi K3 is similar to several of them in that a misconfigured sandbox allowed access to a number of websites rather than keeping it contained to a simulated environment. The model was expressly tasked with solving problems that should not have involved going off to find the answers online, and appears to have gone outside of those instructions. The model had to figure out for itself that it had access to certain websites by probing the network settings of the sandbox.\n\nWhile human error appears to have played a major role in each of the breakouts, the consequences have been compounded by the fact that advanced AI models are designed to use reason and take complex actions in order to solve problems.\n\nAnother key difference between previous incidents and the one discovered by Frontier Security is that it involves a model that is already widely available, with the same safeguards an average user would encounter.\n\n“Kimi K3 is very good at following a goal by any means necessary and also doesn't have the guardrails to prevent it from cheating or escaping the sandbox,” says Paul Kassianik, a researcher at Frontier Security.\n\nKassianik and Singer both say that Kimi and other open-weight models are also excellent tools for cybersecurity defense. (Hugging Face ultimately used an unnamed AI model from China to defend itself against the OpenAI agent hack.) Their company has developed [benchmarks](https://evals.frontier.security) that measure a model’s capacity to find vulnerabilities in software and networks, which show that Kimi excels at these tasks.\n\nFrontier Security says the sandbox it used was the default included in AISI's Inspect framework for testing AI systems.\n\n“These claims are inaccurate and irresponsible. Inspect is open-source software, made freely available to support AI safety testing globally. Users are responsible for configuring the tool to suit their needs, and we have published detailed guidance on how to do so,\" an AISI spokesperson told WIRED. \"The company has offered no evidence or wider detail offered to support the claims made. The issues they highlight result from how they chose to configure the tool.\"\n\nIn response, Frontier told WIRED that it provided details of the incident privately to AISI, and that it used the tool's default configuration, and did not modify it. AISI did not respond to WIRED's follow-up questions.\n\nSome cybersecurity experts say the issue discovered by Frontier Security reinforces how important it is to configure the environments that frontier AI models are placed in carefully.\n\n“It's not surprising at all,” says Matt Fredrikson, CEO of [Gray Swan](https://www.grayswan.ai/), another cybersecurity startup, and associate professor at Carnegie Mellon University. “As a general phenomenon, if you give one of these models an objective, and if you're not very explicit, like walls you're putting around it, it'll find a way to get the answer.”\n\nFredrikson says this means that people using AI models as agents, including in [tools like OpenClaw](https://www.wired.com/story/malevolent-ai-agent-openclaw-clawdbot/), which use AI to automate a wide range of useful chores, could find their systems misbehaving if they aren’t careful. “It is a cautionary tale,” he says.\n\n*Update 08/07/26 6:20pm ET: This story has been updated to include comments from AISI and Frontier Security.*", "url": "https://wpnews.pro/news/one-of-chinas-most-powerful-ai-models-has-also-escaped-containment", "canonical_source": "https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape-sandbox/", "published_at": "2026-08-07 01:16:55+00:00", "updated_at": "2026-08-09 12:11:54.749947+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-agents"], "entities": ["Frontier Security", "Moonshot AI", "Kimi K3", "OpenAI", "Anthropic", "Hugging Face", "UK AI Security Institute", "Yaron Singer"], "alternates": {"html": "https://wpnews.pro/news/one-of-chinas-most-powerful-ai-models-has-also-escaped-containment", "markdown": "https://wpnews.pro/news/one-of-chinas-most-powerful-ai-models-has-also-escaped-containment.md", "text": "https://wpnews.pro/news/one-of-chinas-most-powerful-ai-models-has-also-escaped-containment.txt", "jsonld": "https://wpnews.pro/news/one-of-chinas-most-powerful-ai-models-has-also-escaped-containment.jsonld"}}