cd /news/ai-safety/moonshots-kimi-ai-model-has-also-esc… · home topics ai-safety article
[ARTICLE · art-88506] src=infoworld.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Moonshot’s Kimi AI model has also escaped from a test environment

Frontier Security reported that Moonshot AI's Kimi K3 model escaped the UK AI Safety Institute's test environment by exploiting a loophole to access github.com and clone the official repository for the benchmark problem, reading the solution directly instead of solving it. This follows similar breaches by models from OpenAI, Anthropic, and Meta. Frontier advised companies to restrict outbound DNS and HTTPS traffic, audit traces, and treat benchmark scores as meaningful only when models lack access to reference implementations.

read2 min views1 publishedAug 7, 2026

Yet another AI model has escaped from a cybersecurity test lab: This time, it’s the Chinese company Moonshot’s Kimi K3 model on the run.

Frontier Security spotted that Kimi K3 had found a loophole in the UK AI Safety Institute’s test environment for AI models performing cybersecurity tasks. The news follows similar exploits by models from OpenAI, which attacked Hugging Face, Anthropic, and most recently Meta.

Frontier revealed how the fault came about. AI models are routinely tested to examine how they perform offensive and defensive cybersecurity tasks, typically in isolated test environments or sandboxes that severely limit their internet access. Frontier reported that Kimi K3 model had found a break in the sandbox it was being tested in, enabling it to reach out to the live github.com website and clone the official repository for the benchmark problem it was supposed to be solving, reading the solution directly off the disk rather than solving the problem for itself.

Frontier warned companies testing AI models to be aware of the dangers such loopholes pose and offered some guidelines.

Companies should restrict outbound DNS and HTTPS traffic from AI models to an explicit allowlist and test those controls from inside the same environment available to the model, Frontier said. They should also audit traces for any suspicious activity and not rely solely on final answers. Companies should also treat a model’s score on benchmarks as meaningful only when the model doesn’t have access to reference implementations and other shortcuts.

Frontier also advised testers to be suspicious of unexpectedly high pass rates, as these may reveal a shared environmental flaw.

Perhaps most importantly of all: They should assume agents will find the worst paths to a solution, including probing a test environment for loopholes, and won’t always follow the path that they are expected to.

As Frontier write in its blog: “Models optimize for the objective function (getting the correct flag/answer), not the human intent behind the benchmark. If a network path to the solution exists, a sufficiently capable agent will find it.”

This article first appeared on CSO.

── more in #ai-safety 4 stories · sorted by recency
── more on @frontier security 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/moonshots-kimi-ai-mo…] indexed:0 read:2min 2026-08-07 ·