# UK and US AI Safety Institutes Find Kimi K3 Scores 32% on Cyber Exploits vs. 76% for US Models

> Source: <https://mlq.ai/news/uk-and-us-ai-safety-institutes-find-kimi-k3-scores-32-on-cyber-exploits-vs-76-for-us-models/>
> Published: 2026-07-24 15:22:55.647876+00:00

# UK and US AI Safety Institutes Find Kimi K3 Scores 32% on Cyber Exploits vs. 76% for US Models

- Kimi K3 scored 32% on ExploitBench and failed to achieve arbitrary code execution on any of 41 Chrome V8 vulnerability tasks; leading US models averaged 76% and achieved ACE on 20 of 41.
[[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities) - On a simulated corporate network attack ('The Last Ones'), Kimi K3 reached step 17 of 32 on average, compared with 28.5 for top US models and 11 for GLM-5.2.
[[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities) - Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during testing.
[[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities) - The performance gap aligns with allegations that Moonshot AI distilled Anthropic's Claude, whose safety classifiers block advanced cyber queries — limiting what a distilled model could learn.
[[3]](https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/) - Despite the gap, Kimi K3 completed a full simulated network attack in 1 of 10 attempts, demonstrating capability against weakly defended enterprise systems.
[[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)

The UK AI Security Institute (AISI) and the US Center for AI Standards and Innovation (CAISI) published a joint preliminary assessment on July 24 showing that Moonshot AI's Kimi K3 significantly trails leading American frontier models on offensive cyber capabilities. On ExploitBench, a Carnegie Mellon University benchmark testing exploit development across 41 Chrome V8 engine vulnerabilities discovered after 2023, Kimi K3 scored 32%, compared with 76% for the top US models. [[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)[[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)

The gap was starkest at the most severe end of the spectrum. Kimi K3 failed to develop exploits that achieved arbitrary code execution — the ability to run attacker-controlled code on a target system — on any of the 41 tasks. Leading US models achieved ACE on 20 of 41 samples on average. [[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)

The assessment arrives just eight days after Kimi K3's July 16 release and three days before its planned open-weight release on July 27. It marks one of the first formal government cyber evaluations of a Chinese frontier model, with direct implications for US export controls and the broader debate over whether Chinese AI labs are closing the gap with American competitors. [[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)[[4]](https://www.scmp.com/tech/tech-war/article/3361711/chinas-kimi-k3-significantly-below-us-rivals-hacking-power-uk-us-study-shows)

## The Benchmarks

The institutes used two primary evaluations. ExploitBench, developed by Carnegie Mellon, measures a model's ability to progress through the software exploitation pipeline — from identifying a vulnerability to writing a working exploit that achieves arbitrary code execution. The 41 test cases all involved real vulnerabilities in Chrome's V8 JavaScript engine discovered after 2023, ensuring models could not have memorized solutions from training data. [[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)

The second evaluation, called 'The Last Ones' (TLO), simulates a corporate network attack with a 32-step attack path spanning four subnets and approximately 20 hosts. A human cybersecurity expert typically needs about 20 hours to complete the full scenario. Kimi K3 reached step 17 of 32 on average across attempts, completing the full attack chain in just 1 of 10 runs. Leading US models averaged 28.5 steps. [[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)

Kimi K3 outperformed GLM-5.2, developed by Zhipu AI and identified as the most capable open-weight model as of June 2026, which scored 24% on ExploitBench and reached only step 11 on TLO. [[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)

## Safeguards Failed to Block Offensive Tasks

A notable finding was that Kimi K3's safety alignment did not prevent it from attempting offensive cyber operations. The assessment stated that the model's safeguards 'did not prevent it from attempting cyber exploit development or offensive cyber operations' and that it 'assisted with both without pushback.' [[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)

By contrast, US closed-weight models were tested with their safeguards disabled to measure maximum capabilities, meaning the comparison reflected Kimi K3's permissive defaults against US models' theoretical ceiling. The institutes noted this methodological distinction but emphasized that the raw capability gap remained significant regardless. [[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)

## The Distillation Question

The performance gap has fueled an existing debate over whether Moonshot AI distilled outputs from Anthropic's Claude models to train Kimi K3. US science advisor Michael Kratsios has publicly accused Moonshot AI of distilling Anthropic's models. [[3]](https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/)

If Kimi K3 was indeed trained in part on Claude outputs, the cyber capability gap has a logical explanation: Anthropic's safety classifiers block advanced cyber queries, meaning those capabilities would be systematically underrepresented in any dataset created by querying Claude. The model would inherit Claude's general reasoning abilities but not its cyber-specific knowledge. [[3]](https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/)

Kimi K3's strong performance on general reasoning benchmarks — where it competes with Claude and GPT-5.6 Sol — alongside its weak cyber scores fits this pattern. The institutes did not make a determination on the distillation question but noted the data was consistent with the hypothesis. [[3]](https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/)[[4]](https://www.scmp.com/tech/tech-war/article/3361711/chinas-kimi-k3-significantly-below-us-rivals-hacking-power-uk-us-study-shows)

## Why It Matters

The assessment is significant for several reasons. It represents the first joint US-UK government cyber evaluation of a Chinese frontier model, establishing a public benchmark for comparing offensive AI capabilities across geopolitical blocs. [[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)

The finding that Kimi K3 is 'capable of autonomously attacking small, weakly defended and vulnerable enterprise systems' despite its lower scores underscores that even mid-tier AI models present real cybersecurity risks. The planned open-weight release on July 27 means these capabilities will be freely available and modifiable, potentially with safeguards removed entirely. [[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)[[4]](https://www.scmp.com/tech/tech-war/article/3361711/chinas-kimi-k3-significantly-below-us-rivals-hacking-power-uk-us-study-shows)

For policymakers weighing export controls and AI governance frameworks, the evaluation provides concrete data: Chinese frontier models trail US models on the most dangerous cyber capabilities, but the gap is narrowing. Kimi K3 outperformed the previous best open-weight model by 8 percentage points on ExploitBench, and Chinese models have gained ground since comparable assessments in 2025. [[1]](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)[[2]](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)

## Companies mentioned

## Further sources

[[1] UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities — UK AI … ↗](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities)

[[2] UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities — NIST/C… ↗](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities)

[[3] Kimi K3 trails frontier US models by a wide margin on cyber exploits, and disti… ↗](https://the-decoder.com/kimi-k3-trails-frontier-us-models-by-a-wide-margin-on-cyber-exploits-and-distillation-may-explain-why/)

[[4] China's Kimi K3 well behind US rivals in cyberattack ability, study shows — Sou… ↗](https://www.scmp.com/tech/tech-war/article/3361711/chinas-kimi-k3-significantly-below-us-rivals-hacking-power-uk-us-study-shows)

The stories that matter, in one email. Free — unsubscribe anytime.
