cd /news/artificial-intelligence/chinas-kimi-k3-significantly-below-u… · home topics artificial-intelligence article
[ARTICLE · art-71567] src=scmp.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

China’s Kimi K3 ‘significantly below’ US rivals in hacking power, study shows

Chinese unicorn Moonshot AI's Kimi K3 model scored 32.2% on the ExploitBench benchmark, far below the 76.2% average of top unnamed US models, according to a joint study by the UK Artificial Intelligence Security Institute and the US Centre for AI Standards and Innovation. The model failed to achieve arbitrary code execution on any of 41 tasks, while leading US models succeeded on 20 tasks, challenging concerns about the rapid rise of Chinese open-source AI.

read1 min views1 publishedJul 24, 2026
China’s Kimi K3 ‘significantly below’ US rivals in hacking power, study shows
Image: Scmp (auto-discovered)

Kimi K3’s overall score of 32.2 per cent compares with top, unnamed US models that average 76.2 per cent

Chinese unicorn Moonshot AI’s Kimi K3 model trails far behind top American rivals in its ability to launch cyberattacks, joint British-US government research shows, challenging Washington’s brewing anxiety over the rapid rise of Chinese open-source artificial intelligence.

The model, currently considered China’s most powerful large language model, performs “significantly below the most recent frontier cyber-capable models”, according to a report published on Thursday by the UK Artificial Intelligence Security Institute (AISI) and the US Centre for AI Standards and Innovation (CAISI).

The AISI is a research arm under the UK Department for Science, Innovation and Technology, while the CAISI operates within the US Department of Commerce’s National Institute of Standards and Technology.

To evaluate capabilities, the two government bodies put Kimi K3 through ExploitBench, a public benchmark assessing an AI’s ability to develop exploits for cybersecurity vulnerabilities.

Kimi K3 achieved an overall score of 32.2 per cent, outperforming domestic rival Zhipu AI’s GLM-5.2 at 24.4 per cent, but lagging well behind top, unnamed US models that averaged 76.2 per cent.

Notably, Kimi K3 failed to achieve arbitrary code execution – the highest-level exploit granting full control of a target system – across all 41 ExploitBench tasks, whereas leading US models achieved it on 20 tasks.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @moonshot ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/chinas-kimi-k3-signi…] indexed:0 read:1min 2026-07-24 ·