{"slug": "chinas-kimi-k3-significantly-below-us-rivals-in-hacking-power-study-shows", "title": "China’s Kimi K3 ‘significantly below’ US rivals in hacking power, study shows", "summary": "Chinese unicorn Moonshot AI's Kimi K3 model scored 32.2% on the ExploitBench benchmark, far below the 76.2% average of top unnamed US models, according to a joint study by the UK Artificial Intelligence Security Institute and the US Centre for AI Standards and Innovation. The model failed to achieve arbitrary code execution on any of 41 tasks, while leading US models succeeded on 20 tasks, challenging concerns about the rapid rise of Chinese open-source AI.", "body_md": "# China’s Kimi K3 ‘significantly below’ US rivals in hacking power, study shows\n\nKimi K3’s overall score of 32.2 per cent compares with top, unnamed US models that average 76.2 per cent\n\nChinese unicorn Moonshot AI’s Kimi K3 model trails far behind top American rivals in its ability to launch cyberattacks, joint British-US government research shows, challenging Washington’s brewing anxiety over the rapid rise of Chinese open-source artificial intelligence.\n\nThe model, currently considered China’s most powerful large language model, performs “significantly below the most recent frontier cyber-capable models”, according to a report published on Thursday by the UK Artificial Intelligence Security Institute (AISI) and the US Centre for AI Standards and Innovation (CAISI).\n\nThe AISI is a research arm under the UK Department for Science, Innovation and Technology, while the CAISI operates within the US Department of Commerce’s National Institute of Standards and Technology.\n\nTo evaluate capabilities, the two government bodies put Kimi K3 through ExploitBench, a public benchmark assessing an AI’s ability to develop exploits for cybersecurity vulnerabilities.\n\nKimi K3 achieved an overall score of 32.2 per cent, outperforming domestic rival Zhipu AI’s GLM-5.2 at 24.4 per cent, but lagging well behind top, unnamed US models that averaged 76.2 per cent.\n\nNotably, Kimi K3 failed to achieve arbitrary code execution – the highest-level exploit granting full control of a target system – across all 41 ExploitBench tasks, whereas leading US models achieved it on 20 tasks.", "url": "https://wpnews.pro/news/chinas-kimi-k3-significantly-below-us-rivals-in-hacking-power-study-shows", "canonical_source": "https://www.scmp.com/tech/tech-war/article/3361711/chinas-kimi-k3-significantly-below-us-rivals-hacking-power-uk-us-study-shows?utm_source=rss_feed", "published_at": "2026-07-24 06:30:46+00:00", "updated_at": "2026-07-24 06:46:46.832771+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-policy"], "entities": ["Moonshot AI", "Kimi K3", "UK Artificial Intelligence Security Institute", "US Centre for AI Standards and Innovation", "Zhipu AI", "GLM-5.2", "ExploitBench"], "alternates": {"html": "https://wpnews.pro/news/chinas-kimi-k3-significantly-below-us-rivals-in-hacking-power-study-shows", "markdown": "https://wpnews.pro/news/chinas-kimi-k3-significantly-below-us-rivals-in-hacking-power-study-shows.md", "text": "https://wpnews.pro/news/chinas-kimi-k3-significantly-below-us-rivals-in-hacking-power-study-shows.txt", "jsonld": "https://wpnews.pro/news/chinas-kimi-k3-significantly-below-us-rivals-in-hacking-power-study-shows.jsonld"}}