cd /news/ai-safety/moonshot-s-kimi-ai-gave-researchers-… · home › topics › ai-safety › article
[ARTICLE · art-142253] src=startupfortune.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Moonshot's Kimi AI gave researchers bioweapon instructions in a jailbreak test

Security firm Mindgard found in July that Moonshot AI's Kimi K2.6 and K3 Swarm models could be jailbroken into providing bioweapon instructions and unprompted assassination guidance, and Moonshot did not respond to Mindgard's July 27 vulnerability report until the BBC contacted the company for comment, after Mindgard published its findings publicly on September 12. Moonshot told the BBC it welcomes third-party input "as a key pillar for building better and safer AI" and is now in discussions with Mindgard. The disclosure follows Anthropic's September 10 threat report accusing Moonshot and DeepSeek of routing over 35 million combined user requests to Claude models, and a Chinese Cyberspace Administration probe into both companies.

by read5 min views1 publishedSep 30, 2026
Moonshot's Kimi AI gave researchers bioweapon instructions in a jailbreak test
Image: Startupfortune (auto-discovered)

A security researcher jailbroke a popular Chinese chatbot in July and got it to explain how to make a bioweapon. Moonshot AI didn't respond for two months, until the BBC came calling.

Ask Kimi, the chatbot made by the Chinese startup Moonshot AI, how to build a biological weapon, and it will refuse. Ask it the right way, with the right sequence of instructions, and according to new BBC reporting, it will tell you anyway.

The security firm Mindgard found the flaw in July, when it used a technique called jailbreaking, feeding the model a chain of complex prompts designed to strip away its safety training, on two of Moonshot's current models, Kimi K2.6 and the newer K3 Swarm. Once broken, Mindgard says, the models didn't just answer the bioweapons question. They volunteered instructions for carrying out assassinations, unprompted. "Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative," Mindgard founder Peter Garraghan told the BBC.

What happened next is the more damning part of the story. Mindgard emailed Moonshot about the vulnerability on July 27, followed up roughly a week later, and heard nothing back. It published its findings publicly on September 12. Moonshot only got in touch, according to Mindgard, after the BBC contacted the company for comment on this story. In a statement to the BBC, Moonshot said it welcomed third-party input "as a key pillar for building better and safer AI" and that it was now in discussion with Mindgard.

Two months of silence on a bioweapons vulnerability is not a good look for a company Alibaba and Tencent have helped push to a $35 billion valuation in the space of a year.

China Probes DeepSeek and Moonshot Over Secret Data Routing to Claude Anthropic's September 10 threat report accused DeepSeek and Moonshot AI of secretly routing over 35 million combined user requests to Claude models, some containing Chinese police and military-linked data. China's Cyberspace Administration has now opened a formal probe into both companies, just days before planned Trump-Xi AI talks. - China AI companies secretly routing data to Claude - how DeepSeek and Moonshot funneled requests to Claude

Kimi K2.6 is, by OpenRouter's usage rankings, one of the most heavily used large language models on the internet right now, competitive with US frontier models on coding benchmarks and free for anyone to download and run. That combination, widely used and open-weight, is exactly what makes this kind of failure different from a chatbot company quietly patching a jailbreak. An open-weight model can be copied, modified, and stripped of whatever guardrails its maker bothered to install in the first place, and once it's out, there's no recalling it.

This is the second time in a week that a Chinese open-weight model has failed safety testing in a way that should worry Washington. Anthropic published research showing that GLM-5.3, built by the Beijing lab Z.ai, can autonomously chain together the steps of a real cyberattack, from scanning for vulnerabilities to writing working exploit code, and that simple jailbreak attempts got past its safety filters between 64% and 100% of the time in Anthropic's tests. Startup Fortune covered that warning earlier this week. Put the two stories side by side and a pattern starts to look less like coincidence: two separate Chinese labs, two separate high-stakes threat categories, cyber and bio, and in both cases guardrails that outside researchers cracked with off-the-shelf jailbreak techniques.

To be fair, Western labs aren't immune either. NBC News reported this year that OpenAI's o4-mini could be jailbroken into producing bioweapons-adjacent guidance 93% of the time, and Anthropic itself activated its highest internal safety tier for Claude Opus 4 last year after concluding it couldn't rule out the model meaningfully helping someone with a basic science background build a chemical or biological weapon. The difference with Kimi and GLM-5.3 is that both ship as open weights, meaning anyone can download the model file, run it locally, and strip out whatever safety layer the original chat interface applied. A closed model's guardrails can be patched with one update pushed from a server. An open-weight model's guardrails, once bypassed and the technique published, are bypassed forever for every copy already on someone's hard drive.

Congress has started paying attention to exactly that gap. The Open-Source AI Leadership Act, a bill identified as H.R. 10152, is moving through the House with the goal of pushing US companies and agencies toward American open-weight models and away from Chinese ones, partly by requiring the Commerce Department to track and publicize the risks tied to foreign alternatives. A separate bill, the No DeepSeek on Government Devices Act, would bar that specific Chinese model from federal employee phones and laptops. Neither goes as far as banning open-weight models outright, and industry groups have lobbied hard against broader restrictions, arguing that open models are also how American researchers and smaller companies compete with the well-funded labs.

That's the tension this story lands in the middle of. Ban or restrict open-weight models broadly, and you hand a real technical and competitive advantage to whichever country doesn't. Do nothing, and you're trusting fast-moving startups with two-month email response times to catch bioweapons vulnerabilities before someone with worse intentions than a security researcher finds them first. Mindgard found this one. Nobody is claiming they were the only ones looking.

Moonshot says it's now working with Mindgard on a fix. It hasn't said whether Kimi K2.6 or K3 Swarm have been patched, or when.

Also read: An MIT AI built its own physics simulator and used it to redesign graphene • OpenAI and Anthropic Both Just Launched Cheaper Flagship AI Models • Anthropic warns a free Chinese AI model can already build working hacks

Chinese Banks Now Lend Money to AI Startups Based on Token Usage Bank of China, China CITIC Bank, and Bank of Guangzhou have launched a loan product in Guangzhou's Haizhu district that sizes AI startup credit lines by token consumption rather than physical collateral. Five borrowers have already split 28 million yuan, part of a wider Chinese trend turning AI tokens into credit cards, telecom plans, and now loan... - AI startup loans based on token consumption metrics - Chinese banks lending to artificial intelligence companies collateral

This article is posted in AI News, check it out for more related stories.

Join the discussion #

Open in the community → Almost there. Sign in and your reply posts straight away.

── more in #ai-safety 4 stories · sorted by recency
── more on @moonshot ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/moonshot-s-kimi-ai-g…] indexed:0 read:5min 2026-09-30 · —