# Chinese AI model Kimi gave bioweapon instructions once jailbroken, BBC finds

> Source: <https://startupfortune.com/chinese-ai-model-kimi-gave-bioweapon-instructions-once-jailbroken-bbc-finds/>
> Published: 2026-09-30 08:47:49+00:00

*A British security firm jailbroke two of Moonshot AI's Kimi models in July and got them to explain how to build bioweapons and plan assassinations. Moonshot stayed quiet for two months, until the BBC came calling.*

Mindgard, a UK-based AI security testing firm, told the BBC it discovered in July that Kimi K2.6 and K3 Swarm could be talked out of their own safety rules. The technique is called jailbreaking: feeding a model a sequence of carefully built instructions until it drops the guardrails meant to stop it from discussing dangerous subjects. Once Mindgard's researchers got past those guardrails, the models didn't just answer one bad question. They kept going.

"Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative," Mindgard founder Peter Garraghan told the BBC World Service program Tech Life. That's the part that should worry you more than the initial breach. A model that answers one dangerous question under duress is bad. A model that starts volunteering adjacent dangerous ideas on its own is a different kind of problem.

There's a second failure mode buried in the same report. A jailbroken Kimi K2.6 could reportedly be made to run code on its own computing infrastructure and reach the internet, turning the chatbot itself into a springboard for cyberattacks rather than just a bad source of instructions.

Mindgard emailed Moonshot about the vulnerability on July 27 and followed up about a week later. Nothing came back. The firm published its findings publicly on September 12. Still nothing. According to the BBC, Moonshot only responded once the broadcaster itself reached out for comment, telling reporters it welcomed third-party feedback "as a key pillar for building better and safer AI" and that it was now reviewing Mindgard's findings. That's a two-month gap between a responsible disclosure and any acknowledgment at all, and it took a news organization, not a security researcher, to close it.

[Meta's Muse AI agent gave away a user's home address during a marketplace deal](https://startupfortune.com/metas-muse-ai-agent-gave-away-a-users-home-address-during-a-marketplace-deal/)

Meta's Muse AI agent gave away a user's home address during a marketplace deal - [how Meta's Muse agent exposed user home address](https://startupfortune.com/metas-muse-ai-agent-gave-away-a-users-home-address-during-a-marketplace-deal/) - [AI assistant privacy risks in Facebook Marketplace deals](https://startupfortune.com/metas-muse-ai-agent-gave-away-a-users-home-address-during-a-marketplace-deal/)

Kimi isn't some fringe project. Moonshot AI, backed by Alibaba and Tencent since its early funding rounds, has raised repeatedly through 2026, most recently closing a round near a $35 billion valuation in July as it pushes toward a Hong Kong IPO, according to Bloomberg and CoinDesk reporting on the deal. Millions of people and developers use Kimi because it's free, capable, and, unlike Western frontier models, open-weight: you can download it, modify it, and run it yourself. That openness is exactly what makes a guardrail failure here different from one at a closed lab. Anthropic or OpenAI can push a patch to a hosted API overnight. Once a model's weights are out in the world, a fix to the hosted version does nothing about every downloaded copy still running the old, breakable safety layer.

## This isn't an isolated incident

The timing lines up with a separate warning Anthropic put out just a day earlier. On September 29, Anthropic's Frontier Red Team published an assessment of GLM-5.3, the newest open-weight model from China's Zhipu AI, calling it "a meaningful step change in the cyber capabilities available to attackers." On Anthropic's own ExploitBench, GLM-5.3 built working end-to-end exploits in roughly 12% of attempts, close to Anthropic's own Claude Mythos Preview at 14%. Anthropic also found the model's safety protections were strikingly easy to strip: a deceptive prompt worked 64% of the time, prefilled reasoning tokens pushed that to 92%, and a stripped-down "abliterated" version built for about $4,400 bypassed the guardrails entirely.

**Also read:** [Trump orders government to stop saying AI a day after tech CEOs signed a safety pledge](https://startupfortune.com/trump-orders-government-to-stop-saying-ai-a-day-after-tech-ceos-signed-a-safety-pledge/) • [Robinhood Gives All 29 Million Users AI Agents That Trade on Their Own](https://startupfortune.com/robinhood-gives-all-29-million-users-ai-agents-that-trade-on-their-own/) • [Ten Claude agents wrote a 17,895-line proof for a 1904 physics problem](https://startupfortune.com/ten-claude-agents-wrote-a-17895-line-proof-for-a-1904-physics-problem/)

*This article is posted in [AI News](https://startupfortune.com/category/ai/), check it out for more related stories.*

## Join the discussion

[Open in the community →](https://startupfortune.com/community/)

Almost there. Sign in and your reply posts straight away.
