The same trick a startup uses to build a cheap chatbot is now the center of a national security fight between Washington and Beijing.
On July 22, a senior White House technology official, Michael Kratsios, accused the Chinese startup Moonshot AI of building its new Kimi K3 model through what he called the "industrial distillation" of Anthropic's Fable model. The claim went further than a business dispute. It framed a common machine learning technique as an act of theft against the United States. Moonshot launched Kimi K3 on July 16, and within days a screenshot went viral showing the model identifying itself as "Claude, an AI assistant made by Anthropic." Neither Anthropic nor the White House has published direct evidence tying that behavior to Fable, and full verification wasn't possible until K3's model weights went public on July 27.
Distillation itself is not new, and it is not illegal. It's a training method where a smaller "student" model learns by studying the outputs of a larger, more expensive "teacher" model, absorbing its patterns without needing the same mountain of compute to get there. Startups use it constantly to cut costs. A small team without the budget for a frontier-scale training run can query a model like GPT-5 or Claude thousands of times, then train a cheaper model to mimic those answers. That's exactly what made DeepSeek's R1 launch in early 2025 so disruptive, and it's why US labs now treat the technique as a competitive threat rather than a neutral tool.
The accusations didn't start with Moonshot. On February 12, OpenAI sent a memo to the House Select Committee on China accusing DeepSeek of running an "ongoing effort to free-ride" on US frontier models, saying accounts tied to DeepSeek employees built code to query OpenAI's models through obfuscated third-party routers designed to mask their source.
Eleven days later, Anthropic went further. On February 23, the company published a report identifying a coordinated campaign by DeepSeek, Moonshot, and MiniMax to extract Claude's capabilities through more than 16 million exchanges routed through roughly 24,000 fraudulent accounts, in violation of Anthropic's terms of service. According to Anthropic, MiniMax generated the bulk of the traffic, over 13 million exchanges. Moonshot ran about 3.4 million, focused on agentic reasoning and computer-use workflows. DeepSeek's share was the smallest, around 150,000 exchanges, aimed at chain-of-thought reasoning and what Anthropic's report described as "censorship-safe alternatives" to politically sensitive queries.
The pattern didn't stop with the original teacher models. Zhipu AI released its GLM-5 model in 2026 claiming frontier-level performance, and its own technical report showed heavy reliance on DeepSeek's architecture, the same lab both OpenAI and Anthropic say distilled US models. Capabilities that started inside a US frontier lab can pass through one Chinese model and land in the next one downstream, several steps removed from anything an export license was ever meant to cover.
Why Washington can't just ban the chips anymore #
For years, the US strategy for slowing China's AI progress ran through hardware: restrict advanced chips, restrict the equipment that makes them, and assume compute scarcity does the rest. Distillation breaks that assumption, because a student model can approach a teacher's performance without anywhere near the teacher's training budget. You don't need a fleet of GPUs to copy behavior, you need access to the outputs. So the policy response has started shifting to knowledge itself, not just silicon. The FY2026 National Defense Authorization Act, in Section 1532, bans the Defense Department and its contractors from using "covered AI" tied to DeepSeek or its parent High-Flyer, or any entity in which High-Flyer holds a stake of 20% or more, with removal required within 30 days of enactment and only narrow waivers allowed for research or counterterrorism work. Separately, US officials believe DeepSeek trained a recent model using Nvidia's restricted Blackwell chips, a claim that, if confirmed, would mark a direct export control violation rather than a policy gray area.
None of this makes distillation illegal for the startups using it every day to keep costs down. It does mean the terms of service you agreed to when you started querying a frontier model's API are becoming a live enforcement question, not fine print. If you're building on top of someone else's model outputs at scale, expect the access restrictions, the rate limits, and the legal language around "derivative use" to tighten well before any new export rule ever reaches Congress.
Also read: Optical Networking Stocks Lumentum Ciena and Corning Are Beating the AI Trade • Apple Tests Banned Chinese Memory Chips Days Before Senate Deadline • Moody's warns AI rush leaves banks dependent on a handful of tech giants