The AI boom has been fueled by the idea that all data is fair game, potential fodder for the training of new models. Little wonder then that ownership and intellectual property theft have become such sticky issues within the tech industry.
On Thursday, Chinese robotics startup JoyIn—developer of a model called Aether—accused OpenAI of distilling its AI systems and stealing its website’s cosmic design aesthetic for the launch and branding of GPT-6 Astra. In an open letter written in Chinese, JoyIn CEO Guo Renjie said his company had started the process of filing a lawsuit, according to a translation from CNBC. (Gizmodo has not independently verified the report.)
Distillation is a common industry practice in which developers use a larger, more advanced model as a sort of “teacher” to help train a smaller “student” model. But it’s become a controversial subject in recent months, as major American AI companies, including OpenAI and Anthropic, have repeatedly claimed that competitors in China have been using their proprietary models to conduct large-scale, illicit distillation efforts in order to gain a competitive edge.
In a statement published Tuesday, the U.S. Cybersecurity & Infrastructure Security Agency (CISA) said that some of the biggest Chinese AI firms—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.ai—“are conducting systematic extraction of proprietary functionalities and capabilities of U.S. AI companies’ models through industrial-scale knowledge distillation campaigns that form the core—not merely a supplement—of their AI development strategy.” All this was likely being carried out with the knowledge and blessing of the Chinese government, according to CISA. Anthropic also said in a Thursday report that Moonshot, developer of an AI model called Kimi, had been carrying out a mechanical Turk-style trick to make its customers think they were getting responses from Kimi, when in fact they’d been generated by Claude.
Guo’s letter to OpenAI is the first known case of a Chinese developer accusing an American company of distillation. It’s far from the first time that OpenAI has been accused of theft, though. The company has had to fend off a litany of lawsuits since the launch of ChatGPT a little under four years ago, filed by artists and news publishers arguing that their copyrighted work was used illegally to train new AI models. (OpenAI has repeatedly defended itself by invoking the so-called “fair use” standard of U.S. copyright law.) In July, Apple sued the company, claiming it had stolen company secrets in order to get its newborn hardware business off the ground. (Again, OpenAI has denied any wrongdoing.) Earlier this week, a NYU mathematics professor questioned whether his work on Codex had been surreptitiously used to train the model that OpenAI used to solve the Navier-Stokes “Millennium Problem.” (OpenAI told the Australian Broadcasting Corporation it “can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”)
The irony of AI companies calling foul on “illicit” distillation when much of their business model depends on the nonconsensual scraping of online content seems to be lost on OpenAI and Anthropic, both of which are preparing for what are expected to be record-breaking IPOs. At least for the time being, the line between “fair use” and theft within the unregulated AI industry is as hazy as the inner-workings of models themselves.