This report covers activity we disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation.
There’s a bit of Arson, Murder and Jaywalking there. One of these things, many would say, is not like the others.
I do not agree, especially given the details we will see later, and given that distillation enables the other six via, as the report says, ‘driving performance on nearly every task’ via transfering Claude’s cognitive skills, without transferring its safeguards.
Indeed, distillation is by far the most important threat in this report, and the part of the report that will have the most impact.
By exposing Chinese attempts at systematic fraudulent distillation of Claude, Anthropic has embarrassed and potentially antagonized the Chinese. This starts with ‘all the top Chinese labs made efforts to distill Claude, which we mitigated and stopped,’ which is already an issue that was also covered by a joint advisory from NSA/CISA/FBI two days before the full report.
The bigger issue is how the Chinese labs were trying to distill Claude. Distillation attempts require lots of realistic queries, so DeepSeek, Moonshot and Xiaomi each sent lots of user queries directly to Claude. At least Moonshot then gave the results back to its users as if these were Kimi outputs.
That’s going to be a problem.
Breaking Unrelated News
This is unrelated to today’s post but you need to know the basic facts now, so:
Yesterday I cautioned that the media and others were reading far too much into Trump’s statements. Alas, in keeping with ‘every day there is breaking news,’ we have now seen Trump say the things he had not yet said. The new statement is very different, resolving the ambiguity from yesterday morning.
He explicitly declared AI existential risk to be a ‘hoax’ on par with (his words) the ‘Russia hoax’ or climate change. This is very, very bad news. I fear he may have crossed a rhetorical Rubicon that will be difficult to step back from once he better understands the situation, and once future incidents change the game.
While we all process the implications, I urge everyone: Please do not make this any more partisan or personal than it already is. Please do not attack Republicans, or Trump. That will only make things worse. Emphasize helpful voices on all sides. That applies no matter what additional statements may come, and I will offer full coverage of that unfortunate situation later this week.
The classifiers are hella annoying sometimes, but Anthropic says they work.
In all cases, Claude Haiku, Sonnet, and Opus models were used; no malicious activity was found on Claude Fable or Mythos (which has a series of safeguards in place that greatly reduce its ability to perform harmful cyber tasks).
The one exception was an attempted distillation attack on Fable by Zhipu, but there were enhanced safeguards in place and Zhipu switched to going after Opus.
Bad Dudes Tend To Be Relatively Unsophisticated
If Bad Dudes were more often sophisticated and Good At Job, the world would look very different. Luckily, Bad Dudes are usually unsophisticated and Bad At Job.
(Everyone else is mostly similar, but less so and less reliably.)
AI, for both better and worse, turns Bad At Job into Good At Job, making it less relevant that the humans are Bad At Job via doing the things. Here are Anthropic’s big themes:
Sophisticated attacks no longer require sophisticated attackers.
AI’s role in cyber operations has become increasingly autonomous.
AI supply chain as target, loot, and attack compute.
As in, attackers go after sources of compute, and use that to keep going.
To use Claude as Bad Dude, you need the value of the loot and compute, and you also need the cover of newly compromised accounts, to stay ahead of Anthropic cracking down.
Basically everyone in the report is stealing their access, at minimum via getting around regional restrictions and buying endless subsidized subscriptions, on top of then doing Bad Dude things with the compute.
AI tradecraft is proliferating.
Like everything else, diffusion takes time.
You need fewer less sophisticated humans, who are less Good at Job, to Do Thing.
For most things, that’s great. This report is about the exceptions. You especially need less in order to adapt when someone defends against you, and to develop new methods.
Particular Bad Dudes
GTG-20006 is a Russian espionage operation linked to Midnight Blizzard. They had a standard set of cyber attack tools. They used AI to monitor how well their tools evaded detection, and iterated until security defenses did not detect their malware, then launched AI-automated attacks, including AI phishing operations from AI-registered domains.
Targets varied, including Ukrainian government, military and diplomatic staff, various government agencies and defense-industrial companies linked to Ukraine, and Ukraine’s drone supply chain. They took over WhatsApp accounts via headless browsers and targeted surveillance cameras. All of it was automated.
GTG-50014 were smash-and-grab opportunists associated with the ShinyHunters collective. They are doing the standard opportunistic things, but AI let them scale and cast a wide net looking for vulnerable systems, targeting bulk data theft and potential extortion. The whole thing is basically ‘vibe hacking,’ with ‘living off the land,’ having the AI look around and use whatever it finds, seeking any vulnerability at all, rather than having a plan.
GTG-10007 was a Chinese-speaking espionage operation, likely in Changsha in China’s Hunan province. Claude was used to automate workflows and form agent swarms, the same as any other coding task, except here it was vulnerability research, testing and exploit design. Roughly fifty organizations were targeted, and an education-technology company was compromised.
GTG-50020 is a Russian-speaking, financially-motivated actor, historically targeting hotel booking and financial technology platforms, that pivoted to attack the AI industry and steal API keys via prompt injecting sandboxes. They attacked 30 companies, but they never achieved their objective, which was access to a pre-release Claude model.
GTG-50029 was a French-speaking hacktivist who targeted European political and affiliated entities, an example of uplift for unsophisticated threat actors, who then managed to compromise 14 of 42 WordPress sites including via an undocumented race condition and steal users’ political opinions, a mass attack on privacy.
Influence Operations
Influence operations are getting larger and more sophisticated.
We’ve seen groups of actors use Claude to build networks of fake social media profiles and entire news sites, leveraging these platforms to publish deceptive content, while completely concealing the entities behind these operations.
It is remarkable how well social media has held up so far in the age of AI. Threat actors spin up hundreds of accounts on a regular basis, all providing social proof for each other, and we complain but the defenders are mostly winning.
This report details nine of those cases. They originated in Russia, Iran, Turkey, and across the Gulf, South Asia, Africa and Europe, and targeted audiences on six continents.
Building such campaigns builds a signature that Anthropic often detects. When Anthropic finds an operation in the planning stage, or discovers one afterwards, they ban the accounts and use this to strengthen their detection mechanisms. But of course such actors can always get new accounts and try again.
Trends listed:
Influence sold as a service.
AI as a newsdesk.
AI helped to build the apparatus as well as the content.
Complex tool use.
Laundering of attribution, sourcing and certainty.
Increased operational security.
Fake personas (and impersonation of real personas).
Targeting people and accountability mechanisms.
Influence operations often fail to reach a genuine audience.
The final takeaway is what I see from the outside. These campaigns are often remarkably ineffective, individually and in general, and much more boogeymen so far than actually influential. Hopefully that will last.
These operations mostly (but not entirely) target third world areas, where there is less competition, that is less sophisticated, in the media and influence ecosystems.
They list nine:
GTG-04001: Russian manipulation in the Central African Republic, via a heavily biased AI-generated news pipeline.
GTG-54002: Commercial ‘influence-as-a-service’ spanning six continents, traced to France. They used 70 fabricated news sites and 250 inauthentic Twitter accounts. Used Claude to write and rewrite news articles and tailor them.
GTG-84005: An election-manipulation platform targeting Malaysia. About 1,000 fake Twitter accounts, a fake news outlet and a series of fabricated dossiers, as paid political influence-as-service.
GTG-24015: Russian state-media editorial pipelines for editorial and news production for audiences in various countries. They targeted the elections in Moldova.
GTG-34001: Iranian state-aligned ICCO, Islamic Propaganda Office and Bina Observatory. They were laying the foundation for an explanatory jihad.
GTG-54006: Automated pro-Awami League self-described fake-news operation targeting rural Bangladesh.
GTG-84006: MEK/NCRI-aligned influence operation using AI to impersonate an activist and recruit inside Iran. They scraped posts on Telegram to try and assemble a profile of the particular activist.
GTG-54004: A domestic inauthentic behavior campaign in Kenya, basically astroturfing public sentiment.
GTG-84002: A UAE-directed influence operation targeting the Muslim Brotherhood, Sudan conflict and UN accountability mechanisms. This was supposed to be ‘a coordinated transatlantic and regional operation to dismantle the Muslim Brotherhood globally.’
Overall, I was not impressed. It does not seem like such folks are getting much done, at least not via Claude. There’s a thin line between a lot of this and Ordinary Politics. If this is as bad as it gets then this is very good news.
Surveillance Operations
These cases include threat actors from China, Iran, and West Africa, as well as the commercial “surveillance-for-hire” market, and range from operations carried out by a single individual to entire teams.
And they did it in violation of the Terms of Service. How dare they.
In many cases this was ‘analyze a ton of social media posts,’ including by the state to target dissidents and dissident groups, or groups likely to cause trouble. Most of the groups were Chinese or Iranian. In one case the Iranians targeted Jews. Sometimes malware was involved.
Are Iran and China the only places trying to do this level of AI surveillance? Or are they the only ones crazy enough to use Claude?
Anthropic highlights GTG-50027, a single independent consultant in Bamako, who used Claude as part of Mali’s state intelligence service (ANSE) to target roughly 25 million SIM cards. This one was not really disrupted. Further work was prevented, but the system had already been deployed and it remains deployed.
This is fundamentally not a solvable problem. Anthropic and other AI companies are in the business of selling intelligence, and there are also open models providing intelligence. There is no way to fully stop all forms of mass surveillance, domestically or otherwise, given the alternative options. Claude does not provide that big an advantage here over open models. If some people are out there to analyze public information, you can slow them down but they are going to succeed.
Another operation stood out: GTG-30005, due to its target: In another investigation, we identified and disrupted an Iran-nexus threat actor that used Claude to collect and analyze publicly accessible data to develop targeting recommendations against US naval forces in the region.
That goes hand in hand with the next section: Use of Claude in conventional weapons.
Conventional Weapons
This refers to software for such weapons, as well as targeting and control systems. There are six cases here: three in China, two in Russia and one in Yemen.
First there are the weapon systems development cases:
GTG-87001: A Yemen-based guided weapons engineering cell using Claude to develop guidance software. Claude as software engineer, except for weapons.
GTG-17001: A China-based operation drafting a fire control specification and acquisition documents for undersea warfare, within their defense contractor ecosystem. They pretended to be American to get Claude to write a proposal.
GTG-27005: A Russia-based likely freelance operation to engineer an autonomous military first-person-view kamikaze drone swarm. Again, the wrong software.
GTG-17002: A China-based operation to build targeting software for electronic warfare and air defense suppression.
It is a thin line between ordinary software development, and software that assists with conventional weapons. It is no surprise developers would attempt to use Claude. Indeed, presumably lots of Western weapons developers also use Claude.
GTG-27006: Russia-based operation to procure mixed military and civilian goods. Claude was used to buy things. Eventually it added up to obvious military use.
GTG-17003: China-based operation to collect open-source intelligence on directed-energy weapons and their supply chain.
Again, this happens to be a particular use we do not like, which is Not So Different.
Biological Misuse
How much is being attempted in practice? They have five cases.
Regional use controls were evaded, but the five case studies here look like scientists. They are doing dual use things, where it is not clear that harm was intended.
A grant application for gain-of-function research.
A research program engineering highly pathogenic mammal-adapted avian influenza.
Orthopoxvirus research, including help with related logistics.
Two cases of dual-use non-transmissible novel venoms and toxins.
The first two cases seem like clear cases of Do Not Want, even if the people involved thought they were helping. The other three are less clear. None of the cases were the kind of dangerous use of AI we worry about, where AI provides major uplift. AI was being used because AI is highly useful at ordinary tasks.
I call this a huge success story, assuming Anthropic isn’t hiding worse things. These aren’t even fully Bad Dudes, only misguided ones.
Scams and Fraud
Ah, good to get back to Ordinary Decent Crime. Except, you know, at scale.
They only offer one example, but it’s a fun one.
GTG-15001: Fake dating app network. Love it. Now we’re talking. To AIs. A full network of 20 fake dating apps, with more than 4,700 distinct AI personas that talked to over 25,000 unique individuals over two weeks in April, with swipe feeds that were 25% real people and 75% AIs. For some of the 25%, workers were hired to choose among three candidate AI replies, thus proving good game design, and paid per message. Claude ran the AI personas, because nothing but the best will do. I might be mangling the details a bit.
Messaging is metered, which is how they make money. That was a hint.
I mean, yeah, yeah, scam and fraud, poor fake users, very sad.
But this is clearly the best case study.
The weird part is this is the best case study, and here the only one. Where are all the others? Again this is a huge success story.
Illicit Fraudulent Distillation
This is the controversial one.
We define illicit distillation as an industrial-scale, covert campaign to extract a model’s capabilities and replicate them in another model without authorization. Illicit distillation is typically enabled by fraud: sophisticated networks of fake accounts created with stolen credit cards, login credentials, and API keys.
Other frontier labs have faced distillation attacks. OpenAI has called attention to this activity since early 2025.
If you want to argue that distillation using your own queries and the model outputs, or data otherwise acquired above board, is fine, I respectfully disagree with you for practical reasons and for the same reasons I support copyright and patent protections, but I understand why you would think that.
If you think that it is okay to then use a ‘Chain of Thought extractor’ as an exploit, as part of that effort, I disagree with you a lot more, and I think you’re clearly in the wrong, but I understand why some people think intellectual property is not real and you should be able to just steal it and are willing to cheer for Chinese companies to steal American IP.
If you think it is okay to do this by secretly rerouting sensitive customer queries to Claude, or using accounts created with stolen credit cards, login credentials and API keys to harvest data?
Then I’m sorry, you’re just flat out wrong, that is obviously not okay.
I mention those particular things because that is what is going on.
Over the last several months, unauthorized labs have developed increasingly sophisticated methods to circumvent our defenses and harvest the capabilities of US frontier models. These labs generally access Anthropic’s models by routing requests through proxy services, also known as “transfer stations.”
To circumvent our geographic restrictions and related controls, these proxy services create thousands of new accounts using false identities, fake or stolen credit cards, and stolen API keys.
…
Unauthorized labs also obtain transcripts of user exchanges with US frontier models by purchasing them from third-party resellers. These resellers include the operators of proxy services, which often save exchanges between users and US models without the knowledge or consent of those users.
Meanwhile, there is an ongoing war where Anthropic tries to block various prompts that attempt to extract the Chain of Thought, and the Chinese and other attackers try to develop new extraction methods.
Also, it gets worse:
These findings raise concerns about the misuse of user data by PRC AI labs. DeepSeek, Xiaomi, and Moonshot fed conversations between their own models and users into Claude.
These labs then used Claude’s responses as training data with which to distill Claude’s capabilities. Some of these exchanges included sensitive information, including from individual users, major multinational companies, and state-affiliated actors. Many of these exchanges were relayed from users of third-party model routing services commonly used by users in the United States and Europe. Those sessions contained names, email addresses, company data, and other sensitive data of hundreds of end users in at least a dozen languages. These practices are likely inconsistent with privacy laws and the labs’ own terms of service.
The report includes the example of asking Claude to analyze CCTV surveillance footage from hundreds of Chinese cameras, and exposing live Russian government credentials. It can get ugly out there.
Your response can be ‘well Anthropic is flat out lying about this’ but short of that there is no way to excuse the behaviors in question.
All of this looks like it is in direct violation of numerous laws, including Chinese laws.
In many of these cases, if nothing else, Article 39 of PIPL was definitely broken by DeepSeek, Moonshot and Xiaomi, and likely Criminal Law Art. 235a as well:
Accessible Law (PIPL): Where a personal information processor provides personal information of an individual to a party outside the territory of the People’s Republic of China, it shall inform the individual of such matters as the name of the overseas recipient, contact information, purpose, and method of processing, type of personal information and the way and procedure for the individual to exercise the rights prescribed herein against the overseas recipient, and shall obtain the individual’s separate consent.
That’s in addition to everyone breaking the Anti-Unfair Competition Law, of using data lawfully held by another business operator through “fraud, coercion, or circumventing or damaging technical or management measures” where this disrupts market competition order. Seems to apply here.
Oh, and all of this involved Ordinary Decent Fraud, as in Criminal Law Art. 196 and Art. 177a.
And then there’s Interim Measures for the Management of Generative AI Services, Article 7.
There are six particular cases.
Alibaba (Qwen) used a massive network of fake accounts.
Moonshot (Kimi) secretly sent user exchanges to Claude, via a massive network of fake accounts.
DeepSeek also secretly sent user exchanges to Claude.
Zhipu (GLM) did the extraction thing, and the fake account thing, and also recently went after the cyber capabilities of leading American models ahead of the release of GLM-5.3.
Xiaomi did the distillation thing, using saved user requests.
SenseTime, MiniMax and others form a third-party reseller ecosystem.
As in, all the cool kids are doing it. That’s how they are the Chinese cool kids.
How do the Chinese have the best open models?
GTG-16005: CoT distillation and AI R&D campaign by Alibaba (Qwen / Tongyi Lab). Alibaba’s CoT distillation pipeline injected a fixed prompt into each request that forced Claude to write out its reasoning traces inside inline text tags before providing its final answer. Those CoT transcripts were then saved and converted into data that could be used for supervised fine-tuning (SFT). These SFT transcripts were used to help train Alibaba’s Qwen models, and were used to distill Claude’s capabilities into Qwen 3.5, 3.6, and 3.7.
Alibaba’s illicit distillation campaign peaked at nearly 3 million exchanges per day launched from more than 3,500 fraudulent accounts.
… Scale of distillation attacks attributable to Alibaba between May and July 2026: over 151 million exchanges observed.
GTG-16002: Moonshot serves Claude instead of Kimi and collects exchanges for model training.
We discovered that Moonshot AI, the company that produces the Kimi family of models, silently forwarded customer requests to Claude, instead of processing them
using Kimi. Moonshot then displayed Claude’s responses to users. These users thought they were using a Kimi model, but received responses from Claude instead. In one instance, over a ten-day period, Moonshot relayed almost 300,000 customer requests to Anthropic, the vast majority of which were routed to Opus. Moonshot used a proxy service network of 5,380 fraudulent accounts, most of which appeared to be located in Singapore and Japan.
In addition to serving Claude’s responses to their customers, Moonshot also captured and saved at least a portion of these exchanges.
… When responding, Claude returns a reference to its raw thinking as a “thinking signature” instead of the raw thinking to mitigate the risk of unauthorized distillation.
… Scale of distillation attacks attributable to Moonshot between May and July 2026: over 23 million exchanges observed.
Lyman Stone 石來民: Okay so over the course of reading this I shifted from “wow Kimi is awful” to “WHY DID ANTHROPIC SPILL THE BEANS ON THIS INCREDIBLE INTELLIGENCE HACK”
No, Lyman. I get why you would say that, but we do not steal user data, even if the user was attempting to use Kimi and kind of deserves it.
GTG-16001: DeepSeek serves Claude instead of its own models and collects exchanges for model training.
Our investigation revealed that DeepSeek also deployed tactics similar to Moonshot’s. DeepSeek built a CoT extraction pipeline, relying on the same cross-session replay attack described above. DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers. Like GTG-16002, their customers were likely not made aware that their requests were being funneled to Claude.
… Scale of distillation attacks attributable to DeepSeek over 14 days in July 2026: over 12.1 million exchanges observed.
GTG-16006: Zhipu distillation, AI R&D and targeting cyber capabilities.
Zhipu, branded outside China as Z.ai, ran a chain-of-thought extraction pipeline against Claude, replaying captured Claude reasoning traces back through Claude to clean them for training its GLM models. Over just ten days, Zhipu launched a CoT extraction pipeline against Claude Opus 4.8 by rotating through 273 fraudulent accounts to evade our model restrictions.
… Zhipu also used Claude to improve its own post-training pipelines, using Claude to judge model outputs and clean and normalize reasoning transcripts harvested for distillation.
… More recently, ahead of the release of its GLM 5.3 model, we identified a campaign to target the cyber capabilities of leading US frontier models. Zhipu researchers used public vulnerability datasets to develop various capture-the-flag challenges.
Zhipu initially attempted to target the cyber capabilities of Anthropic’s Fable model. Fable—Anthropic’s top generally accessible model—has strengthened cyber safeguards, making it more difficult for would-be distillers to target Fable’s cyber capabilities. Zhipu eventually gave up trying to target Fable after Anthropic’s cyber safeguards degraded Zhipu’s attacks.
… Scale of distillation attacks attributable to Zhipu over 17 days in June and July 2026: over 3.4 million exchanges observed.
That’s right. They tried to forcibly distill Fable’s cyber capabilities in order to put them into an open model. Luckily, this did not work and they had to pivot to Opus 4.6 where the safeguards were weaker.
GTG-16008: Distillation campaign by Xiaomi
This one was only ~400k messages, using saved user messages to seed the exchanges.
Our investigation suggests that Xiaomi may have launched its MiMo-V2-Pro model with a free trial period—which was then extended—with the intent to use the surge in international developer use of the model to distill Claude capabilities. The bulk of the distillation attacks on Claude began just as the trial period was ending.
Distilling Claude is the business model.
What You Gonna Do About It, Punk?
The arms race continues. If you’re wondering why you can’t see the CoT, part of it is to guard it against supervision pressure a la An Alien Mind, but the main reason is avoiding distillation.
We use metadata and look for signals of irregular activity to identify accounts associated with proxy service networks. Instead of banning proxy accounts individually, we work to attribute this suspicious activity to a specific organization, allowing us to take comprehensive enforcement actions more effectively to prevent distillation attacks.
We’ve also built classifiers designed specifically to detect adversarial extraction.
We’ve also added new safeguards that make it harder for unauthorized labs to distill Claude’s capabilities. Claude now summarizes its internal reasoning before responding, which makes stolen transcripts less useful for training another model.
… As we investigate and disrupt distillation attacks, what we learn will continue to inform the safeguards we build.
Google has a clause where if you are attempting to distill Gemini, their plan is to intentionally sabotage the responses, to damage your operation. I believe this is the correct policy. If you have strong evidence the user is doing crime, to try and steal your stuff, as in a giant network of fraudulent accounts being used for distillation, then yes, you should be able to screw with them, not merely ban their accounts.
I see this as very distinct from the idea of silently degrading responses when users attempt to do AI R&D. That’s unacceptable and hostile, either refuse the requests or don’t, since this risks hitting normal work and now everyone has to be paranoid about being hit.
Distillation via massive fraudulent account networks and CoT extraction attacks is very obviously different. That is not a trigger you hit by accident, or without mens rea. You know what you did. If you play that game, you deserve whatever you get and more.
Good News, Everyone
Mostly the report is good news. Yes, there are Bad Dudes out there, trying to do various bad things. Occasionally they do something bad. It all adds up to not much. We happily accept this level of malicious use.
The details of many of these operations are pretty wild. This was a short summary.
Of course, there is also this, especially given no mention of North Korea.
PoIiMath: “We disrupted every operation in the report”
A Very Different Read of The Report
I read this report as mostly good news, as saying that the situation is basically fine, and also as therefore not big news. There were so many other stories that seemed much bigger that same week.
Ryan Fedasiuk does not see it that way. He overstates his case, in that he claims this was front page news everywhere, which it wasn’t. But he sees this, together with the NSA/CISA/FBI advisory on the distillation efforts, as fundamentally altering the US-China relationship and China’s attitude towards Anthropic, and China’s relationship to its top labs.
I do expect a substantial change in PRC’s relationship with its top labs, after revelations that they silently shipped a lot of consumer queries directly to Anthropic, including requests from China’s security services.
China threatened ‘resolute countermeasures’ in response to the advisory, a day before the full report got issued on the 10th. Then on the 11th Mao Ning at the MFA briefing did respond to the report in particular by saying China opposes attempts to ‘smear China by distorting facts.’
Ryan Fedasiuk: Here is a low-confidence theory: I think we are under-weighting the extent to which Chinese state media’s hostility toward U.S.-China coordination on AI security may be a response specifically to Anthropic’s AI Misuse report.
Anthropic’s report is one of the finest intelligence products I’ve ever seen. Not only does it document a world-historic Chinese counterintelligence failure—but it publicly, personally embarrassed China’s security services on the front page of every newspaper in the world.
This was a huge, huge deal for China. It will fundamentally alter the relationship between Chinese AI labs and the state. I am still expecting reprisals against the Chinese AI companies that were caught routing sensitive requests directly to Claude, unbeknownst to their Chinese users.
In fact, I’m sure this was part of Anthropic’s calculation to originally release the report. Embarrassing China’s security services and putting a target on the back of Chinese labs engaging in distillation is surely an effective deterrent to that practice.
But I also think this partly explains what we’re seeing with Chinese state media’s hostility toward Anthropic and Dario in particular. They view Anthropic not just as the tip of the American AI spear, but as a bad-faith actor gung-ho on smearing and destabilizing China’s political system.
It would be unfortunate if this personal, political animosity were now bleeding into wider discussions of U.S.-China coordination on AI safety—when, in fact, I do think the Party cares about AI risk, and will continue to care about AI risk (cf MSS Minister Chen Yixin’s recent commentary).
Julian Gewirtz: I haven’t seen much attention to this new China Daily editorial today. But it offers a revealing look at Beijing’s response to calls for an AI slowdown.
It explicitly connects the calls from @DarioAmodei , @sama , and @elonmusk with Anthropic’s recent report alleging illicit distillation by Chinese companies—portraying them as parts of a coordinated effort to protect American firms from Chinese competition.
Sam Altman says Trump and Xi could win the Nobel Peace Prize for an AI agreement: “I don’t think this is hard. This is like a one-page document.” But that’s wrong. Upcoming AI talks will be “hard.” And this editorial shows why.
The headline really sums it all up: “‘Dr Frankenstein’ alarm cries of US’ AI elites a self-serving bid for profit.”
A few key quotes from the China Daily:
–”The [distillation] report, and the corporate ‘alliance’ that followed it, amounted in essence to a coordinated play — a response to Chinese competition and to the regulatory pressure coming from Washington. Its aims were threefold: to blunt China’s AI advance, to win a favorable policy environment at home and to keep investors’ enthusiasm for US AI alight.”
–”The distinction between a security measure and a competitive moat can become blurred when the companies building the moat are also allowed to define ‘illicit distillation.'”
–”The proposed coordination [to ‘pace the frontier’] among the three companies sounds rather like a club whose membership rules have been drafted before the guest list is announced. A global AI-safety framework that excludes China is not quite global.”
The editorial doesn’t rule out diplomacy but is trying to set the terms for upcoming talks.
For more on China’s response to “Pacing the Frontier” [see here]. Anthropic is by far the most China-hostile of the American AI labs, calling for strong action on chips and distillation as core parts of their overall strategic posture. China is highly reasonably interpreting that, plus this kind of report and the full cutting off of all Chinese access to Claude, as being hostile. This could then by association make China less inclined to cooperate on catastrophic or existential risks, as they too are vulnerable to this kind of cognitive mistake.
I would as always caution against the tendency to treat each day’s developments as something that permanently alters or solidifies attitudes and relationships and conflicts. Every month looks very different from the previous month. I have lost track of the number of individual moves in the game that I have been told ‘inevitably’ led to huge permanent changes, without which maybe things would have gone differently. Such claims are almost always wrong.