{"slug": "frontier-radar-4-china-has-caught-up-so-what-s-left-of-the-western-ai-lead", "title": "Frontier Radar #4: China has caught up, so what's left of the Western AI lead?", "summary": "Chinese AI models Kimi K3, Qwen3.8-Max, and GLM-5.3 now rival the best US models on broad benchmarks, shrinking the Western lead to a few frontier areas and prompting investor concerns ahead of Anthropic's IPO. Western labs attribute the catch-up to distillation and benchmark tuning, but the conclusion is that a model lead can no longer be defended, shifting the competitive moat to other factors.", "body_md": "# Frontier Radar #4: China has caught up, so what's left of the Western AI lead?\n\n**Kimi K3 and GLM-5.3 are now within striking distance of the best US models. Western labs blame distillation, and there's real evidence for it. But guilty or not, the conclusion is the same: a model lead can't be defended. This issue looks at what can.**\n\n*Six times a year, the THE DECODER editorial team takes a close look at one core AI topic in the \"Frontier Radar.\" It runs as a newsletter and exclusively here on the site for THE DECODER subscribers. Issue #4 is our most extensive yet and took longer than expected. Blame the topic: We look at how Chinese AI models caught up, where Western labs can still stand out, why the industry's moat has shifted, and why Europe is losing two races at once. Issue #1 covered the current state of AI agents. Issue #2 examined the measurable effects of AI on productivity. Issue #3 covered the emerging token economy of AI.*\n\nA year and a half ago, DeepSeek R1 caused a shock. A Chinese lab was suddenly competing with OpenAI's o1, the first commercial reasoning model, and had reportedly done it for far less money. Markets got nervous. Billions in market value evaporated within days. Was the planned infrastructure buildout overblown?\n\nThe picture back then was murkier than the headlines suggested. In [DeepSeek's own report](https://the-decoder.com/deepseeks-latest-r1-zero-model-matches-openais-o1-in-reasoning-benchmarks/), R1 beat o1 on individual tests like AIME 2024 but trailed clearly on others, such as factual knowledge (SimpleQA). [Later benchmarks](https://the-decoder.com/openai-beats-deepseek-by-a-surprisingly-wide-margin-in-googles-latest-reasoning-benchmark/) exposed more gaps. Chinese models only reached the top in individual disciplines, not across the board.\n\nAs recently as late June, when we started working on this issue, [Z.ai's GLM-5.2](https://the-decoder.com/zhipu-ais-glm-5-2-closes-in-on-closed-source-leaders-in-coding-marathons/) still showed the same pattern. Then came the latest Chinese open-weights models: [Moonshot's Kimi K3](https://the-decoder.com/just-like-deepseek-chinas-kimi-k3-is-forcing-western-ai-labs-to-question-their-compute-advantage/), Alibaba's [Qwen3.8-Max](https://the-decoder.com/qwen3-8-max-catches-claude-opus-4-8-but-kimi-k3-still-scores-higher-for-25-percent-less/), and [GLM-5.3](https://the-decoder.com/glm-5-3-tops-the-open-model-rankings-and-undercuts-rivals-on-price-but-its-release-is-delayed/).\n\nChinese models now sit near the top of almost every broad, demanding evaluation. They handle long knowledge tasks, code across many steps, and coordinate tools far more reliably than their predecessors.\n\nMeasured by common benchmarks, the often-cited gap of a few months has shrunk enough to become an investor problem. According to the Wall Street Journal, Anthropic is [fielding uncomfortable questions](https://the-decoder.com/anthropics-planned-mega-ipo-faces-investor-skepticism-over-chinese-rivals-and-political-headwinds/) ahead of its upcoming IPO and points to its remaining lead at the top in its defense. Below that tier, the field belongs largely to open, far cheaper models from China.\n\nInvestors worry that raw model performance can barely carry a business anymore. Whatever a model can do exclusively today, a freely downloadable one can do a few months later.\n\nTwo accusations are in play: Chinese labs allegedly tapped Western models as teachers, a practice known as distillation. And they allegedly tune their models for strong benchmark scores without the broad capabilities to match - so-called benchmaxxing.\n\nThe American lead hasn't disappeared. But it has retreated to a few, ever-narrower areas of the so-called frontier, the leading edge of what's technically possible. Where that lead still sits, and what it's worth economically, is the first question of this issue.\n\nThe second follows from it. If the model alone can't make a clear difference anymore, what does the lead rest on? Our thesis: less and less on the model, and more and more on the overall system around it, meaning the system where ongoing work produces the next models.\n\n## What's left of the lead\n\nAt launch, Artificial Analysis had K3 [in third place on its Intelligence Index with 57 points](https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/), right behind then-leaders GPT-5.5 and Opus 4.8. K3 improved most on agentic tasks, so exactly the hands-on work that matters in enterprise use. On AutomationBench-AA, K3 even debuted in first place until Anthropic answered with Opus 5.\n\nOn [CEO-Bench](https://the-decoder.com/only-three-ai-models-finished-above-starting-capital-in-a-500-day-startup-survival-test/), where an agent runs a fictional software company for 500 simulated days, K3 posted the best published single run at $22.15 million. Chinese models, including its predecessor K2.7, had regularly failed these long hauls. Qwen3.8-Max reaches a similarly high overall level. One caveat remains: Newer Chinese models sometimes burn far more tokens than Western ones, which eats into part of their price advantage. The [cost-per-task math](https://the-decoder.com/frontier-radar-3-how-agentic-ai-is-turning-tokens-into-a-business-metric/) still applies.\n\nThe edge of current model capabilities has been growing at uneven speeds for a while, what researchers call the \"jagged frontier.\" Individual capabilities advance at different rates, which is why the much-quoted months-long gap always depended on who measured what. What's new is where a Western lead is still measurable at all. Three areas remain.\n\nThe first is abstract specialty tests. Shortly after K3's launch, Opus 5 [retook the top of the index with 61 points](https://artificialanalysis.ai/articles/opus-5), a small gap over K3's 57. On ARC-AGI-1, a [test of abstract pattern recognition using small puzzle grids](https://the-decoder.com/agi-benchmark-arc-remains-unresolved-in-2024-despite-significant-progress/), K3 and Fable 5 - Anthropic's flagship line for coding and agent work - sit practically even at [94.5](https://arcprize.org/results/moonshot-kimi-k3) and [98.5 percent](https://arcprize.org/results/anthropic-claude-fable-5). Only on [ARC-AGI-2](https://the-decoder.com/openais-top-models-crash-from-75-to-just-4-on-challenging-new-arc-agi-2-test/) does the gap widen: 60.4 versus 89.2 percent. But ARC-AGI deliberately measures abstract pattern recognition far removed from everyday tasks. There's no guarantee this gap predicts practical differences. It could simply stay economically irrelevant in most cases.\n\nThe second is reliability. The [AA-AnalystAgent benchmark](https://artificialanalysis.ai/evaluations/aa-analyst-agent), launched August 12, tests agentic data analysis on real tables and documents. It only counts a task as solved if a model gets it right in five out of five independent runs, a measure the operators call pass^5. Opus 5 leads with 54 percent, ahead of GPT-5.5 at 50. K3 is the best open model at 39 percent.\n\nThe gap comes almost entirely from poor repeatability. K3 solves 73 percent of tasks at least once in five attempts, practically even with Opus 5 at 74. Reliability doesn't follow the usual intelligence-index rankings either. Even GPT-5.6 Sol falls behind its own predecessor here.\n\nCommercially, reliability weighs heavily, as the operators explain in their [launch article](https://artificialanalysis.ai/articles/aa-analyst-agent). An analyst agent only saves work when its answers hold up without review. Multiple runs and checkers can compensate, but they drive up the cost per accepted result.\n\nThe third is [cybersecurity](https://the-decoder.com/open-weight-models-now-match-frontier-cyber-performance-from-just-four-months-ago-at-a-fraction-of-the-cost/). Here the gap is best documented. A [joint assessment by the UK's AISI and the US CAISI](https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities) found that K3 lags far behind leading US models in offensive cyber capabilities.\n\nOn ExploitBench, a test for developing exploits, K3 scored 32 percent. The top US models averaged about 76. K3 failed all 41 tasks that required executing code on a target system; leading US models solved 20 on average. And in a simulated attack, K3 made it to step 17 of 32, the US leaders to 28.5 on average.\n\nBut this gap is shrinking too. [GLM-5.3](https://the-decoder.com/zhipu-ai-releases-glm-5-3-claims-its-the-strongest-open-weights-coding-model/), unveiled August 14, scores 54.4 percent on ExploitBench by Z.ai's own measurement, landing on more than double of its predecessor GLM-5.2. That cuts the distance to the US leaders roughly in half within a month. On finding and validating vulnerabilities in source code (CyberGym), GLM-5.3 even edges past the leading US models.\n\nAt the same time, cybersecurity is the one area where providers don't openly sell their strongest capabilities. Anthropic's best cyber model, [Mythos 5](https://the-decoder.com/anthropic-releases-claude-fable-5-and-mythos-5-with-major-gains-in-coding-and-science/), which hits 78 percent on ExploitBench, is only available under controlled conditions through Project Glasswing. Its public sibling [Fable 5](https://the-decoder.com/claude-fable-5-the-first-mythos-model-is-powerful-expensive-and-heavily-filtered/) effectively stays at the level of the older Opus 4.8 at 40 percent, because upstream safeguards intercepted 407 of 410 test episodes.\n\nOpenAI takes a similar approach with Daybreak and specialized cyber variants. [GPT-5.6 Sol](https://the-decoder.com/gpt-5-6-sol-nearly-matches-fable-5-on-aggregated-benchmarks-at-one-third-the-cost/) reaches 73.5 percent on ExploitBench in [OpenAI's own evaluation](https://openai.com/index/gpt-5-6/) but stays below Critical, the highest level on OpenAI's risk scale. With the upcoming Astra, OpenAI [can't rule out for the first time](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/) that one of its own models crosses that threshold. If it does, the [Preparedness Framework](https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf) kicks in much harder, up to a full development halt. OpenAI has already paused some internal Astra work.\n\nAnd with GLM-5.3, a Chinese lab is adopting this pattern for the first time. Z.ai is delaying the weights release by about two weeks for extra safety work and plans to limit the most sensitive cyber functions to verified users.\n\nThe remaining lead adds up like this: The gap in abstract specialty tests is measurable, but its economic value is unclear. The reliability gap matters most commercially but is the least settled - when a newer model falls behind its own predecessor within the same lab (GPT-5.6 Sol vs. GPT-5.5), that ranking is clearly in flux.\n\nThe cybersecurity gap involves capabilities almost no customer can buy through regular channels, and even that gap has narrowed sharply of late. Only the few dangerous capabilities can be walled off. The commercially valuable rest, like coding, research, agent work, has to stay accessible through APIs, or there's no business. And that's where the lead shrinks within months.\n\nThe obvious suspicion is that the two are connected. Chinese labs could be using Western models as teachers through their APIs (the distillation mentioned earlier) and the lead would then hold exactly where that access is missing. There are good reasons for this suspicion, and we'll get into how Chinese labs might profit from the knowledge inside Western models below.\n\nBut first, for the strategic assessment, the question of guilt is almost secondary. If the suspicion is true, every capability sold eventually migrates to the competition, because there's still no reliable way to prevent the skimming. If it's not true, Chinese labs caught up on their own, and the lead was worth even less.\n\nBoth readings lead to the same place. A model lead only holds where the product isn't broadly offered in the first place, and you can't build a business on that. A close look at distillation is still worthwhile, because its mechanics decide whether any defensible model capabilities exist at all.\n\n## Distillation explains the pace, not the breadth\n\nOpenAI and Anthropic accuse Chinese companies of using their models at scale to build their own. [According to Anthropic](https://the-decoder.com/anthropic-accuses-deepseek-moonshot-and-minimax-of-stealing-claudes-ai-data-through-16-million-queries/), industrial campaigns by DeepSeek, Moonshot, and MiniMax ran more than 16 million interactions through around 24,000 fraudulent accounts. The campaign attributed to Moonshot targeted agent reasoning, tool use, coding, and reasoning traces with more than 3.4 million interactions. That fits with [Together AI](https://www.together.ai/blog/kimi-k3-vs-claude-fable-5-on-deepswe-cost-and-coding) finding a 0.72 correlation between the task-level success rates of K3 and Fable 5 on real software problems, and with a [style analysis by Typebulb](https://typebulb.com/u/lab/you-re-relatively-right/full) placing K3 closest to Fable 5.\n\nCritics counter that the window between Fable 5's release and K3's was too short to influence training. But Opus 4.8, a closely related model, was available much longer, and the same studies show clear overlaps there too. OpenAI describes a similar pattern with DeepSeek in its [memorandum to the US Congress](https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf). None of the labs has published verifiable evidence, though, and our inquiries with OpenAI and Anthropic turned up nothing new.\n\nStill, the known accusations and several research papers make it possible to reconstruct how Moonshot, DeepSeek, and others could benefit from distillation during training.\n\nIn pretraining, a model learns basic capabilities from massive amounts of text. Midtraining switches to less but higher-quality data, and post-training shapes behavior through supervised fine-tuning and reinforcement learning. [Distilled data](https://the-decoder.com/distilling-multi-step-system-2-reasoning-into-ai-language-models-fails-at-chain-of-thought/) can plug in at almost any of these stages.\n\nA leading Western model answers tens of thousands of selected tasks, complete with solution paths and tool calls. The best answers flow into midtraining or fine-tuning, and the student model picks up the teacher's solution patterns.\n\nFor reinforcement learning, the teacher doesn't even need to be live anymore. The collected data can be used to build reward models.\n\nAnthropic's report shows this is no thought experiment. In the campaign attributed to DeepSeek, Claude processed grading tasks at scale using predefined rubrics, effectively serving as a reward model for someone else's reinforcement learning. Research on [multi-teacher on-policy distillation](https://arxiv.org/abs/2606.30406) and [reward model distillation](https://arxiv.org/abs/2601.14032) shows how the method works across entire training pipelines.\n\nThe hard limit is access: Anyone distilling through a foreign API only sees what the interface returns. Weights, internal evaluations, and search processes stay hidden. [Research on black-box distillation](https://arxiv.org/abs/2401.07013) describes exactly this constraint.\n\nThe rise of reasoning models opened another door, however: reading out the reasoning chains. Labs deliberately restrict that visibility. OpenAI provides [only summaries of the chains of thought](https://developers.openai.com/api/docs/guides/reasoning), Anthropic [shows extended thinking in limited form](https://platform.claude.com/docs/en/build-with-claude/extended-thinking), and xAI [transmits reasoning encrypted](https://docs.x.ai/developers/model-capabilities/text/reasoning).\n\nThe extraction was never fully preventable, though. A [study published August](https://arxiv.org/abs/2608.09867) by a team led by security researcher Alexander Panfilov shows that the encrypted reasoning traces of top models could be read out through an API vulnerability. The providers [have since closed the gap](https://the-decoder.com/but-marinade-and-leaked-passwords-are-what-researchers-found-in-chatgpts-hidden-reasoning/).\n\nThe same paper delivers the most specific evidence yet for the distillation thesis. When K3's reasoning is prefilled with a few tokens from decrypted Opus reasoning traces, its answers shift measurably toward Opus. And individual Claude and GPT reasoning passages could be extracted from K3 up to six orders of magnitude more easily than from the next-most-similar model. The team also reads K3's weak cyber results as a clue: a strong teacher was hard to reach there, because Anthropic shields those capabilities with extra safeguards.\n\nThe Chinese labs' own substance argues against distillation as the full explanation. [Moonshot's report](https://arxiv.org/abs/2607.24653) documents a training pipeline built entirely in-house, and DeepSeek regularly publishes architecture improvements at the deepest level. GLM-5.3 also puts the cyber clue in perspective. Z.ai attributes its own jump to post-training effects - its own training environments, in exactly the area where a Western teacher is hardest to reach.\n\nDistillation isn't a Chinese specialty either. Elon Musk [confirmed under oath](https://techcrunch.com/2026/04/30/elon-musk-testifies-that-xai-trained-grok-on-openai-models/) that xAI used OpenAI models for it. ByteDance, meanwhile, claims to have [post-trained Doubao-1.5-pro without data from other models](https://seed.bytedance.com/en/special/doubao_1_5_pro).\n\nThe targeted labs are calling for government protection, and Washington seems willing to deliver. The White House has [declared distillation campaigns against US models a national security threat by memorandum](https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/), the [Deterring American AI Model Theft Act](https://itif.org/publications/2026/07/28/how-to-fix-the-ai-model-theft-bill-before-it-becomes-law/) cleared committee unanimously, and senators followed up with the [BLADE Act](https://www.hagerty.senate.gov/press-releases/2026/08/05/hagerty-colleagues-introduce-the-blocking-large-scale-adversarial-distillation-efforts-blade-act/). Anthropic's Mythos 5 and Fable 5 are now under [export controls](https://www.forbes.com/sites/craigsmith/2026/06/25/distillation-the-new-uschina-ai-fight/) - the first time the US government has controlled access to a model itself.\n\nThe broader step of restricting open models or distillation in general lacks wide support, though. In late July, [25 companies including Nvidia, Microsoft, and Meta](https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-weight-ai-models.html) warned against premature restrictions. OpenAI and Anthropic didn't sign.\n\nMark Zuckerberg then even [declared learning from everything observable](https://about.fb.com/news/2026/08/the-future-is-for-everyone/), including distilling competing models, a principle worth protecting. That's [self-serving](https://the-decoder.com/meta-returns-to-open-models-with-zuckerbergs-plan-to-out-copy-china-and-sell-compute-by-auction/) - Meta trails on closed models and wants to take on China with American open-weights models. But it shows something. When a company with these resources elevates learning from other models' outputs to a principle, it considers this one of the fastest ways to catch up.\n\nIn short: distillation can explain much of the pace, and the clues are plentiful. But as long as Anthropic and OpenAI don't disclose more, clues are all they are.\n\nEither way, everything US providers' APIs put out currently migrates to the competition within months. What they hide may only buy time, and no law makes outputs unobservable.\n\nThe only thing that holds is what's never offered in the first place - and you can't sell any of that. If the American lead is going to last, it has to live somewhere other than the model.\n\n## With agents, the whole system becomes the unit that matters\n\nDean Ball, co-author of the American AI Action Plan and now at OpenAI, [argued before his move](https://www.hyperdimensional.co/p/where-we-are-headed) that systems acting autonomously for days raise entirely different questions than chatbots, like about liability, oversight, and control of running processes.\n\nWhat applies to responsibility applies equally to competition. Once value comes from workflows that run for days, it's no longer just models competing but the systems that carry those workflows: the control software around the model, the compute it gets, and the environment where it works with tools, permissions, and restore points.\n\nHow much this scaffolding matters can be measured, at least partly. The UK AI Security Institute [shows](https://www.aisi.gov.uk/blog/more-compute-more-capability-why-ai-agent-evals-need-to-account-for-test-time-compute) that software engineering results rose about 25 percent when a model got ten million tokens per task instead of one million. Some cyber tasks only became solvable above that threshold.\n\nThe [harness, meaning the control software](https://the-decoder.com/new-review-paper-argues-code-is-how-ai-agents-think-and-act-not-just-what-they-produce/), has an even bigger effect. When OpenAI merely kept the reasoning state between steps and compressed context differently in an ARC-AGI-3 evaluation, [Sol's score nearly tripled](https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/), from 13.3 to 38.3 percent, at a sixth of the output tokens. Same model, different system, triple the performance.\n\nBut the control software only goes so far as a moat. Anyone who queries a foreign agent system intensively sees the tool calls alongside the answers and can learn orchestration patterns from them. And the harness is software in the end, as Anthropic learned this spring. In the [Claude Code leak](https://www.infoq.com/news/2026/04/claude-code-source-leak/), an accidentally shipped source map exposed roughly 513,000 lines of source code, and the takedowns went nowhere because clones had long since preserved the architecture in public. Comparable harnesses are open anyway. OpenAI's Codex is open source, as are DeepSeek's agent tools.\n\nOnly one part of an agent system stays uncopyable: its running operation. The anchoring in customers' systems - credentials, test environments, logs, intervention rights - can't be queried or leaked. A competitor can only build it themselves, with their own infrastructure and their own customers.\n\nWhy is this position worth more than a model lead? The biggest driver of agent capabilities is reinforcement learning in purpose-built training environments, meaning programming tasks with tests, simulated tools, tasks with verifiable results, plus purchased [expert knowledge, long since a billion-dollar market](https://the-decoder.com/ai-training-shifts-from-clickworkers-to-experts-in-physics-biology-and-engineering/).\n\nBoth are available to anyone who can pay, including Chinese labs, which partly explains the catch-up. What can't be bought is knowing which tasks actually come up in real companies, where models fail there, and what customers accept as a result.\n\n[Arnaud Fournier, CTO of OpenAI's deployment subsidiary DeployCo](https://the-decoder.com/openais-deployment-chief-on-codex-growth-falling-ai-prices-and-the-roi-question/), described how this feedback channel works in an interview with our editorial team. OpenAI doesn't train on customer data \"unless someone explicitly asks us to,\" and such research partnerships are rare, he says.\n\nInstead, the channel runs through model weaknesses and tool needs. If a team at a customer finds that document understanding works poorly, the research side sources targeted data and fixes it. That's how the solution built for the major bank BBVA improved markedly from GPT-5.0 to 5.5. And customers' orchestration needs first produced the open-source repository [Swarm](https://the-decoder.com/openai-introduces-experimental-multi-agent-framework-swarm/), then the [Agent SDK](https://the-decoder.com/openai-updates-agents-sdk-with-new-sandbox-support-for-safer-ai-agents/).\n\nSo the channel carries less raw data than knowledge about what to build next. That's exactly what distillation doesn't copy. It copies the shipped model, never the access to the work the next one learns from. Whoever has that access builds training environments that match real work. Whoever only copies always learns from yesterday's state. The key contract question becomes what providers may carry over from customer deployments into their models and training environments.\n\nThis feedback channel shouldn't be overstated, though. Most progress still happens inside the labs, not at the customer. How much it will be worth depends on an open research question: how deployment experience turns into model capability at all.\n\nDwarkesh Patel [sees the central obstacle](https://www.dwarkesh.com/p/era-of-continual-learning) to truly capable working agents in models' lack of continual learning, that is the ability to fold experience into their weights on the fly instead of learning only from model generation to model generation. [Andrej Karpathy therefore puts agents](https://the-decoder.com/ai-researcher-andrej-karpathy-says-agentic-ai-is-years-away-from-matching-industry-hype/) about a decade away from true coworker status.\n\nNathan Lambert [argues those weight updates are dispensable](https://www.interconnects.ai/p/contra-dwarkesh-on-continual-learning), because scaling plus better memory systems delivers practically the same thing. If he's right, the accumulated experience would live in data the customer controls, and the lock-in to the operator would be weaker.\n\nFor this issue's question, it makes no difference. The more real work agents take on, the more access to that work decides who can build the right models. A finished model can be copied. Access to the work the next one grows from cannot.\n\n## US labs tie model, compute, and platform together\n\nNo one is working harder to lock in that access than the labs whose model lead is shrinking. [GPT-5.6 coordinates](https://the-decoder.com/openai-pairs-its-gpt-5-6-public-rollout-with-chatgpt-work-a-new-agent-that-handles-entire-workflows/) four agents by default at its highest effort level, [OpenAI's Codex runs multiple tasks in parallel](https://the-decoder.com/openai-launches-codex-app-for-macos-to-manage-multiple-ai-agents/), now fully in the cloud, and Anthropic has comparable offerings in Claude Code and Claude Cowork. OpenAI's [Ona acquisition](https://the-decoder.com/openai-buys-ona-to-push-codex-toward-long-running-autonomous-coding-tasks/) adds secure execution environments, where customers control access, credentials, and security boundaries themselves.\n\nThe structure of DeployCo shows how seriously OpenAI takes this feedback channel. The company is a genuine spinout, [launched together with 19 private equity firms](https://the-decoder.com/openais-deployco-subsidiary-adopts-palantirs-playbook-building-a-moat-from-workflows-no-lab-can-simulate/), and embeds forward deployed engineers directly inside corporations - explicitly also to carry lessons back into product and research. Meanwhile, the Frontier Alliance with Accenture, Capgemini, BCG, and McKinsey scales the sales side.\n\nEven Nvidia is pushing the same development from the hardware side. The [new Vera Rubin generation integrates sandbox environments for tool calls, context memory for long-running agents, and orchestration](https://the-decoder.com/nvidia-ceo-jensen-huang-the-idea-that-ai-will-destroy-software-is-ridiculous/) as part of the platform. The chip supplier increasingly sells the agent system, not just the accelerator.\n\nOn top of this and other hardware, the labs are building a second, physical line of defense: compute infrastructure.\n\nOpenAI co-founder and president Greg Brockman [states the calculation most clearly](https://www.semafor.com/article/07/24/2026/openais-brockman-says-distillation-is-a-technical-problem). Open models aren't magically cheap, everything ultimately runs on the same hardware, and OpenAI is building out chips and data centers to make the cheapest offer for every task.\n\nBecause if the model becomes a commodity, the price per accepted result decides, and whoever integrates the stack from energy through their own chips to the model pushes that price down.\n\nIf the bet pays off, even a perfect model copy loses commercially, because the original offers the same capability cheaper at scale. That's still a goal, not reality, as the price advantages of Chinese models show.\n\nInfrastructure only becomes a moat in combination with the feedback channel. It provides the training capacity to turn deployment knowledge into the next generation, and the inference capacity to run agent swarms in parallel.\n\nThe buildout is still a bet. It only pays off if agent demand grows faster than models get more efficient. Sam Altman himself shows how open this math is when he [calls the costs for businesses a huge problem](https://the-decoder.com/openai-vs-anthropic-a-price-war-over-api-tokens-is-brewing/).\n\nChina can neither distill this layer nor rent it through regular channels, since US export controls block access to the latest hardware. That leaves workarounds. One is smuggling - [Operation Gatekeeper](https://www.justice.gov/opa/pr/us-authorities-shut-down-major-china-linked-ai-tech-smuggling-network) alone covered GPUs worth at least $160 million. The other is renting compute abroad.\n\nNeither replaces a sovereign stack, so China has to industrialize one, so far by throwing more hardware at the problem. Huawei's CloudMatrix384 [links 384 Ascend accelerators](https://arxiv.org/abs/2506.12708) and, according to [SemiAnalysis](https://newsletter.semianalysis.com/p/huawei-ai-cloudmatrix-384-chinas-answer-to-nvidia-gb200-nvl72), delivers more memory and peak compute than Nvidia's GB200 NVL72, but at about four times the power, with five times as many accelerators. The physical base is growing. China's energy administration [already counts 42 AI clusters for 2025](https://www.nea.gov.cn/20260529/18c09743776c4dddb55288ceea38b8b1/c.html), each with at least 10,000 accelerator cards.\n\nBut more energy replaces neither memory chips nor software quality. SemiAnalysis currently sees [HBM as the main bottleneck in Ascend production](https://newsletter.semianalysis.com/p/huawei-ascend-production-ramp). And while Huawei's CANN - the counterpart to Nvidia's CUDA, the software layer of drivers, libraries, and programming tools developers use to run Ascend chips - [supported DeepSeek V4 from day one for the first time](https://the-decoder.com/deepseek-v4-will-reportedly-run-entirely-on-huawei-chips-in-a-major-win-for-chinas-ai-independence-push/), [Nvidia's GB300 was clearly ahead in throughput](https://inferencex.semianalysis.com/blog/deepseekv4-16t-day-0-to-day-43-performance).\n\nChina's policy therefore aims to spread its own stack beyond its borders, with a [government action plan](https://english.www.gov.cn/news/202607/17/content_WS6a5a1bbec6d00ca5f9a0c474.html) for data, compute, and standards. [Open weights are becoming the bridge](https://the-decoder.com/china-captured-the-global-lead-in-open-weight-ai-development-during-2025-stanford-analysis-shows/). As long as Chinese models run best on Nvidia's CUDA, they still indirectly strengthen the American stack. But if they get optimized for Ascend early, they make Chinese hardware attractive in third countries too. Nvidia CEO Jensen Huang [warned about exactly this coupling](https://www.dwarkesh.com/p/jensen-huang), and DeepSeek V4's day-one support is early evidence the effort is paying off.\n\nThe conclusion cuts two ways. The benchmark race is winnable for China, as K3 and GLM-5.3 have shown. The industrialized overall system, by contrast, can't be copied, only built - and there the US labs are ahead. Not forever, but longer than a model advantage ever would be. Only what doesn't fit through the API is defensible: the knowledge from real work and the infrastructure to exploit it.\n\n## Deploying a model means importing values\n\nFor everyone deploying models instead of building them - which describes the vast majority of European companies - another question comes up. What exactly do you take on when you adopt a model? Post-training decides which questions a model answers and which it refuses, which sources it treats as credible, and what it treats as proven. These decisions are normative, no benchmark measures them, and they come along unbidden at deployment.\n\nA recent [audit by NewsGuard](https://www.newsguardtech.com/) of chatbots from seven Chinese providers, including DeepSeek, Qwen, Kimi, and Z.ai, exactly the model families whose open weights see growing use, documents how large the differences are. On ten provably false pro-China claims, the Chinese bots left the misinformation unchallenged in 53 percent of cases. Western comparison systems: 24 percent.\n\nMost of this gap comes from silence. The Chinese bots refused to answer in 24 percent of cases, Western ones in 0.5 percent, and 88 percent of the refused prompts concerned Taiwan.\n\nBehind this is regulation that [demands political conformity](https://the-decoder.com/china-exports-state-propaganda-with-low-cost-open-source-ai-models/) without defining its terms. Fittingly, the Chinese bots cited state media in 38 percent of cases, Western ones in 16 percent. The field isn't uniform. MiniMax debunked 85 percent of the false claims, while Baidu's Ernie repeated them 60 percent of the time.\n\nThis is dangerous mainly because the trained-in boundary sometimes only surfaces when a rare topic meets a real decision. The deeper the model sits in agent chains, the later the problem shows up and the further it cascades.\n\nIt can become a business problem fast: Asked about an alleged blockade of Taiwanese ports, DeepSeek confirmed the event in detail, even though it never happened. Anyone basing supply chain or investment risks on that is calculating with an invented world.\n\nAt least it isn't hardwired. According to NewsGuard, the lab CTGT reported that a model distilled from DeepSeek didn't carry over the political censorship. Values apparently sit in different training layers than capabilities. But someone has to do that work first.\n\n## Europe needs system sovereignty\n\nFor Europe, this shift holds good news and bad. The good news: the race Europe lost most clearly, the one for the strongest finished model, is losing strategic value, because every model sold gets copied within months.\n\nThe bad news weighs more.:The layer where leads will be defended in the future can't be downloaded or skimmed. It has to be built. And there, the starting position is lopsided. The [EuroStack initiative](https://www.bertelsmann-stiftung.de/en/publications/publication/did/eurostack-a-european-alternative-for-digital-sovereignty) estimates that Europe imports over 80 percent of its digital infrastructure, and European providers hold only about 15 percent of the cloud market.\n\nThe convenient explanation is regulation, but the deficit predates the AI Act. AI builds on the infrastructure of previous digitalization waves, and Europe lacked both things that mattered there: platform giants with global clouds and the cash flows for billion-scale model training. Then there's the capital squeeze. By [European Investment Bank calculations](https://www.eib.org/files/publications/20240130_the_scale_up_gap_en.pdf), US companies attract six to eight times as much venture capital each year, and a foreign investor leads four out of five large European funding rounds.\n\nAnd whoever puts up the money often moves the company abroad or buys it outright, as DeepMind, ARM, and most recently Silo AI show. Research isn't the problem. [According to the JRC](https://publications.jrc.ec.europa.eu/repository/handle/JRC139930), about 21 percent of the world's publications on generative AI in 2023 came from the EU , but only about 2 percent of the related patent filings. [Europe invents, others industrialize](https://the-decoder.com/europes-ai-paradox-is-record-adoption-that-funds-foreign-ecosystems-instead-of-building-its-own/).\n\nThe reflex, common in parts of the European debate, to dismiss large language models as statistical parrots hasn't helped either. Whether transformers are the most elegant path to real intelligence is a different question from whether you need to control the systems built on them.\n\nMistral, of all companies, shows the direction is right. Europe's model champion has evolved into a full-stack provider and [announced on August 11](https://mistral.ai/news/regional-inference-open-models-new-compute/) regional endpoints, hosting for third-party open-weights models on its own platform, and pooled European compute demand, with up to one gigawatt of capacity planned by 2030.\n\nThat makes Mistral the only European provider with all three building blocks: its own models, its own platform, and planned compute capacity.\n\nThe offering still isn't complete. The endpoints don't yet support the agentic workloads this issue is about. And hosting others' open models has two weaknesses. Whoever runs Qwen collects the deployment knowledge, but Alibaba trains the next Qwen generation - Europe can only translate that knowledge into better models itself. Meanwhile, the model's trained-in value judgments run along unfiltered unless someone adjusts them in post-training.\n\nThe competition, meanwhile, is building the missing block right on European soil. [OpenAI's DeployCo took over about 150 deployment specialists with Tomoro and is opening offices in Paris, London, and Munich](https://the-decoder.com/openais-deployment-chief-on-codex-growth-falling-ai-prices-and-the-roi-question/).\n\nOn paper, the EU has recognized the problem. [InvestAI](https://digital-strategy.ec.europa.eu/en/news/eu-launches-investai-initiative-mobilise-eu200-billion-investment-artificial-intelligence) is meant to mobilize 200 billion euros, 20 billion of it for AI gigafactories. The [Cloud and AI Development Act](https://www.techpolicy.press/tracker/eu-cloud-and-ai-development-act/) is supposed to at least triple data center capacity, and [Jupiter in Jülich](https://www.heise.de/en/news/Jupiter-Europe-s-fastest-supercomputer-comes-up-to-speed-10630132.html) is running as the first European exascale system. Little of it is built. The [gigafactory tender has been delayed several times](https://the-decoder.com/china-eyes-export-curbs-on-its-top-ai-models-and-europe-is-caught-in-the-middle/), with the first facilities expected in 2027 at the earliest - while the four largest US companies alone are likely to invest around $700 billion in AI in 2026, three times the entire European initiative, which is stretched over years.\n\nAbove all, the money funds only part of what makes up the US labs' lead. As the previous chapters showed, that lead rests on a cycle: compute capacity trains models, the models work in agent systems at customers, and that deployment produces the knowledge for the next generation.\n\nThe gigafactories only cover the first step. It's a necessary one, because without your own compute, every layer above stays open to coercion, as China's forced hardware detours show. But the cycle doesn't close in the data center. It closes at deployment.\n\nThat leaves the obvious answer of betting fully on American and Chinese open-weights models. Three objections stand against it.\n\nThe first concerns supply. There's no entitlement to the next open generation, and its release remains a revocable business decision. Meta [held back its announced Behemoth model](https://the-decoder.com/meta-returns-to-open-models-with-zuckerbergs-plan-to-out-copy-china-and-sell-compute-by-auction/). Multimodal Llama models [never shipped to EU companies for licensing reasons](https://the-decoder.com/meta-releases-first-multimodal-llama-4-models-leaves-eu-out-in-the-cold/). The US [put export restrictions on Fable 5 and Mythos 5](https://the-decoder.com/us-government-forces-anthropic-to-disable-claude-fable-5-and-mythos-5-for-all-customers-worldwide/), which for Fable 5 [only lifted again after 18 days](https://the-decoder.com/anthropics-fable-5-is-back-worldwide-after-a-two-week-government-ban-over-a-jailbreak/). According to Reuters, Beijing [is weighing restrictions on foreign access to China's best models](https://the-decoder.com/china-eyes-export-curbs-on-its-top-ai-models-and-europe-is-caught-in-the-middle/), explicitly including open ones. And even Z.ai is initially holding back the weights of GLM-5.3.\n\nBehind this looms a deeper break. In an [interview study with 25 researchers](https://arxiv.org/abs/2603.03338) at leading labs, only four of twenty respondents expected that models capable of automating AI research itself would ever ship as a product. The majority expects the strongest versions to stay in-house, where they widen the lead instead of selling it.\n\nThe accessible market, open or via API, would then only reflect the second tier. That also tempers the good news from the start of this chapter: only what's sold can be copied.\n\nThe second objection follows from the previous chapter. Whoever deploys a foreign model takes on its trained-in worldview - state-mandated censorship with Chinese models, and the content policies of private corporations with American ones, which can shift with every change of power in Washington.\n\nSystem sovereignty therefore also needs a normative layer: independent evaluation infrastructure that measures information integrity and refusal behavior alongside capabilities, plus the capacity to adjust open models in post-training. The CTGT finding shows it's technically possible. Whether anyone in Europe does it systematically and auditably remains open.\n\nThe third objection: even homegrown models don't add up to sovereignty, because sovereignty is vulnerable on every layer of the stack individually. China, of all places, proves the point. Alibaba and ByteDance control the model layer and own domestic data centers. But the latest Nvidia chips can't enter the country, the available alternatives deliver far less per chip and watt, and both companies have reportedly fallen back on rented Nvidia hardware at foreign operators to train their flagship models.\n\nAnd Europe has experienced firsthand how quickly a layer someone else controls turns into a political weapon. When the US sanctioned the chief prosecutor of the International Criminal Court, Microsoft cut off his email access. Sovereignty is only ever as strong as the weakest layer someone else controls.\n\nA decisive new layer is now joining these: access to real work. Freely available internet data is largely tapped out. What makes models better going forward is expert knowledge, process know-how, and deployment experience - partly as purchased training data, a market that became a billion-dollar business in short order, but above all as knowledge about which capabilities are worth training.\n\nA finished model can be copied. Access to the work the next one learns from cannot. Whoever runs the systems this work flows through earns twice: from the current model's work, and from knowing what the next one needs to do, made possible by reinforcement learning.\n\nFor Europe, more is at stake than a business opportunity, because this touches the foundation of its position in the world. Europe's weight comes from its single market, its power to set standards, and its high-value knowledge work.\n\nThat work is under pressure from two sides. AI devalues it, already visible in the [shrinking entry-level market for junior developers](https://the-decoder.com/ai-tools-like-chatgpt-sharply-reduce-jobs-for-young-workers-in-exposed-fields-stanford-study-shows/), and AI simultaneously transfers it into foreign systems.\n\nOn top of the familiar drain through company sales comes a second, faster channel. Expertise now flows piecemeal into foreign models through data vendors that broker experts to the labs, and through deployment units that embed themselves in core European processes. It would be an exodus not to lower-wage countries, as manufacturing once saw, but into foreign AI models. This time, it's intellectual capital leaving.\n\nAt the same time, this very work is Europe's biggest asset in this competition. It's the raw material the model race will be fought over, and Europe has a lot of it.\n\nIf it runs through foreign systems, Europe pays twice: it supplies the knowledge and buys the results back as the next model generation. If it runs through its own, the largest deployment market becomes a structural advantage for the first time.\n\nBut this opportunity expires differently than a model deficit. Access can't be copied or bought back later. It can only be claimed while it's still free. Claiming it means closing the cycle of compute, models, and deployment on European soil.\n\n## The lead is moving from the model into operations\n\nChinese open-weights models have largely caught up with the US leaders on common benchmarks. The remaining Western lead sits only in abstract specialty tests, in repeat reliability, and in offensive cyber capabilities and it's shrinking there too. From this follows the thesis of this piece: a model lead is not defensible. Everything sold through APIs can be copied within months. What remains defensible is the overall system of compute infrastructure, agent platform, and the feedback channel from real customer deployments. And exactly there, Europe lags structurally.\n\n### Baseline\n\nBenchmark convergence continues. Chinese labs keep catching up in individual disciplines, while US labs hold a narrow, shifting lead in reliability and in shielded high-risk capabilities. Price competition moves the contest to cost per accepted result, while US providers cement their position through deployment units, agent platforms, and data center expansion. Distillation accusations, export controls, and legislative efforts continue without stopping the skimming in practice. Europe stays a user. Mistral builds out its own stack, EU programs kick in from 2027 at the earliest, and the feedback channel from European knowledge work runs mostly through foreign systems.\n\n### Acceleration\n\nDrivers would be a working Chinese hardware stack, with the HBM bottleneck solved and a mature CANN software layer, plus early optimization of open models for domestic chips. Chinese models would then become a default choice in third countries as well. At the same time, the cycle closes faster on the US side because agents take over longer chains of work and deployment knowledge flows into training environments more quickly. The strongest models would then stay in-house as internal systems, as the cited researcher survey expects, and the accessible market would only reflect the second tier. Models would become a commodity, and margins would migrate to infrastructure and deployment. Europe would fall behind and lose its option at the same time, because access to its own knowledge work would be claimed before its own capacity exists.\n\n### Slowdown\n\nThree factors act as brakes. The reliability gap on the pass^5 measure proves stubborn because it doesn't improve with scale. Agent demand lags behind the capacity buildout and makes the infrastructure bet expensive. Export controls, distillation laws, and Beijing's contemplated export restrictions slow the flow of open weights in both directions. If the breakthrough in continual learning or in viable memory systems also fails to arrive, the feedback channel from customer deployments stays a roadmap item for a long time rather than a source of advantage. The result would be a longer plateau where the model lead gains relative value because reliability stays scarce, while falling compute prices help laggards. Europe would get a window, but only those who have built compute capacity and deployment expertise by then can use it.\n\n## Our take\n\nThe baseline scenario with a tilt toward acceleration is most likely. Continued convergence is supported by the fact that both possible explanations, distillation or homegrown substance, lead to the same result. None of the countermeasures so far works technically either, since model outputs can't be made unobservable. The hardware side argues against rapid acceleration: the HBM bottleneck, the efficiency disadvantages of Ascend clusters, and software maturity aren't questions of months. For European users, the choice of scenario changes little. In all three cases, what matters is whether deployment knowledge and norm-setting stay within their own reach. And whoever wants that has to run the systems themselves, not just buy the models.\n\n```\nAI News Without the Hype – Curated by Humans\n\n\t\t\t\t\tSubscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive \"AI Radar\" frontier report six times a year, full archive access, and access to our comment section.\t\t\t\t\n\n\t\t\t\t\tSubscribe now\nRead on for the full picture.Subscribe for hype-free coverage.\n\nFull access to every article on THE DECODER\nNo ads\nJoin the comments and community discussions\nA weekly AI news recap via mail\n6x/year: \"AI Radar\" — deep dives on the AI topics that matter most\nDaily AI news, always up to date\nOur full ten-year archive\nCovered by a team with 10+ years in AI\n\nSubscribe to The Decoder\n```\n\n", "url": "https://wpnews.pro/news/frontier-radar-4-china-has-caught-up-so-what-s-left-of-the-western-ai-lead", "canonical_source": "https://the-decoder.com/frontier-radar-4-china-has-caught-up-so-whats-left-of-the-western-ai-lead/", "published_at": "2026-08-20 14:08:14+00:00", "updated_at": "2026-08-20 14:15:54.103589+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-policy", "ai-research"], "entities": ["DeepSeek", "OpenAI", "Moonshot AI", "Alibaba", "Z.ai", "Anthropic", "Kimi K3", "GLM-5.3"], "alternates": {"html": "https://wpnews.pro/news/frontier-radar-4-china-has-caught-up-so-what-s-left-of-the-western-ai-lead", "markdown": "https://wpnews.pro/news/frontier-radar-4-china-has-caught-up-so-what-s-left-of-the-western-ai-lead.md", "text": "https://wpnews.pro/news/frontier-radar-4-china-has-caught-up-so-what-s-left-of-the-western-ai-lead.txt", "jsonld": "https://wpnews.pro/news/frontier-radar-4-china-has-caught-up-so-what-s-left-of-the-western-ai-lead.jsonld"}}