# US vs China AI: the lead is basically gone

> Source: <https://promptcube3.com/en/news/4865/>
> Published: 2026-08-04 00:26:28+00:00

# **US vs China AI: the lead is basically gone**

Look at the trajectory. When GPT-4 launched, the best Chinese models were clearly a tier or two behind. Then [DeepSeek](/en/tags/deepseek/)'s R1 and V3 landed, and suddenly the reasoning gap shrank to a few quarters. Qwen became the default open family for a huge amount of agentic and AI workflow work, not just in China but globally. That doesn't happen if the underlying research isn't genuinely strong.

A few things I think are driving this:

**Compute efficiency under constraint.** Export controls didn't slow China down — they forced labs to optimize hard. MoE routing, aggressive quantization, and custom inference stacks are areas where Chinese teams now have deep, production-hardened experience. Necessity did its thing.**Open-weight distribution.** The Chinese open-source ecosystem is massive. Qwen, DeepSeek, GLM, Hunyuan — several of these match or beat US open-weight releases on math, code, and reasoning benchmarks. For real-world deployments, that's what actually matters.**Application density.** LLM agent and AI workflow deployments are shipping into Chinese manufacturing, finance, and education at a pace I haven't seen elsewhere. Models improve fastest when they're constantly in production with real feedback loops.

## Where the US still leads

**Frontier training at scale:** the absolute best closed models are still trained in the US on the largest GPU clusters.**The CUDA moat and full tooling stack:** PyTorch, Triton, the whole inference ecosystem.**Foundational research velocity.**

But none of those feel permanent. Chinese labs are actively building toolchains that bypass CUDA dependence, and the agentic tooling gap is much smaller than most people assume.

## What this means if you're building

Stop assuming the best model is American. For real-world AI workflow deployments that need strong open weights and a good cost-per-token, the Chinese model families are often the cheaper, better choice — especially if your stack is partly or fully in Chinese. A lot of the LLM agent work I've seen in the last year runs fine on these models with noticeably lower inference bills.

The conclusion I've landed on after following this closely: the US is no longer

[The Race to Beat Cheap AI from China: What It Really Takes 2h ago](/en/news/4909/)

[The AI Boom Is Reshaping the US Economy — What I'm Noticing 2h ago](/en/news/4907/)

[From rogue model to asset: taming a Chinese LLM in our lab 4h ago](/en/news/4898/)

[15 Attorneys General vs OpenAI: The Regulatory Push 4h ago](/en/news/4894/)

[CUDA's Moat Is Weakening, and AI Coding Agents Are the Pickaxe 5h ago](/en/news/4888/)

[If Astra Really Solved 10 Open Math Problems, Here's the Catch 7h ago](/en/news/4874/)

[Next FutureSearch: AI forecasting you can verify, now out of beta →](/en/news/4863/)
