cd /news/artificial-intelligence/the-china-ai-gap-is-three-questions-… · home topics artificial-intelligence article
[ARTICLE · art-135736] src=hellochinatech.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

The China AI Gap Is Three Questions Now

On Terminal-Bench 4.0, the benchmark co-hosted by Stanford, Harbor, and the Laude Institute for multi-step software tasks, the best Chinese entry is Zhipu's GLM-5.3 at 41.8% as of September 19, while OpenAI's GPT-6 Astra holds the top score at 58.2%, a 16.4 percentage-point gap. On Artificial Analysis's DeepSWE repository-level coding benchmark, Kimi K3 and Codex both scored 68%, and Kimi K3 scored higher on SWE-Atlas-QnA, showing the China-US AI gap splits by task type rather than a single ranking. The comparison follows Anthropic's September 10, 2026 threat-intelligence report alleging seven China-based AI labs, including Alibaba and Moonshot, extracted Claude capabilities across over 151 million exchanges attributed to Alibaba alone over three months, and Nvidia's July 2025 disclosure that it generated 5 million training samples using DeepSeek's R1 model for its Nemotron family.

by read3 min views1 publishedSep 21, 2026
The China AI Gap Is Three Questions Now
Image: Hellochinatech (auto-discovered)

Put Kimi K3 and Codex on one coding benchmark and they finish with the same score. Put them inside a terminal environment and the gap widens sharply. Ask a different question and a different country leads: on AI video generation, Chinese models occupy most of the top ten. Ask whose training data came from whose model and the answer depends on which direction you look.

As of September 19, on Terminal-Bench 4.0, the benchmark co-hosted by Stanford, Harbor, and the Laude Institute for multi-step software tasks, the best Chinese entry is Zhipu’s GLM-5.3 at 41.8%. The top score belongs to OpenAI’s GPT-6 Astra at 58.2%. On a different type of coding evaluation, the picture changes. Artificial Analysis tested Kimi K3 and Codex on DeepSWE, a benchmark of original, repository-level engineering tasks. Both scored 68%. On SWE-Atlas-QnA, Kimi K3 scored higher.

On one repository-level coding benchmark, Kimi K3 matches Codex. On terminal-based agent execution, the gap is much wider. The split is not simply short versus long. It is between different forms of agent work.

Two data points frame what follows. In July 2025, Nvidia published a blog post describing how it generated 5 million training samples using DeepSeek’s R1 model to train its own Nemotron family. Fourteen months later, on September 10, 2026, Anthropic published a report alleging that seven China-based AI labs, including Alibaba and Moonshot, had conducted large-scale campaigns to extract Claude’s capabilities. Anthropic attributed over 151 million exchanges to Alibaba alone over a three-month period.

Everyone is learning from everyone. The rules are not the same.

A venture investor who recently visited Silicon Valley wrote that researchers at several American labs had started splitting the China gap by training stage and modality. They no longer rank the two countries on a single scale. The observation matches what the leaderboards show. Three questions do a better job.

Three Questions, Not One Ranking #

In July, DeepSeek founder Liang Wenfeng laid out a theory in a nearly four-hour investor meeting. “All the differences we see, including talent, model capability, and applications, can be attributed to differences in compute resources.” Every gap traces back to one variable: chips.

That theory explains part of what the leaderboards show. It does not explain all of it. If compute were the only variable, China would trail everywhere. It does not. If compute did not matter, China would not trail anywhere. It does.

Three questions sort the evidence better than one theory.

Can the model do a whole job? On Terminal-Bench, which tests agents inside an interactive terminal, the best-listed Chinese system scores 16.4 percentage points below the top US entry. Bloomberg reported, citing people familiar with the matter, that Anthropic’s revenue run rate reached $65 billion by the end of July, up from $47 billion in May. Anthropic had separately disclosed that Claude Code’s run rate exceeded $2.5 billion in February. Coding and agent products are a material part of the business. Chinese labs trail on a capability that has already become commercially significant for Anthropic.

Who decides what good looks like? On AI video generation, Chinese models hold five of the top ten non-preliminary spots on LMArena’s image-to-video arena (as of September 15). At least seven of the top ten on Artificial Analysis’s text-to-video board are built on Chinese base models. On image editing, where human judgment also matters, China does not lead. Seedream 5.0 Pro, ByteDance’s best entry, ranks ninth.

Whose model taught it? Nvidia openly trained its Nemotron models on DeepSeek’s output. DeepSeek’s licence says anyone may use its weights for “distillation for training other LLMs.” OpenAI’s business terms prohibit using output to develop competing AI models. Training on a competitor’s output is permitted by the licence in one direction and a contract violation in the other.

Each question points to a different mechanism. Why might compute and training infrastructure matter more in terminal-based agent work? Why might platform ownership matter more in video than in image editing? And how do different model licences shape the market for distillation? Upgrade to paid for the full analysis.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @zhipu 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-china-ai-gap-is-…] indexed:0 read:3min 2026-09-21 ·