# DeepSeek leads surge in low cost Chinese open-weight models on US platform

> Source: <https://www.scmp.com/tech/tech-trends/article/3365204/deepseek-leads-surge-low-cost-chinese-open-weight-models-us-platform?utm_source=rss_feed>
> Published: 2026-08-25 10:30:07+00:00

# DeepSeek leads surge in low cost Chinese open-weight models on US platform

Open models reached a record share of 62 per cent on a US web development platform on Saturday, a reversal of figures in June

The usage of open-weight AI models from China hit a record high on a popular US web development platform, driven largely by DeepSeek’s latest lightweight model, as business demand for Anthropic’s cutting-edge Fable 5 stalls due to high costs.

Open-weight models accounted for 54 per cent of token volume on Vercel’s AI Gateway on Tuesday, surpassing the 46 per cent consumed by proprietary models, according to data on its website.

Open models reached a record share of 62 per cent on the platform on Saturday, dwarfing closed systems’ 38 per cent, Vercel CEO Guillermo Rauch said on social media.

This was in stark contrast to the breakdown on June 24, when open-weight models accounted for only 28 per cent of Vercel AI Gateway’s token volume, compared with 72 per cent for closed ones, according to Rauch.

[DeepSeek-V4-Flash](https://www.scmp.com/tech/tech-trends/article/3363129/deepseek-signals-significant-price-hike-amid-surge-demand-low-cost-ai-models?module=inline&pgtype=article)was the most-used model in terms of token volume, while Chinese models dominated the top five, including Step 3.7 Flash from StepFun, GLM-5.2 from Zhipu, or Z.ai, and DeepSeek-V4-Flash’s updated 0731 version, which ranked second, fourth and fifth, respectively. The third spot was taken by OpenAI’s GPT-5.6 Luna, Vercel data showed.

[choosing cheaper and open models](https://www.scmp.com/tech/tech-trends/article/3364085/open-weight-chinese-ai-models-gain-foothold-europe-despite-brussels-trepidation?module=inline&pgtype=article)for production workloads, particularly for autonomous agents that can consume large numbers of tokens for reasoning, writing code and calling software tools.
