cd /news/artificial-intelligence/8-chinese-ai-models-on-the-same-pi-c… · home topics artificial-intelligence article
[ARTICLE · art-114404] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

8 Chinese AI models on the same Pi coding-agent task: what we measured

A developer benchmarked eight Chinese AI models on the same coding-agent task using Pi as the agent, finding all completed the task successfully. The test, published on Vancine, provides model-by-model runtime and token data, with the developer noting it is a narrow test and not a general model ranking.

read2 min views1 publishedAug 28, 2026

Most model comparisons try to answer a question that is too broad: “Which model is best?” We wanted a smaller, reproducible question instead:

What happens when eight current Chinese AI models receive the same coding task through the same agent?

We used Pi as the coding agent and ran one isolated JavaScript task across these models:

Each run started from its own copy of the fixture. The test directory was kept unchanged, the work directory was checked for unexpected files, and raw run evidence was stored separately from the task workspace. The same Pi provider configuration and task contract were used for every model.

This is deliberately a narrow test. It does not measure architecture work, long-horizon debugging, frontend judgment, or performance on a real production repository.

All eight models completed the task successfully. Across the complete run we recorded:

The public page includes the model-by-model table, runtime and token measurements, the Pi configuration, methodology notes, and a downloadable JSON file:

View the full benchmark and data A small benchmark cannot tell you which model is generally better. It can still answer useful operational questions:

For this task, the answer to the first question was yes for all eight models. The differences are in the detailed run data, not a winner label. The total above is the audited amount recorded for these eight runs, not a forecast for arbitrary coding work. Agent cost depends heavily on task length, retries, context growth, and tool behavior. A real repository can be much more expensive than this small fixture.

The benchmark page includes the Pi setup and downloadable structured results. If you repeat it, keep the task, tests, agent version, model IDs, timeout, and evidence rules fixed. Otherwise you are comparing different experiments.

Disclosure: I operate Vancine, the OpenAI-compatible API used for these runs. The page is published as product evidence, and the result should not be read as a general model ranking.

I would especially value feedback on the harness and on what the next coding-agent task should test.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @pi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/8-chinese-ai-models-…] indexed:0 read:2min 2026-08-28 ·