cd /news/artificial-intelligence/xing4-0-29b-a4b · home › topics › artificial-intelligence › article
[ARTICLE · art-139652] src=tokenstead.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Xing4.0-29B-A4B

China Telecom AI released Xing4.0-29B-A4B, an open-weight agentic mixture-of-experts model under Apache 2.0 that was trained entirely on Huawei Ascend 910C NPUs with the MindSpore framework and no Nvidia hardware. The 29B-parameter model activates 4B parameters per token across 64 routed experts, supports 256K native context extensible to 512K, and posts vendor-run scores of 75.0 on SWE-bench Verified and 57.5 on Terminal-Bench 2.1, versus 30.0 for Gemma4-26B-A4B and 51.5 for Qwen3.6-35B-A3B. China Telecom says it runs the model in its group-level customer service platform for multi-step inquiry resolution, and a 4-bit quantized build needs 15 GB of GPU memory, putting it on a single RTX 3090 or 4090.

read4 min views1 publishedSep 25, 2026
Xing4.0-29B-A4B
Image: Tokenstead (auto-discovered)

MoE enthusiast An agent model from a phone company, trained without Nvidia. Xing4.0-29B-A4B is an open-weight agentic MoE from China Telecom AI (the Xing series, formerly TeleChat). It does the agent loop: plan multi-step tasks, call tools, chew through long documents, hand back finished work. The distinguishing fact is the training hardware: the whole run happened on Huawei Ascend NPUs with the MindSpore framework, Ascend 910C clusters, no Nvidia cards anywhere in it. It is the first model of this size with a chips-to-framework all-domestic chain behind it.

What runs where. 29B total parameters, 4B active per token, 64 routed experts (4 active + 1 shared) with MLA attention, 256K context native, extensible to 512K. In bf16 the weights are 62.4 GB (server territory). The deployment story is quantization: the vendor says a 4-bit quantized build needs 15 GB of GPU memory, and the math agrees (29B x 4 bits is 14.5 GB), which puts it on one RTX 3090 or 4090 with room left for KV cache. A community GGUF (IQ4_NL, 20.1 GB file) is already on HuggingFace with 10,880 downloads. vLLM, SGLang, and KTransformers are supported for serving; LLaMA-Factory and MindFormers for fine-tuning.

Benchmarks, labeled. SWE-bench Verified 75.0 and Terminal-Bench 2.1 57.5 are vendor-run numbers from the model card (SWE-agent harness, 210K context window); Terminal-Bench 2.1 at 57.5 versus 30.0 for Gemma4-26B-A4B and 51.5 for Qwen3.6-35B-A3B is the standout row. SuperCLUE agent capability 93.52 puts it third, less than one point behind the top two Qwen models. Treat all of these as vendor-measured until third parties replicate.

In production, not just on a chart. China Telecom runs it in its group-level customer service platform for multi-step inquiry resolution, and in mid-screen interactive scenarios. The company says larger Xing models are coming.

The market signal. A state telecom shipping a competitive small agent model, open weights, Apache 2.0, is a statement about where compute sovereignty is going: model quality is now achievable without touching Nvidia silicon, and the release PR is aimed at developers running agents on their desktops.

  • 29.0B
  • 256k
  • apache 2.0
  • 🇨🇳 China
  • Sep 2026

Scores #

Guides covering Xing4.0-29B-A4B #

Save your hardware and every model page answers the real question: will it run on your machine, and how fast?

Join free - save your rig →

Run it locally #

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

The reference hardware

22 reference configs, drawn in-house. Scroll for more.

Can you run it? - reference rigs

| Rig | Q4_K_M | BF16 |

|---|---|---|
| NVIDIA Jetson Orin NX 16GB |                          [no -> cloud](#cloud-pricing)  |                          [no -> cloud](#cloud-pricing)  | 
| Single GTX 1080 Ti (11GB) | offload |                          [no -> cloud](#cloud-pricing)  | 

| 4x H100 80GB (320GB) | fast 2653.7t/s | fast 855.8t/s | | NVIDIA DGX Station 748GB | fast 1584.3t/s | fast 510.9t/s | | 8x RTX 3090 rack (192GB) | fast 1483.2t/s | fast 478.3t/s | | 4x RTX 5090 (128GB) | fast 1419.6t/s | fast 457.8t/s | | AMD Instinct MI300X (192GB) | fast 1054.5t/s | fast 340.1t/s | | 4x RTX 4090 (96GB) | fast 798.5t/s | fast 257.5t/s | | 2x RTX 5090 (64GB) | fast 709.8t/s | tight | | 2x RTX 3090 (48GB) | fast 370.8t/s | offload | | Single RTX 5090 (32GB) | fast 354.9t/s | no -> cloud | | RTX PRO 6000 Blackwell (96GB) | fast 354.9t/s | fast 114.5t/s | | Mac Studio M4 Ultra 192GB | fast 235.9t/s | fast 76.1t/s | | Mac Studio M4 Ultra 512GB | fast 235.9t/s | fast 76.1t/s | | Single RTX 4090 (24GB) | fast 199.6t/s | no -> cloud | | MacBook Pro M5 Max 128GB | fast 132.7t/s | fast 42.8t/s | | Dual EPYC 9004 + 768GB DDR5-4800 | fast 91.3t/s | fast 29.4t/s | | DGX Spark 128GB unified | fast 54.1t/s | ok 17.4t/s | | Ryzen AI Max+ 395 128GB | fast 50.7t/s | ok 16.4t/s | | Jetson AGX Orin 64GB | fast 40.6t/s | tight | | Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 40.6t/s | ok 13.1t/s | | Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 30.4t/s | ok 9.8t/s |

Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options #

Or run it in the cloud #

    No per-token API provider pricing tracked for Xing4.0-29B-A4B yet.
        For flagship list prices, see the
        [calculator](https://tokenstead.ai/calculator).

Inference cost over time #

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @china telecom ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/xing4-0-29b-a4b] indexed:0 read:4min 2026-09-25 · —