MoE enthusiast An agent model from a phone company, trained without Nvidia. Xing4.0-29B-A4B is an open-weight agentic MoE from China Telecom AI (the Xing series, formerly TeleChat). It does the agent loop: plan multi-step tasks, call tools, chew through long documents, hand back finished work. The distinguishing fact is the training hardware: the whole run happened on Huawei Ascend NPUs with the MindSpore framework, Ascend 910C clusters, no Nvidia cards anywhere in it. It is the first model of this size with a chips-to-framework all-domestic chain behind it.
What runs where. 29B total parameters, 4B active per token, 64 routed experts (4 active + 1 shared) with MLA attention, 256K context native, extensible to 512K. In bf16 the weights are 62.4 GB (server territory). The deployment story is quantization: the vendor says a 4-bit quantized build needs 15 GB of GPU memory, and the math agrees (29B x 4 bits is 14.5 GB), which puts it on one RTX 3090 or 4090 with room left for KV cache. A community GGUF (IQ4_NL, 20.1 GB file) is already on HuggingFace with 10,880 downloads. vLLM, SGLang, and KTransformers are supported for serving; LLaMA-Factory and MindFormers for fine-tuning.
Benchmarks, labeled. SWE-bench Verified 75.0 and Terminal-Bench 2.1 57.5 are vendor-run numbers from the model card (SWE-agent harness, 210K context window); Terminal-Bench 2.1 at 57.5 versus 30.0 for Gemma4-26B-A4B and 51.5 for Qwen3.6-35B-A3B is the standout row. SuperCLUE agent capability 93.52 puts it third, less than one point behind the top two Qwen models. Treat all of these as vendor-measured until third parties replicate.
In production, not just on a chart. China Telecom runs it in its group-level customer service platform for multi-step inquiry resolution, and in mid-screen interactive scenarios. The company says larger Xing models are coming.
The market signal. A state telecom shipping a competitive small agent model, open weights, Apache 2.0, is a statement about where compute sovereignty is going: model quality is now achievable without touching Nvidia silicon, and the release PR is aimed at developers running agents on their desktops.
- 29.0B
- 256k
- apache 2.0
- 🇨🇳 China
- Sep 2026
Scores #
Guides covering Xing4.0-29B-A4B #
Save your hardware and every model page answers the real question: will it run on your machine, and how fast?
Run it locally #
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
The reference hardware
22 reference configs, drawn in-house. Scroll for more.
Can you run it? - reference rigs
| Rig | Q4_K_M | BF16 |
|---|---|---|
| NVIDIA Jetson Orin NX 16GB | [no -> cloud](#cloud-pricing) | [no -> cloud](#cloud-pricing) |
| Single GTX 1080 Ti (11GB) | offload | [no -> cloud](#cloud-pricing) |
| 4x H100 80GB (320GB) | fast 2653.7t/s | fast 855.8t/s | | NVIDIA DGX Station 748GB | fast 1584.3t/s | fast 510.9t/s | | 8x RTX 3090 rack (192GB) | fast 1483.2t/s | fast 478.3t/s | | 4x RTX 5090 (128GB) | fast 1419.6t/s | fast 457.8t/s | | AMD Instinct MI300X (192GB) | fast 1054.5t/s | fast 340.1t/s | | 4x RTX 4090 (96GB) | fast 798.5t/s | fast 257.5t/s | | 2x RTX 5090 (64GB) | fast 709.8t/s | tight | | 2x RTX 3090 (48GB) | fast 370.8t/s | offload | | Single RTX 5090 (32GB) | fast 354.9t/s | no -> cloud | | RTX PRO 6000 Blackwell (96GB) | fast 354.9t/s | fast 114.5t/s | | Mac Studio M4 Ultra 192GB | fast 235.9t/s | fast 76.1t/s | | Mac Studio M4 Ultra 512GB | fast 235.9t/s | fast 76.1t/s | | Single RTX 4090 (24GB) | fast 199.6t/s | no -> cloud | | MacBook Pro M5 Max 128GB | fast 132.7t/s | fast 42.8t/s | | Dual EPYC 9004 + 768GB DDR5-4800 | fast 91.3t/s | fast 29.4t/s | | DGX Spark 128GB unified | fast 54.1t/s | ok 17.4t/s | | Ryzen AI Max+ 395 128GB | fast 50.7t/s | ok 16.4t/s | | Jetson AGX Orin 64GB | fast 40.6t/s | tight | | Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 40.6t/s | ok 13.1t/s | | Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 30.4t/s | ok 9.8t/s |
Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options #
Or run it in the cloud #
No per-token API provider pricing tracked for Xing4.0-29B-A4B yet.
For flagship list prices, see the
[calculator](https://tokenstead.ai/calculator).
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.