Xing4.0-29B-A4B China Telecom AI released Xing4.0-29B-A4B, an open-weight agentic mixture-of-experts model under Apache 2.0 that was trained entirely on Huawei Ascend 910C NPUs with the MindSpore framework and no Nvidia hardware. The 29B-parameter model activates 4B parameters per token across 64 routed experts, supports 256K native context extensible to 512K, and posts vendor-run scores of 75.0 on SWE-bench Verified and 57.5 on Terminal-Bench 2.1, versus 30.0 for Gemma4-26B-A4B and 51.5 for Qwen3.6-35B-A3B. China Telecom says it runs the model in its group-level customer service platform for multi-step inquiry resolution, and a 4-bit quantized build needs 15 GB of GPU memory, putting it on a single RTX 3090 or 4090. Xing4.0-29B-A4B MoE enthusiast An agent model from a phone company, trained without Nvidia. Xing4.0-29B-A4B is an open-weight agentic MoE from China Telecom AI the Xing series, formerly TeleChat . It does the agent loop: plan multi-step tasks, call tools, chew through long documents, hand back finished work. The distinguishing fact is the training hardware: the whole run happened on Huawei Ascend NPUs with the MindSpore framework, Ascend 910C clusters, no Nvidia cards anywhere in it. It is the first model of this size with a chips-to-framework all-domestic chain behind it. What runs where. 29B total parameters, 4B active per token, 64 routed experts 4 active + 1 shared with MLA attention, 256K context native, extensible to 512K. In bf16 the weights are 62.4 GB server territory . The deployment story is quantization: the vendor says a 4-bit quantized build needs 15 GB of GPU memory, and the math agrees 29B x 4 bits is 14.5 GB , which puts it on one RTX 3090 or 4090 with room left for KV cache. A community GGUF IQ4 NL, 20.1 GB file is already on HuggingFace with 10,880 downloads. vLLM, SGLang, and KTransformers are supported for serving; LLaMA-Factory and MindFormers for fine-tuning. Benchmarks, labeled. SWE-bench Verified 75.0 and Terminal-Bench 2.1 57.5 are vendor-run numbers from the model card SWE-agent harness, 210K context window ; Terminal-Bench 2.1 at 57.5 versus 30.0 for Gemma4-26B-A4B and 51.5 for Qwen3.6-35B-A3B is the standout row. SuperCLUE agent capability 93.52 puts it third, less than one point behind the top two Qwen models. Treat all of these as vendor-measured until third parties replicate. In production, not just on a chart. China Telecom runs it in its group-level customer service platform for multi-step inquiry resolution, and in mid-screen interactive scenarios. The company says larger Xing models are coming. The market signal. A state telecom shipping a competitive small agent model, open weights, Apache 2.0, is a statement about where compute sovereignty is going: model quality is now achievable without touching Nvidia silicon, and the release PR is aimed at developers running agents on their desktops. - 29.0B - 256k - apache 2.0 - 🇨🇳 China - Sep 2026 Scores Guides covering Xing4.0-29B-A4B Save your hardware and every model page answers the real question: will it run on your machine, and how fast? Join free - save your rig → https://tokenstead.ai/login?return to=%2Fonboarding Run it locally Per-quant memory needs and a static "can you run it?" reference - no rig entry required The reference hardware 22 reference configs, drawn in-house. Scroll for more. Can you run it? - reference rigs | Rig | Q4 K M | BF16 | |---|---|---| | NVIDIA Jetson Orin NX 16GB | no - cloud cloud-pricing | no - cloud cloud-pricing | | Single GTX 1080 Ti 11GB | offload | no - cloud cloud-pricing | | 4x H100 80GB 320GB | fast 2653.7t/s | fast 855.8t/s | | NVIDIA DGX Station 748GB | fast 1584.3t/s | fast 510.9t/s | | 8x RTX 3090 rack 192GB | fast 1483.2t/s | fast 478.3t/s | | 4x RTX 5090 128GB | fast 1419.6t/s | fast 457.8t/s | | AMD Instinct MI300X 192GB | fast 1054.5t/s | fast 340.1t/s | | 4x RTX 4090 96GB | fast 798.5t/s | fast 257.5t/s | | 2x RTX 5090 64GB | fast 709.8t/s | tight | | 2x RTX 3090 48GB | fast 370.8t/s | offload | | Single RTX 5090 32GB | fast 354.9t/s | no - cloud cloud-pricing | | RTX PRO 6000 Blackwell 96GB | fast 354.9t/s | fast 114.5t/s | | Mac Studio M4 Ultra 192GB | fast 235.9t/s | fast 76.1t/s | | Mac Studio M4 Ultra 512GB | fast 235.9t/s | fast 76.1t/s | | Single RTX 4090 24GB | fast 199.6t/s | no - cloud cloud-pricing | | MacBook Pro M5 Max 128GB | fast 132.7t/s | fast 42.8t/s | | Dual EPYC 9004 + 768GB DDR5-4800 | fast 91.3t/s | fast 29.4t/s | | DGX Spark 128GB unified | fast 54.1t/s | ok 17.4t/s | | Ryzen AI Max+ 395 128GB | fast 50.7t/s | ok 16.4t/s | | Jetson AGX Orin 64GB | fast 40.6t/s | tight | | Epyc + 512GB DDR4-3200 + 2x RTX 3090 | fast 40.6t/s | ok 13.1t/s | | Epyc + 512GB DDR4-2400 + 2x RTX 3090 | fast 30.4t/s | ok 9.8t/s | Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast =20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark. Download options Or run it in the cloud No per-token API provider pricing tracked for Xing4.0-29B-A4B yet. For flagship list prices, see the calculator https://tokenstead.ai/calculator . Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.