Alibaba's TSMC-Built 5nm RISC-V Chip, XuanTie C950, Now Runs Qwen-3.8 27B Model Natively, Unlocking Massive Vertical Integration Tailwinds Alibaba has brought day-zero support for its Qwen-3.8 27B AI model to its RISC-V-based XuanTie C950 chip, achieving 30 tokens per second decode speed and a 1.9-second time to first token. The chip, reportedly fabricated by TSMC on a 5nm node, is the first RISC-V processor designed to run billion-parameter LLMs natively without GPUs, enabling Alibaba to vertically integrate its AI software and hardware and reduce reliance on external chip vendors. Alibaba appears to be emulating NVIDIA's well-established vertical integration playbook by bringing day-zero support for its highly capable Qwen-3.8 27B AI model - one that can run on just 32GB of VRAM https://x.com/tphuang/status/2089551401084944622 - to its bespoke RISC-V chip, called XuanTie C950. Alibaba can now run its Qwen-series AI models on its own chips, carving out a hefty moat for itself in one of the world's most competitive AI markets Alibaba unveiled the XuanTie C950 in March 2026, marketing the chip as a RISC-V-based offering for edge AI. Unlike typical ASICs, the XuanTie C950 does not rely on GPUs for AI workloads. Instead, the chip is basically a server-grade 64-bit RISC-V processor, replete with 64 compute cores located on a single piece of silicon, with clock frequencies that are scalable up to 3.20GHz, and where multiple clusters - 8 cores per cluster - are linked together natively using high-speed AMBA CHI fabrics. What's more, to handle demanding AI workloads, matrix and vector acceleration engines are embedded directly into the chip https://circuitdigest.com/news/alibaba-unveiled-xuantie-c950-high-performance-risc-v-core-for-edge-ai , eliminating the need for GPUs. Alibaba's XuanTie C950 features standard L1 caches, a flexible and highly configurable L2 cache, and supports an optional shared L3 cache to prevent inter-core communication bottlenecks. The chip also utilizes hardware-level intelligent data prefetching algorithms to load memory strings into the cache hierarchy before the execution engine requests them. Other important details include: - The chip is based on the open-source RISC-V ISA, which allows Alibaba to bypass licensing fees associated with the x86 architecture or ARM's designs. This also allows for greater customization. - The chip utilizes an 8-instruction decode width, allowing the core to read and process a large volume of commands simultaneously. - The chip features a 16-stage pipeline, which strikes a balance between maintaining high clock speeds and efficiently executing complex server and AI workloads by breaking down each instruction into 16 parts. - Unlike GPUs, the XuanTie C950 runs a single inference thread per socket, making it better suited for edge deployment and private inference rather than high-concurrency public APIs. - The combination of the custom pipeline and the integrated acceleration engines makes the XuanTie C950 the first RISC-V processor designed to run billion-parameter Large Language Models LLMs completely natively, and without the need for emulation or translation layers - the hardware units and instruction set extensions are designed to directly execute the core operations required by small- and medium-sized AI models. - Critically, the XuanTie C950 is believed to be fabricated by TSMC on its 5nm node https://www.trendforce.com/news/2026/03/25/news-alibaba-unveils-risc-v-xuantie-c950-cpu-for-ai-agents-5nm-chip-reportedly-made-by-tsmc/ , though Alibaba has issued no direct confirmation. This brings us to the core of today's topic. Alibaba has now brought day-zero support for its latest Qwen-3.8 27B model to the XuanTie C950 chip, offering decode speeds of 30 tokens per second, and a Time To First Token TTFT of just 1.9 seconds. For the benefit of those who might not be aware, the Qwen-3.8 27B https://wccftech.com/deepseeks-peak-hour-pricing-betrays-where-its-users-really-live-calming-us-fears-of-a-china-ai-takeover-even-as-alibabas-qwen-models-bury-meta-on-hugging-face/ is a 27-billion-parameter open-weight AI model that sports coding capabilities that are similar to Opus 4.5, and yet can run on a single MacBook. By bringing day-zero support for this model to the XuanTie C950, Alibaba is not only trying to lock customers within its own ecosystem but also substantially expanding the optionality around its compute footprint. After all, Alibaba can easily pair the C950 with other AI accelerators within its data centers to efficiently handle inference-related workloads. Follow Wccftech on Google https://profile.google.com/cp/Cg0vZy8xMWM3NDB2MmIyGgA to get more of our news coverage in your feeds.