cd /news/ai-chips/huawei-pulls-forward-ascend-960-road… · home topics ai-chips article
[ARTICLE · art-133718] src=mlq.ai ↗ pub= topic=ai-chips verified=true sentiment=· neutral

Huawei pulls forward Ascend 960 roadmap as it targets million-processor AI systems

Huawei pulled forward its Ascend 960 AI accelerator roadmap at its Huawei Connect conference in Shanghai on September 17, 2026, moving the training-focused Ascend 960DT to Q1 2027 and the inference-focused Ascend 960PR to Q3 2027 from a previously disclosed Q4 2027 target for the 960 family. Huawei lists the 960DT at 2 FP8 petaflops and 4 FP4 petaflops with 288 GB of memory, and the 960PR at 2 FP8 petaflops and 8 FP4 petaflops for inference with 192 GB of memory, a figure Tom's Hardware reported is twice Huawei's 2025 estimate. Huawei also introduced Peerium, a system architecture built on its UnifiedBus interconnect with a stated path to one million accelerator cards, alongside an Ascend 960 supernode using near-packaged optics; the figures are Huawei roadmap claims, not independent benchmarks, and Huawei has not publicly identified the fabrication process or memory suppliers for the 960 family.

read5 min views3 publishedSep 18, 2026
Huawei pulls forward Ascend 960 roadmap as it targets million-processor AI systems
Image: Mlq (auto-discovered)
  • Huawei now targets the Ascend 960DT for Q1 2027 and the inference-focused Ascend 960PR for Q3 2027, versus Q4 2027 for the 960 family in its 2025 roadmap. <sup>[1]</sup>
  • Huawei lists the 960DT at 2 FP8 petaflops and 4 FP4 petaflops, while the 960PR is listed at 2 FP8 petaflops and 8 FP4 petaflops for inference. <sup>[2]</sup>
  • The company's Peerium architecture uses UnifiedBus to connect processors, memory, storage and networking equipment, with a stated path to one million accelerator cards. <sup>[3]</sup>
  • The figures are Huawei roadmap claims, not independent benchmark results, and the company has not publicly identified the fabrication process or memory suppliers for the 960 family. <sup>[4]</sup>

Huawei has accelerated its next generation of Ascend AI accelerators as it tries to make large networks of Chinese-made chips operate as one computer. The Ascend 960DT is now scheduled for the first quarter of 2027, followed by the inference-focused Ascend 960PR in the third quarter, Huawei said at its Huawei Connect conference in Shanghai on September 17, 2026. [5]

The change brings the 960 family forward from Huawei's Q4 2027 target disclosed in 2025. The company also introduced Peerium, a system architecture built around its UnifiedBus interconnect, and announced an Ascend 960 supernode using near-packaged optics. [6]

Two 960 variants, different workload priorities #

Huawei's roadmap separates the 960 family by workload. The 960DT is designed for model training and the decode stage of inference. Huawei lists it at 2 FP8 petaflops and 4 FP4 petaflops, with 288 GB of memory, 9.6 TB/s of memory bandwidth and 2.2 TB/s of interconnect bandwidth. [7]

The 960PR targets prefill and recommendation workloads. Huawei's figures put it at 2 FP8 petaflops and 8 FP4 petaflops for inference, with 192 GB of memory, 2.4 TB/s of memory bandwidth and a 2.2-TB/s interconnect. Tom's Hardware reported that the 8-FP4-petaflop figure is twice Huawei's 2025 estimate. [8]

Huawei also added the Ascend 980 to the longer-range roadmap for 2029. Tom's Hardware reported 14 FP4 petaflops for the 2028 Ascend 970 and preliminary figures of 28 FP4 petaflops and 384 GB of memory for the 980. Those later figures should be treated as projections rather than shipping specifications. [9]

The system architecture is the larger claim #

Peerium is Huawei's attempt to address the limits of scaling individual accelerator servers. Huawei describes UnifiedBus as a common protocol for connecting CPUs, NPUs, memory, SSDs, network interface cards and switches. The architecture uses unified memory addressing and peer interconnects so that processors can operate as one logical system. [10]

Huawei's announced Ascend 960 supernode uses its Hi-ONE near-packaged optical engine and liquid cooling. The company says the system can scale to 4,096 accelerator cards and deliver 8 exaflops of FP8 or 16 exaflops of FP4 performance. It says multiple supernodes can be connected into a cluster of up to 512,000 cards, with a further topology supporting one million cards. [11]

These are system-level claims from Huawei, not results from an independent benchmark. The company also says an Atlas 950 SuperCluster with 256,000 cards is already being deployed and that the Atlas 960 system is under testing, but it has not named the customer or disclosed utilization, model workload or achieved throughput. [12]

Fabrication and memory remain the key unknowns #

Huawei's current roadmap materials do not specify the logic process used for the Ascend 960 or identify the memory manufacturer. Earlier Huawei materials described proprietary HiBL and HiZQ memory products for the Ascend 950 family, including a 144 GB HiZQ 2.0 configuration with 4 TB/s of bandwidth, but that disclosure does not establish the sourcing or production scale of the 960 memory configurations. [13]

Independent analysis has identified Chinese advanced-logic capacity, high-bandwidth memory and advanced packaging as likely constraints on Huawei's accelerator output. Epoch AI's modeling argues that Huawei's projected gains will depend heavily on packaging and memory supply, but those are analytical estimates rather than Huawei disclosures or independent production data. [14]

The software picture is similarly incomplete. Huawei says its CANN stack is fully open source and has highlighted work with PyTorch, Triton, vLLM and verl. That may improve access for developers, but it does not demonstrate parity with Nvidia's CUDA libraries, tools or production support. A 2026 field study of Ascend deployments reported that researchers needed source-level patches and operational workarounds for some inference workloads. [15]

A narrower Nvidia comparison #

Huawei's announced per-chip figures trail Nvidia's published Rubin specifications. Nvidia lists a Rubin GPU at 35 FP4 petaflops for dense training and 50 FP4 petaflops for inference, with 288 GB of HBM4 and up to 22 TB/s of memory bandwidth. Huawei's 960DT is listed at 4 FP4 petaflops, while its 960PR is listed at 8 FP4 petaflops for inference. [16]

The comparison is not a complete measure of system performance. Huawei is emphasizing scale, optical interconnect and unified memory, while Nvidia's figures are for a different architecture and precision stack. Neither company's peak figures establish delivered throughput, energy efficiency, reliability or cost on the same model and workload. [17]

The verified development is therefore a faster Ascend schedule and a more ambitious system architecture, not a demonstrated performance reversal against Nvidia. Huawei's ability to turn the roadmap into reliable, memory-equipped systems at commercial volume will be the more consequential test in 2027. [18]

Companies mentioned #

Further sources #

The stories that matter, in one email. Free — unsubscribe anytime.

── more in #ai-chips 4 stories · sorted by recency
── more on @huawei 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/huawei-pulls-forward…] indexed:0 read:5min 2026-09-18 ·