cd /news/ai-chips/huawei-unveils-ai-cluster-capable-of… · home topics ai-chips article
[ARTICLE · art-135213] src=cryptobriefing.com ↗ pub= topic=ai-chips verified=true sentiment=↑ positive

Huawei unveils AI cluster capable of linking 4,096 chips into a single machine

Huawei unveiled the Atlas 960E SuperPoD at HUAWEI CONNECT 2026 in Shanghai on September 17, an AI cluster that links 4,096 Ascend neural processing units into a single logical machine with unified memory addressing, targeting models up to 10 trillion parameters. The system delivers 8 EFLOPS at FP8 and 16 EFLOPS at FP4 precision with up to 1 petabyte of high-bandwidth memory, and Huawei said its Hi-ONE Near-Packaged Optics engine pushes 7.2 terabits per second per module while cutting power consumption by more than 550 kilowatts per pod. Huawei also pulled the Ascend 960DT training chip forward to Q1 2027, three quarters ahead of schedule, and stated a long-term goal of scaling to 1 million NPUs using multi-rail topology.

read3 min views1 publishedSep 20, 2026
Huawei unveils AI cluster capable of linking 4,096 chips into a single machine
Image: Cryptobriefing (auto-discovered)

Photo: Jakub Pabis / Pexels

The Atlas 960E SuperPoD targets 10-trillion-parameter models with up to 16 EFLOPS of compute and a petabyte of high-bandwidth memory

Huawei just made its boldest play yet in the AI hardware arms race. At HUAWEI CONNECT 2026 in Shanghai on September 17, the company revealed the Atlas 960E SuperPoD, an AI computing cluster that stitches together 4,096 of its Ascend neural processing units into what it describes as a single logical machine with unified memory addressing.

Think of it like this: instead of 4,096 separate brains trying to coordinate over walkie-talkies, the system makes them behave like one enormous brain sharing one enormous memory. That distinction matters a lot when you’re training models with up to 10 trillion parameters.

The numbers behind the machine #

The Atlas 960E delivers 8 EFLOPS at FP8 precision and 16 EFLOPS at FP4 precision. The full configuration supports up to 1 petabyte of high-bandwidth memory, which is roughly the equivalent of storing 250 million high-resolution photos in the fastest memory tier available.

To keep all those chips talking to each other without bottlenecks, Huawei developed what it calls the Hi-ONE Near-Packaged Optics engine. It’s an industry first, according to the company, pushing 7.2 terabits per second per module. That optical interconnect technology also cuts power consumption by more than 550 kilowatts per pod.

Compared to previous Atlas generations, the 960E architecture is projected to deliver between 2.3 and 4 times the training and inference throughput for large models. Huawei also claims a 99.8% operational availability rate.

The company didn’t stop at the SuperPoD announcement. It revealed that the Ascend 960DT training chip, originally slated for a later release, has been pulled forward to Q1 2027, arriving three quarters ahead of schedule.

AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

Scaling ambitions and the competitive landscape #

Perhaps the most striking long-term detail is Huawei’s stated goal of scaling to 1 million NPUs using multi-rail topology.

Alongside the Atlas 960E, Huawei also announced upgrades to its TaiShan 950 SuperPoD and introduced the OceanStor M900 storage solution.

Why the power efficiency angle matters more than it seems #

The 550 kW power reduction per pod deserves more attention than a bullet point. AI data centers are increasingly bumping up against physical power constraints, with new facilities being built next to power plants and some regions running out of available grid capacity for new data center construction.

The 10-trillion-parameter target is worth contextualizing. Today’s largest publicly known models sit in the range of a few trillion parameters. By targeting 10 trillion, Huawei is building hardware for models that don’t widely exist yet.

The accelerated timeline for the Ascend 960DT chip adds another layer of competitive pressure. Pulling a major chip launch forward by three quarters compresses the window between announcement and availability, giving potential customers less reason to commit to alternative platforms while waiting.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #ai-chips 4 stories · sorted by recency
── more on @huawei 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/huawei-unveils-ai-cl…] indexed:0 read:3min 2026-09-20 ·