{"slug": "huawei-unveils-ai-cluster-capable-of-linking-4096-chips-into-a-single-machine", "title": "Huawei unveils AI cluster capable of linking 4,096 chips into a single machine", "summary": "Huawei unveiled the Atlas 960E SuperPoD at HUAWEI CONNECT 2026 in Shanghai on September 17, an AI cluster that links 4,096 Ascend neural processing units into a single logical machine with unified memory addressing, targeting models up to 10 trillion parameters. The system delivers 8 EFLOPS at FP8 and 16 EFLOPS at FP4 precision with up to 1 petabyte of high-bandwidth memory, and Huawei said its Hi-ONE Near-Packaged Optics engine pushes 7.2 terabits per second per module while cutting power consumption by more than 550 kilowatts per pod. Huawei also pulled the Ascend 960DT training chip forward to Q1 2027, three quarters ahead of schedule, and stated a long-term goal of scaling to 1 million NPUs using multi-rail topology.", "body_md": "Photo: Jakub Pabis / Pexels\n\n# Huawei unveils AI cluster capable of linking 4,096 chips into a single machine\n\nThe Atlas 960E SuperPoD targets 10-trillion-parameter models with up to 16 EFLOPS of compute and a petabyte of high-bandwidth memory\n\nHuawei just made its boldest play yet in the AI hardware arms race. At HUAWEI CONNECT 2026 in Shanghai on September 17, the company revealed the Atlas 960E SuperPoD, an AI computing cluster that stitches together 4,096 of its Ascend neural processing units into what it describes as a single logical machine with unified memory addressing.\n\nThink of it like this: instead of 4,096 separate brains trying to coordinate over walkie-talkies, the system makes them behave like one enormous brain sharing one enormous memory. That distinction matters a lot when you’re training models with up to 10 trillion parameters.\n\n## The numbers behind the machine\n\nThe Atlas 960E delivers 8 EFLOPS at FP8 precision and 16 EFLOPS at FP4 precision. The full configuration supports up to 1 petabyte of high-bandwidth memory, which is roughly the equivalent of storing 250 million high-resolution photos in the fastest memory tier available.\n\nTo keep all those chips talking to each other without bottlenecks, Huawei developed what it calls the Hi-ONE Near-Packaged Optics engine. It’s an industry first, according to the company, pushing 7.2 terabits per second per module. That optical interconnect technology also cuts power consumption by more than 550 kilowatts per pod.\n\nCompared to previous Atlas generations, the 960E architecture is projected to deliver between 2.3 and 4 times the training and inference throughput for large models. Huawei also claims a 99.8% operational availability rate.\n\nThe company didn’t stop at the SuperPoD announcement. It revealed that the Ascend 960DT training chip, originally slated for a later release, has been pulled forward to Q1 2027, arriving three quarters ahead of schedule.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\n## Scaling ambitions and the competitive landscape\n\nPerhaps the most striking long-term detail is Huawei’s stated goal of scaling to 1 million NPUs using multi-rail topology.\n\nAlongside the Atlas 960E, Huawei also announced upgrades to its TaiShan 950 SuperPoD and introduced the OceanStor M900 storage solution.\n\n## Why the power efficiency angle matters more than it seems\n\nThe 550 kW power reduction per pod deserves more attention than a bullet point. AI data centers are increasingly bumping up against physical power constraints, with new facilities being built next to power plants and some regions running out of available grid capacity for new data center construction.\n\nThe 10-trillion-parameter target is worth contextualizing. Today’s largest publicly known models sit in the range of a few trillion parameters. By targeting 10 trillion, Huawei is building hardware for models that don’t widely exist yet.\n\nThe accelerated timeline for the Ascend 960DT chip adds another layer of competitive pressure. Pulling a major chip launch forward by three quarters compresses the window between announcement and availability, giving potential customers less reason to commit to alternative platforms while waiting.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/huawei-unveils-ai-cluster-capable-of-linking-4096-chips-into-a-single-machine", "canonical_source": "https://cryptobriefing.com/huawei-atlas-960e-ai-cluster-4096-chips/", "published_at": "2026-09-20 16:35:38+00:00", "updated_at": "2026-09-20 16:55:11.609768+00:00", "lang": "en", "topics": ["ai-chips", "ai-infrastructure", "artificial-intelligence", "large-language-models"], "entities": ["Huawei", "Atlas 960E SuperPoD", "Ascend", "Hi-ONE Near-Packaged Optics", "Ascend 960DT", "TaiShan 950 SuperPoD", "OceanStor M900", "HUAWEI CONNECT 2026"], "alternates": {"html": "https://wpnews.pro/news/huawei-unveils-ai-cluster-capable-of-linking-4096-chips-into-a-single-machine", "markdown": "https://wpnews.pro/news/huawei-unveils-ai-cluster-capable-of-linking-4096-chips-into-a-single-machine.md", "text": "https://wpnews.pro/news/huawei-unveils-ai-cluster-capable-of-linking-4096-chips-into-a-single-machine.txt", "jsonld": "https://wpnews.pro/news/huawei-unveils-ai-cluster-capable-of-linking-4096-chips-into-a-single-machine.jsonld"}}