Huawei’s OceanStor KV cache storage for hyper-scale AI data centers Huawei announced the OceanStor M900, a scale-out all-flash storage cluster providing up to 64 PB of KV cache base tier for its Atlas 960 SuperPoDs, which scale to 4,096 NPUs and 8 EFLOPS of FP8 compute. Huawei claims the M900 delivers up to 4 PB of shared L3.5 KV cache capacity and 40 TB/sec of aggregate bandwidth per cluster, cutting access latency from milliseconds to about 60 microseconds and extending SSD read/write lifespan 16-fold, per David Wang, Deputy Chairman of Huawei's Board and Rotating Chairman. The system targets 10-trillion parameter models and million-plus token context windows, where a single accelerator's high-bandwidth memory cannot hold all key-value tokens, positioning Huawei against Nvidia's AI Data Platform and SuperPOD reference designs. Huawei’s OceanStor KV cache storage for hyper-scale AI data centers Huawei https://www.blocksandfiles.com/flash/2026/09/14/beegfs-runs-on-huawai-oceandisk-storage-servers/5296296 has announced the OceanStor M900, a scale-out, all-flash storage cluster providing an up to 64 PB KV cache base tier for its Atlas 960 SuperPoDs; rack-scale AI accelerators similar to Nvidia’s SuperPODs. Nvidia has defined an AI Data Platform https://www.blocksandfiles.com/ai-ml/2026/03/30/nvidia-and-its-partners-kv-cache-extenders/5209284 reference design integrating its GPUs, BlueField NICs, Spectrum-X switches, and AI software with storage systems, which support GPUDirect, RDMA, and KV cache extension with its Dynamo software and CMX scheme. This has a multi-tier design with GPU’s high-bandwidth memory HBM being the top and fastest tier, then the associated X86 servers’s DRAM being level 2, the server’s local SSDs being L3, and intelligent BlueField-4 NIC-connected NVMe SSDs in a flash storage server functioning as L3.5. These can provide a single network hop to a GPU’s HBM for microsecond-class latency. Huawei is now taking on Nvidia’s SuperPOD design with its own SuperPoDs for hyper-scale AI data centers. They are aimed at 10-trillion parameter models and million-plus token context windows, meaning a single accelerator’s high-bandwidth memory cannot contain all the key-value tokens needed and a caching scheme is needed. The company says a single Atlas 960E SuperPoD can scale up to 4,096 NPUs Neural Processing Units , delivering 8 EFLOPS of FP8 compute performance, with up to 1 petabyte of HBM capacity and a 256 TB unified memory pool. With 5,500 Hi-ONE units, which use UnifiedBus near-packaged optics https://www.theregister.com/networks/2026/09/14/huawei-pitches-near-packaged-optics-as-co-packaged-costs-bite/5296164 NPO , it cuts power consumption by over 550 kilowatts, compared to the 48,000 800G optical modules that would traditionally be required to connect all the NPUs while doubling the system's fault-free operating time, achieving 99.8 percent system availability. The SuperPoD has a multi-tier KV caching scheme, with the M900 providing a petabyte-scale KV cache for the L3.5 layer, delivering up to 4 PB of shared L3.5 KV cache capacity and 40 TB/sec of aggregate bandwidth per cluster, for the L3.5 layer, via optical networking.This provides TBs of KV cache capacity per NPU. Huawei claims that, in typical AI programming scenarios, this architecture doubles the inference cluster's token throughput and halves the time to first token TTFT . It says the 40 TB/sec bandwidth is 1.5 times higher than peer solutions, without naming them. We understand it’s generically referring to DDN, Everpure, IBM Storage Scale , MinIO and VAST Data systems with cluster scale bandwidth in the 10 - 25 TB/sec area. The M900 has an integrated architecture featuring the CPU, network controller and NAND controller in one unit, giving SuperPoD NPUs a direct, one-hop connection to the SSDs. This cuts access latency from milliseconds to about 60 microseconds. David Wang, Deputy Chairman of Huaewi’s Board and Rotating Chairman, said: “OceanStor M900 also uses hybrid media and an optimized retention algorithm, extending SSD read/write lifespan by 16-fold. This ensures a higher KV cache hit rate alongside long-term stability and reliability from the ground up.” In a little bit more detail, Huawei says the M9000 has KV-aware adaptive storage technology which predicts the expected lifetime and value of each piece of KV cache data. Based on that prediction, it schedules and places the data across different media tiers; on-chip memory, DRAM, and SSDs. It claims this optimized placement and retention strategy allows the SSDs to handle up to 24 drive writes per day DWPD , a surprisingly high number, and, as a result, SSD endurance is claimed to increase by 16 times, supporting three years of stable operation with fewer drive replacements. It is too early for Huawei to publish M900 data sheets and technology backgrounders, so we know nothing about the node rack unit size, controllers, drive count and capacity, cluster node count or other details. NPU Bootnote NPUs are Huawei Ascend Neural Processing Units which are, Huawei says, designed from the ground up for AI unlike GPUs that started as graphics chips . they use Huawei’s Da Vinci architecture with Cube matrix cores and Vector cores, and support low-precision formats like FP8 and FP4 for faster inference and training of large models. When combined into very large systems via Huawei’s UnifiedBus all-optical interconnect, they can scale to thousands or even hundreds of thousands of NPUs working as one logical machine. Recent products like the Atlas 350 using Ascend 950PR and upcoming Atlas 960 SuperPoDs are positioned as alternatives to Nvidia GPUs in the Chinese market.