{"slug": "weka-s-big-beast-ai-storage-products", "title": "WEKA's big beast AI storage products", "summary": "WEKA announced an exabyte-scale WEKAPod all-flash array running its sixth-generation NeuralMesh software, claiming 1.1 exabytes of effective capacity in a single 56-RU rack. The company says the system delivers 10x higher token throughput and 7x more tokens per GPU on Oracle Cloud Infrastructure compared to Nvidia CMX, positioning it against high-end AI storage arrays from DDN, Everpure, NetApp, and VAST Data.", "body_md": "# WEKA's big beast AI storage products\n\nWEKA has announced an exabyte-scale WEKAPod all-flash array for its sixth generation NeuralMesh software to feed and receive massive amounts of data for production-scale, accelerate computing AI training and inference workloads. This WEKApod places WEKA in the same high-end AI storage array camp as systems from DDN, Everpure, NetApp and VAST Data but with claimed class-leading density; 1.1 exabyte effective capacity in a 56 RU rack.\n\nThe NeuralMesh v6 software provides native hyperscale multi-tenancy, a unified file-and-S3-object protocol stack built on NVMe, intelligent data mobility with replication and remote caching support, always-on data reduction with performance guarantees, Kubernetes-native operations, and unified observability. WEKA says NeuralMesh is used on [Oracle Cloud Infrastructure](https://www.oracle.com/cloud/) (OCI) where, with its Augmented Memory Grid capability, benchmarks on OCI H100 infrastructure have demonstrated 10x higher token throughput, 10x more concurrent users served, and 7x more tokens per GPU in production deployments, compared to Nvidia CMX. CoreWeave gets 4.2x more tokens per GPU, and up to 6x TTFT reduction with NeuralMesh compared to CMX.\n\nLiran Zvibel, WEKA’s co-founder and CEO, said: “AI infrastructure built for traditional workloads cannot power the inference era. The constraints are different: rack space, energy, supply chain, and operational density now determine whether AI deployments produce margin or destroy it.\n\n“We built WEKApod to operate inside those constraints rather than against them, on hardware WEKA designed and engineered specifically for inference-era computing, with a supply chain we control. It delivers upfront economics that compound with the best total cost of ownership across the full gamut of AI infrastructure requirements, establishing a new scorecard that we believe every customer running production AI at scale will soon evaluate their infrastructure decisions against.”\n\n**WEKApod**\n\nThe [WEKAPod appliance](https://www.blocksandfiles.com/ai-ml/2025/11/19/wekas-new-appliances-can-run-its-gpu-memory-wall-busting-software/1718977) is a storage server appliance storing and pushing data out to GPU servers such as Nvidia’s SuperPOD. It consists of pre-configured hardware + WEKA Data Platform SW storage nodes, and there were two models in late 2025; PCIe gen 4-using Prime and faster PCIe gen 5-using Nitro. Their AlloyFlash feature combined fast and high-capacity SSD storage tiers. The appliances are built to run WEKA’s NeuralMesh software and provide production-class AI inferencing support.\n\nThis is the third generation of the WEKApod and there are updated Nitro, Prime and new Prime Max configurations using a 2RU chassis.\n\nPrime MAX is the highest capacity system, pushing past 1 EB per rack.\n\nPrime is the optimized price-performance system, with four servers per 2U chassis, 70 NVMe drives that combine QLC and TLC in a single system, all managed by AlloyFlash.\n\nNitro provides extreme performance forAI factories with training deployments and GPU-as-a-service environments running hundreds or thousands of GPUs.\n\nA single WEKApod system can deliver 1.1 exabytes of effective capacity in a single rack, with NeuralMesh data reduction running on a hardware foundation of 441.5 PB raw capacity, making it the first single-rack system to break the exabyte barrier. Throughput reaches 10.2 TB/s per rack. IOPS reach 210 million per rack.\n\nThe new WEKAPods have a PCIe gen 6 internal fabric, a cable-based, backplane-free drive interconnect, NvidiaConnectX SuperNIC networking enabling Spectrum-X Ethernet data plane connectivity, and a software-managed thermal architecture.\n\nWEKA says the system is “built to reduce the datacenter footprint, reclaim critical rack space, and improve tokens per watt, without sacrificing token throughput or workload performance. Per rack unit, WEKApod leads every publicly available alternative in capacity and throughput density, by 267 percent in effective capacity and 114 percent in throughput, over the next-best system in each category.”\n\nThe company says that it controls component sourcing directly rather than inheriting OEM channel dependencies, giving AI cloud providers and enterprises the pricing stability and lead-time predictability that infrastructure planning at scale actually requires.\n\nWEKApod ships factory-installed and pre-validated with the full NeuralMesh feature set, including Augmented Memory Grid, with no dedicated storage team required.\n\n**NeuralMesh 6**\n\n[NeuralMesh](https://www.blocksandfiles.com/ai-ml/2025/10/30/weka-brings-neuralmesh-ai-file-system-to-nvidia-bluefield-4-dpus/1593005) is WEKA’s parallel filesystem software. Its [Augmented Memory Grid](https://www.blocksandfiles.com/ai-ml/2025/03/18/storage-vendors-rally-behind-nvidia-at-gtc-2025/1602277?_gl=1*1sk8a3v*_ga*MTY2OTcyMjAyNS4xNzcwODg2MTMy*_ga_NSDTXHMMN0*czE3ODQ2Mjk0MzQkbzE3NiRnMSR0MTc4NDYyOTY3NyRqNjAkbDAkaDA.) enables AI models to extend GPU server memory for inferencing to the Neural Mesh’s external storage, using it as a KV Cache with microsecond latencies and multi-TBPs bandwidth, and providing up to additional petabytes of memory address space capacity.\n\nNeuralMesh 6 delivers;\n\nComposable Clusters provide full hardware-level isolation, dedicated CPU, memory, and storage drives per tenant for anchor tenants that require guaranteed resources, workload separation, and predictable performance under load.\n\nVirtual Multi-Tenancy provides VPC-like network isolation through WEKA's Virtualized RDMA Data Fabric (VRDF), supporting private VLANs, overlapping IP address spaces, per-tenant quality of service (QoS), per-tenant encryption with independent KMS, and independent LDAP/AD authentication. Virtual MT scales to more than 1,000 isolated logical tenants per cluster, with new tenant provisioning in under 30 minutes.\n\nMultiple NeuralMesh clusters share physical NVMe drives, extracting more parallelism from each device, enabling multi-tenant inference workloads on shared infrastructure.\n\nA single WEKA hardware cluster running 50 Composable Clusters can support up to 50,000 logically isolated tenants on the same physical infrastructure. Data blocks are addressable via S3 and POSIX (files) simultaneously in a single unified namespace. S3 over RDMA enables zero-copy data transfer directly to GPU memory. There can be 2,000 to 5,000 concurrent S3 connections per node, roughly five times the concurrency of conventional S3 architectures.\n\nThe AlloyFlash feature enables combining TLC and QLC NVMe flash within a single cluster, transparently routing latency-sensitive operations to TLC while leveraging QLC at roughly 30-40 percent lower cost per terabyte for bulk-capacity workloads.\n\nThe software provides always-on data reduction “with fingerprinting, similarity hashing, deduplication, and compression. It delivers less than 5 percent write overhead, up to 6x capacity savings on AI training data, and a contractual guarantee of both the data reduction ratio and the performance impact.” Customer-specific pre-sales analysis can project workload-based reduction outcomes before deployment.\n\nThe data reduction runs entirely off the write path, with background-first similarity-based compression and cross-filesystem deduplication. Writes commit at native NVMe speed regardless of capacity fill or background load. [Ironically this is similar to [ExaGrid’s](https://www.blocksandfiles.com/data-protection/2026/01/23/exagrid-launches-all-flash-backup-appliances/4090479) deduplication-free landing zone in its purpose-built backup appliances.) There is no sustained-write cliff during checkpoint-heavy training, the failure mode which inline reduction architectures hit.\n\nWEKA says NeuralMresh is being developed to support complete federation and a global namespace. V6 NeuralMesh asynchronous replication supports AI data mobility, cloudbursting, and multi-site collaboration, with metadata being replicated first, making destination environments immediately browsable; providing an “instantaneous remote caching functionality”. Data is replicated with data reduction, being hydrated on demand. This reduces WAN traffic.\n\nA NeuralMesh Kubernetes Operator automates cluster deployment and lifecycle management in Kubernetes environments, making it suitable for Neoclouds and AI research labs running Kubernetes as their operational standard. A NeuralMesh SaaS-based Observe facility provides unified multi-cluster dashboards, client-level diagnostics, intelligent alerting with configurable thresholds, and routing to Slack, PagerDuty, or email with direct links to relevant dashboards.\n\nAjay Singh, WEKA’s Chief Product Officer, said: “The infrastructure operators running production AI today have been forced to assemble platforms from vendors that were never designed to work together. Separate stacks for file and object, manual data movement between them, multi-tenancy bolted on after the fact. NeuralMesh 6 delivers what they've actually needed all along: a single platform that handles the high-performance file layer and the high-capacity object layer on the same blocks, with native multi-tenancy, intelligent data mobility, and always-on data efficiency built in from the start. This is what production inference infrastructure looks like when it's designed for the workload, not retrofitted for it.”\n\nNeuralMesh 6 will be generally available in the second half of 2026. Existing WEKA customers can upgrade to NeuralMesh 6 at no additional cost through standard upgrade channels", "url": "https://wpnews.pro/news/weka-s-big-beast-ai-storage-products", "canonical_source": "https://www.blocksandfiles.com/ai-ml/2026/07/21/wekas-big-beast-ai-storage-products/5275510", "published_at": "2026-07-21 12:00:00+00:00", "updated_at": "2026-07-21 15:56:25.771360+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-products", "ai-tools"], "entities": ["WEKA", "WEKAPod", "NeuralMesh", "Oracle Cloud Infrastructure", "Nvidia", "CoreWeave", "Liran Zvibel", "Nvidia CMX"], "alternates": {"html": "https://wpnews.pro/news/weka-s-big-beast-ai-storage-products", "markdown": "https://wpnews.pro/news/weka-s-big-beast-ai-storage-products.md", "text": "https://wpnews.pro/news/weka-s-big-beast-ai-storage-products.txt", "jsonld": "https://wpnews.pro/news/weka-s-big-beast-ai-storage-products.jsonld"}}