{"slug": "mlperf-storage-benchmark-updated-for-modern-ai", "title": "MLPerf Storage benchmark updated for modern AI", "summary": "MLCommons released version 3.0 of its MLPerf Storage benchmark, adding an S3 data layer, KV Cache, and vector database workload tests to better reflect modern AI data systems. The update includes four workloads: training with Unet3D and RetinaNet, checkpointing with Llama 3-style models up to 1,250 billion parameters, a VDB vector database test using a 1 million vector Milvus database, and a KVcache test. Nineteen suppliers submitted results, with notable absences including DDN, Dell, Huawei, IBM, NetApp, VAST Data, and WEKA.", "body_md": "# MLPerf Storage benchmark updated for modern AI\n\nV3.0 of the [MLPerf Storage benchmark](https://www.blocksandfiles.com/ai-ml/2025/08/05/arrays-speed-up-for-mlperf-storage-benchmark-v20/101034) adds an S3 data layer to the existing POSIX one, plus KV Cache and vector database workload tests to better reflect the latest AI data systems needs.\n\nThe MLPerf Storage benchmark, developed by [MLCommons](https://mlcommons.org/), originally measured the performance of storage systems for machine learning (ML) workloads, using Unet3D, Cosmoflow, and Resnet50 AI training tasks, with two types of GPU: Nvidia A100 and H100. V3.0 tests restructures the whole thing and looks at four workloads:\n\nTraining with Unet3D and RetinaNet, looking for the maximum number of accelerators with >90 percent utilization, and simulated with Nvidia B200 or AD MI355 GPUs\n\nCheckpointing with 8, 70, 405 and 1,250 billion parameter Llama 3-style models, looking to minimize the duration of 10x writes ad then then 10x reads,\n\nVDB vector database using a 1 million vector @1,536 dimensions Milvus database, looking to maximize queries/s while reporting latency & recall percentage,\n\nKVcache looking to maximize the number of supported conversations\n\nBrian Belgodere, MLPerf Storage working group co-chair, said: “These new additions to the benchmark suite round out the test collection, covering a larger range of AI inference workloads that drive storage needs. On the training side, we have training and checkpointing tests, and on the inference side we now have KV cache and vector database tests. Including tests that decompose monolithic AI systems and focus on specific storage uses and patterns, such as checkpointing, KV caching and vector databases, gives stakeholders a much clearer idea of how to engineer and provision AI systems to minimize storage performance bottlenecks.”\n\nThe new S3 object storage layer is supported for training, checkpointing, and some VDB tests in the version 3.0 benchmark suite. Approximately one-sixth of the total submissions in this round utilized the S3 storage access layer.\n\nCurtis Anderson, MLPerf Storage working group co-chair, said: “As the scale of AI contexts reaches into the trillions, we expect object-based storage systems to emerge as a viable – and possibly preferred – alternative to filesystem-based storage. By enabling S3 support now, we are ensuring that stakeholders will have the performance information they need to make smart decisions.”\n\nNineteen suppliers ran tests and submitted results to this round of the benchmark.\n\nAbsent AI storage-relevant suppliers include DDN, Dell, Huawei, IBM, NetApp, VAST Data and WEKA.\n\nThe MLPerf Storage people say: “Traditional NAS is thin:19 orgs present; legacy enterprise incumbents are mostly missing.”\n\nReflecting the complexity of the AI storage field, there is no one single result. Instead there are four, one for each workload, with several parameters included. However, there is a single results spreadsheet with four workloads, taking up approximately 145 rows and 55 columns, almost 8,000 cells, although many are empty. There are 12 columns just to cover the KVcache workload:\n\nKVCache llama3.1-8b Storage Only - Throughput (tok/s)\n\nKVCache llama3.1-8b Storage Only - Read B/W (GiB/s)\n\nKVCache llama3.1-8b Storage Only - Write B/W (GiB/s)\n\nKVCache llama3.1-8b Storage Only - P95 Read Latency (ms)\n\nKVCache llama3.1-8b Storage + Mem - Throughput (tok/s)\n\nKVCache llama3.1-8b Storage + Mem - Read B/W (GiB/s)\n\nKVCache llama3.1-8b Storage + Mem - Write B/W (GiB/s)\n\nKVCache llama3.1-8b Storage + Mem - P95 Read Latency (ms)\n\nKVCache llama3.1-70b Storage Only - Throughput (tok/s)\n\nKVCache llama3.1-70b Storage Only - Read B/W (GiB/s)\n\nKVCache llama3.1-70b Storage Only - Write B/W (GiB/s)\n\nKVCache llama3.1-70b Storage Only - P95 Read Latency (ms)\n\nAlso, MLCommons suggests results can be normalized to performance per rack unit and also performance per watt as a power-efficiency measure.\n\nAll-in-all this is a highly complex benchmark suite with an intricate set of results and suppliers being selective over their submitted workload results. For example, only NewFW, Samsung, Suzhou Zishan and TTA submitted VDB workload results.\n\nEverpure said its FlashBlade//EXA ranks number one across the 405B parameter and 1.25T parameter model checkpointing and KV cache performance categories. FlashBlade//EXA delivered 877.52 GiB/s write bandwidth (17.74s duration) and 588.28 GiB/s of read bandwidth (28.99s duration) using 30 data nodes for a 1.25T parameter model. This was for 1,024 simulated accelerators. Performance scaled predictably from 327.65 GiB/s of write throughput with 10 data nodes to 877.52 GiB/s with 30 data nodes.\n\nCheck out the v3.0 MLPerf Storage benchmark results [here](https://mlcommons.org/benchmarks/storage/) and get overview supplier test information [here](https://mlcommons.org/wp-content/uploads/2026/08/Final-MLPerf-Storage-v3.0-Supplemental-UNDER-EMBARGO-UNTIL-9_1_26-8_00-AM-PT.pdf).\n\nBootnote\n\nOn-premises submissions for the checkpointing write test achieved a median rate of 14 GB/s/watt, with a maximum of 201. Similarly, submissions for the UNet3D read test achieved a median of 34 GB/s/watt, with a maximum of 277.", "url": "https://wpnews.pro/news/mlperf-storage-benchmark-updated-for-modern-ai", "canonical_source": "https://www.blocksandfiles.com/ai-ml/2026/09/01/mlperf-storage-benchmark-updated-for-modern-ai/5293617", "published_at": "2026-09-01 14:00:00+00:00", "updated_at": "2026-09-01 14:54:51.092975+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-research", "ai-tools"], "entities": ["MLCommons", "MLPerf Storage", "Brian Belgodere", "Curtis Anderson", "Nvidia", "Milvus", "Llama 3"], "alternates": {"html": "https://wpnews.pro/news/mlperf-storage-benchmark-updated-for-modern-ai", "markdown": "https://wpnews.pro/news/mlperf-storage-benchmark-updated-for-modern-ai.md", "text": "https://wpnews.pro/news/mlperf-storage-benchmark-updated-for-modern-ai.txt", "jsonld": "https://wpnews.pro/news/mlperf-storage-benchmark-updated-for-modern-ai.jsonld"}}