# MLPerf Storage benchmark updated for modern AI

> Source: <https://www.blocksandfiles.com/ai-ml/2026/09/01/mlperf-storage-benchmark-updated-for-modern-ai/5293617>
> Published: 2026-09-01 14:00:00+00:00

# MLPerf Storage benchmark updated for modern AI

V3.0 of the [MLPerf Storage benchmark](https://www.blocksandfiles.com/ai-ml/2025/08/05/arrays-speed-up-for-mlperf-storage-benchmark-v20/101034) adds an S3 data layer to the existing POSIX one, plus KV Cache and vector database workload tests to better reflect the latest AI data systems needs.

The MLPerf Storage benchmark, developed by [MLCommons](https://mlcommons.org/), originally measured the performance of storage systems for machine learning (ML) workloads, using Unet3D, Cosmoflow, and Resnet50 AI training tasks, with two types of GPU: Nvidia A100 and H100. V3.0 tests restructures the whole thing and looks at four workloads:

Training with Unet3D and RetinaNet, looking for the maximum number of accelerators with >90 percent utilization, and simulated with Nvidia B200 or AD MI355 GPUs

Checkpointing with 8, 70, 405 and 1,250 billion parameter Llama 3-style models, looking to minimize the duration of 10x writes ad then then 10x reads,

VDB vector database using a 1 million vector @1,536 dimensions Milvus database, looking to maximize queries/s while reporting latency & recall percentage,

KVcache looking to maximize the number of supported conversations

Brian Belgodere, MLPerf Storage working group co-chair, said: “These new additions to the benchmark suite round out the test collection, covering a larger range of AI inference workloads that drive storage needs. On the training side, we have training and checkpointing tests, and on the inference side we now have KV cache and vector database tests. Including tests that decompose monolithic AI systems and focus on specific storage uses and patterns, such as checkpointing, KV caching and vector databases, gives stakeholders a much clearer idea of how to engineer and provision AI systems to minimize storage performance bottlenecks.”

The new S3 object storage layer is supported for training, checkpointing, and some VDB tests in the version 3.0 benchmark suite. Approximately one-sixth of the total submissions in this round utilized the S3 storage access layer.

Curtis Anderson, MLPerf Storage working group co-chair, said: “As the scale of AI contexts reaches into the trillions, we expect object-based storage systems to emerge as a viable – and possibly preferred – alternative to filesystem-based storage. By enabling S3 support now, we are ensuring that stakeholders will have the performance information they need to make smart decisions.”

Nineteen suppliers ran tests and submitted results to this round of the benchmark.

Absent AI storage-relevant suppliers include DDN, Dell, Huawei, IBM, NetApp, VAST Data and WEKA.

The MLPerf Storage people say: “Traditional NAS is thin:19 orgs present; legacy enterprise incumbents are mostly missing.”

Reflecting the complexity of the AI storage field, there is no one single result. Instead there are four, one for each workload, with several parameters included. However, there is a single results spreadsheet with four workloads, taking up approximately 145 rows and 55 columns, almost 8,000 cells, although many are empty. There are 12 columns just to cover the KVcache workload:

KVCache llama3.1-8b Storage Only - Throughput (tok/s)

KVCache llama3.1-8b Storage Only - Read B/W (GiB/s)

KVCache llama3.1-8b Storage Only - Write B/W (GiB/s)

KVCache llama3.1-8b Storage Only - P95 Read Latency (ms)

KVCache llama3.1-8b Storage + Mem - Throughput (tok/s)

KVCache llama3.1-8b Storage + Mem - Read B/W (GiB/s)

KVCache llama3.1-8b Storage + Mem - Write B/W (GiB/s)

KVCache llama3.1-8b Storage + Mem - P95 Read Latency (ms)

KVCache llama3.1-70b Storage Only - Throughput (tok/s)

KVCache llama3.1-70b Storage Only - Read B/W (GiB/s)

KVCache llama3.1-70b Storage Only - Write B/W (GiB/s)

KVCache llama3.1-70b Storage Only - P95 Read Latency (ms)

Also, MLCommons suggests results can be normalized to performance per rack unit and also performance per watt as a power-efficiency measure.

All-in-all this is a highly complex benchmark suite with an intricate set of results and suppliers being selective over their submitted workload results. For example, only NewFW, Samsung, Suzhou Zishan and TTA submitted VDB workload results.

Everpure said its FlashBlade//EXA ranks number one across the 405B parameter and 1.25T parameter model checkpointing and KV cache performance categories. FlashBlade//EXA delivered 877.52 GiB/s write bandwidth (17.74s duration) and 588.28 GiB/s of read bandwidth (28.99s duration) using 30 data nodes for a 1.25T parameter model. This was for 1,024 simulated accelerators. Performance scaled predictably from 327.65 GiB/s of write throughput with 10 data nodes to 877.52 GiB/s with 30 data nodes.

Check out the v3.0 MLPerf Storage benchmark results [here](https://mlcommons.org/benchmarks/storage/) and get overview supplier test information [here](https://mlcommons.org/wp-content/uploads/2026/08/Final-MLPerf-Storage-v3.0-Supplemental-UNDER-EMBARGO-UNTIL-9_1_26-8_00-AM-PT.pdf).

Bootnote

On-premises submissions for the checkpointing write test achieved a median rate of 14 GB/s/watt, with a maximum of 201. Similarly, submissions for the UNet3D read test achieved a median of 34 GB/s/watt, with a maximum of 277.
