MinIO blogs about object caching in AI and says you don’t need a cache if your object storage is fast enough. Blogger Daniel Valdivia, a MinIO architect, compares a CoreWeave LOTA (Local Object Transport Accelerator) caching proxy benchmark with MinIO AIStor object storage performance. CoreWeave ran Warp, an open benchmark test, object data on on 20 GPU nodes, each with 8 GPUs, 1 TiB of LOTA cache per node, and a dedicated network adapter with dual 100 Gbps links for storage access, and achieved S3 GET performance of 18.4 GiB/s per node (19.8 GB/s) . MiniO’s AIStor, using Solidigm QLC SSDS, with no caching achieved c33.5 GiB/s per node (36 GB/s) in a separate benchmark.
Valdiva says: “An erasure-coded object store on capacity-optimized QLC drives, running over plain TCP, delivers more throughput per node than a warm cache on local NVMe (about 33.5 GiB/s versus 18.4 GiB/s), and more than 11 times the cluster throughput of CoreWeave's own cold path. A cache should be faster than what's behind it. When a durable store with parity beats the cache, the architecture is telling you something.”
He explains: “A cache exists because the thing behind it is slow. When your object store is fast, you don't need to make complicated trade-offs to keep GPUs fed. …you shouldn't need a cache on every GPU node to make your object store fast in the first place.”
MinIO is not saying you don't need KV caching systems: “An object cache speeds up S3 GETs of datasets and checkpoints. An inference KV cache stores attention state so the engine doesn't recompute a long prompt every time a user comes back. They have different access patterns, block sizes, and success metrics.” They are two different caching situations.
MinIO supports KV caches and has developed its own MemKV, a shared-nothing KV store on raw NVMe, wired into inference engines like vLLM.
In effect, CoreWeave wouldn't need LOTA if it used MinIO's AIStor instead.