cd /news/ai-infrastructure/multi-tier-storage-rewrites-the-econ… · home topics ai-infrastructure article
[ARTICLE · art-92495] src=siliconangle.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Multi-tier storage rewrites the economics of AI inference

Super Micro Computer Inc. and partners including Intel Corp., Samsung Semiconductor Inc., Western Digital Corp., WekaIO Inc., and Scality Inc. are promoting multi-tier storage architectures that combine flash, object storage, and disk-based capacity to cut AI inference costs and boost GPU productivity. Paul McLeod, product director of storage at Supermicro, said software-defined partners are tuning products to store agentic key-value caches for longer periods, while Intel's QuickAssist Technology moves compression and encryption into hardware to improve CPU efficiency and power savings.

read6 min views1 publishedAug 11, 2026
Multi-tier storage rewrites the economics of AI inference
Image: Siliconangle (auto-discovered)

Multi-tier storage rewrites the economics of AI inference

As inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance.

These architectures combine flash, object storage and disk-based capacity tiers, enabling enterprises to serve training and inference workflows while maximizing GPU productivity and economic savings. Super Micro Computer Inc. has collaborated with its partners to address a basic issue: how to efficiently satisfy the need for AI agents to access key-value, or KV, cache where workflow data is stored.

“What we really see out in the market are monolithic massive solutions to these problems, and as they get smaller, those problems are different,” said Paul McLeod (pictured, top right), product director of storage at Supermicro. “Our software-defined partners have been great at creatively tuning their products to fit some very specific key areas. That’s where we see huge growth to take care of all the agentic KVs that people are trying to store for longer periods of time to bring back into the AI.”

McLeod spoke with theCUBE Research’s Rob Strechay for the Supermicro Open Storage Summit interview series, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. He was joined by Angela Gill (bottom, left), director of partner enablement at Intel Corp.; Jonathan Prout (top, left), director of memory business development at Samsung Semiconductor Inc.; Marc Tanguay (bottom, center), senior product marketing manager for HDDs at Western Digital Corp.; Anthony Lembo (bottom, right), vice president of global systems engineering at WekaIO Inc.; and Greg DiFraia (top, center), senior vice president of AI and alliance partnerships at Scality Inc.

They discussed how enterprises can optimize AI infrastructure economics, where different storage tiers fit into training and inference workflows, and how to maximize GPU productivity without simply buying more hardware. ( Disclosure below.)*

New multi-tier storage solutions

A central focus of the panel was why storage tiering makes sense today because not every AI workload is equal. Agents create a wide range of data demands across different types of storage, such as file, object and flash, as well as GPU and CPU platforms.

Intel has responded to this new reality with solutions such as its QuickAssist Technology, a hardware accelerator built into select Intel processors.

“With technologies like Intel’s QuickAssist, we can move compression and encryption out of software and into hardware,” Gill noted. “That translates directly into more usable CPU cycles … lower latency, and importantly, better power efficiency at scale. Storage is no longer in the slow lane. It has to move at compute speed.”

For storage to move at compute speed, it must be able to adapt to changing processing demands. At Western Digital, this involves retooling its hard disk drive, or HDD, Ultrastar portfolio to provide the capacity and performance needed for high-intensity AI workloads. “Adding storage capacity is not as simple as just adding another rack,” said Tanguay. “We need to work within existing footprints. With Ultrastar, we’re up to our 28 terabyte CMR hard drive, which is the drive that’s been qualified and tested to work with all of Supermicro’s products. We’re moving up to 30 terabytes in our Ultrastar CMR technology hard drives before the end of this year.”

Maximizing GPU resources

Supermicro’s collaboration with its software-defined storage partners has focused on finding innovative solutions to maximize GPU utilization. Scality integrates GPU-direct storage access into its platform by combining high-speed flash media with S3. This enables AI training and inference pipelines to stream data straight from object storage into GPU memory.

“We need to drive utilization up has high as it can go, and in order to do that you’ve got to really think about things differently,” DiFraia said. “We’re doing things on the extreme hot tier with S3 storage and GPU-direct and S3 over RDMA, hot tier, warm tier and even cold tier. Because when we look at the topology for customers, it’s that the life cycle, when you’re talking about tens or hundreds of petabytes or even exabytes, it’s not all going to live in flash.”

The memory required to process KV cache requirements has occasionally exceeded available GPU memory and system resources, creating a bottleneck for AI inferencing. To address this issue, Samsung Semiconductor has developed scalable memory expansion while maintaining GPU performance.

“We’re looking at these different requirements in the KV cache storage tier, and we’ve developed solutions to hit each of these,” Prout explained. “With local on-node performance for GPUs, we’ve developed the PM1723. This is a Gen 6 drive for extreme performance. It’s up to 28.4 gigabytes per second of sequential read throughput and up to 6.6 million random read IOPS.”

Many enterprises face a scenario in which storage for AI inference workloads requires a tradeoff between performance and cost. Addressing this tradeoff requires careful engineering and selection of the optimal storage technology to minimize data movement for each AI-driven need.

“Balancing cost and economics with performance is really tough, especially if you look at these workloads and what they do,” Lembo told theCUBE. “From the GPU server’s perspective, how can I eliminate data movement and land in one place? Operational simplicity is one of the really underrated factors that’s sometimes not considered when you’re trying to optimize infrastructure.”

Stay tuned for the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the Supermicro Open Storage Summit interview series.

( Disclosure: TheCUBE is a paid media partner for the Supermicro Open Storage Summit interview series. Neither Supermicro, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*

Photo: SiliconANGLE

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

About SiliconANGLE Media

SiliconANGLE,

theCUBE Network,

theCUBE Research,

CUBE365,

theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @super micro computer inc. 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/multi-tier-storage-r…] indexed:0 read:6min 2026-08-11 ·