{"slug": "multi-tier-storage-rewrites-the-economics-of-ai-inference", "title": "Multi-tier storage rewrites the economics of AI inference", "summary": "Super Micro Computer Inc. and partners including Intel Corp., Samsung Semiconductor Inc., Western Digital Corp., WekaIO Inc., and Scality Inc. are promoting multi-tier storage architectures that combine flash, object storage, and disk-based capacity to cut AI inference costs and boost GPU productivity. Paul McLeod, product director of storage at Supermicro, said software-defined partners are tuning products to store agentic key-value caches for longer periods, while Intel's QuickAssist Technology moves compression and encryption into hardware to improve CPU efficiency and power savings.", "body_md": "### Multi-tier storage rewrites the economics of AI inference\n\nAs inference becomes the dominant workload in AI infrastructure, multi-tier storage architectures are emerging as a key method for cost control and enhanced performance.\n\nThese [architectures](https://siliconangle.com/2026/08/05/storage-architecture-supermicro-open-storage-summit-supermicroopenstoragesummit/) combine flash, object storage and disk-based capacity tiers, enabling enterprises to serve training and inference workflows while maximizing GPU productivity and economic savings. Super Micro Computer Inc. has collaborated with its partners to address a basic issue: how to efficiently satisfy the need for AI agents to access key-value, or KV, cache where workflow data is stored.\n\n“What we really see out in the market are monolithic massive solutions to these problems, and as they get smaller, those problems are different,” said [Paul McLeod](https://www.linkedin.com/in/paul-mcleod-2230662/) (pictured, top right), product director of storage at Supermicro. “Our software-defined partners have been great at creatively tuning their products to fit some very specific key areas. That’s where we see huge growth to take care of all the agentic KVs that people are trying to store for longer periods of time to bring back into the AI.”\n\nMcLeod spoke with theCUBE Research’s [Rob Strechay](https://www.linkedin.com/in/robstrechay/) for the [Supermicro Open Storage Summit interview series](https://www.thecube.net/events/supermicro/open-storage-summit-2026?utm_source=siliconangle&utm_medium=article&utm_campaign=supermicro_open_storage_summit_2026), during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. He was joined by Angela Gill (bottom, left), director of partner enablement at Intel Corp.; [Jonathan Prout](https://www.linkedin.com/in/jonathanprout/) (top, left), director of memory business development at Samsung Semiconductor Inc.; [Marc Tanguay](https://www.linkedin.com/in/marctanguay/) (bottom, center), senior product marketing manager for HDDs at Western Digital Corp.; [Anthony Lembo](https://www.linkedin.com/in/lemboaw/) (bottom, right), vice president of global systems engineering at WekaIO Inc.; and [Greg DiFraia](https://www.linkedin.com/in/greg-difraia-7914704/) (top, center), senior vice president of AI and alliance partnerships at Scality Inc.\n\nThey discussed how enterprises can optimize AI infrastructure economics, where different storage tiers fit into training and inference workflows, and how to maximize GPU productivity without simply buying more hardware. *(* Disclosure below.)*\n\n### New multi-tier storage solutions\n\nA central focus of the panel was why storage tiering makes sense today because not every AI workload is equal. Agents create a wide range of data demands across different types of storage, such as file, object and flash, as well as GPU and CPU platforms.\n\nIntel has responded to this new reality with solutions such as its [QuickAssist Technology](https://www.intel.com/content/www/us/en/products/details/processors/xeon/features/quick-assist-technology.html), a hardware accelerator built into select Intel processors.\n\n“With technologies like Intel’s QuickAssist, we can move compression and encryption out of software and into hardware,” Gill noted. “That translates directly into more usable CPU cycles … lower latency, and importantly, better power efficiency at scale. Storage is no longer in the slow lane. It has to move at compute speed.”\n\nFor storage to move at compute speed, it must be able to adapt to changing processing demands. At Western Digital, this involves retooling its hard disk drive, or HDD, [Ultrastar](https://www.westerndigital.com/products/data-center-storage) portfolio to provide the capacity and performance needed for high-intensity AI workloads.\n\n“Adding storage capacity is not as simple as just adding another rack,” said Tanguay. “We need to work within existing footprints. With Ultrastar, we’re up to our 28 terabyte CMR hard drive, which is the drive that’s been qualified and tested to work with all of Supermicro’s products. We’re moving up to 30 terabytes in our Ultrastar CMR technology hard drives before the end of this year.”\n\n### Maximizing GPU resources\n\nSupermicro’s collaboration with its software-defined storage partners has focused on finding innovative solutions to maximize GPU utilization. Scality [integrates](https://www.solved.scality.com/gpu-direct-storage/) GPU-direct storage access into its platform by combining high-speed flash media with S3. This enables AI training and inference pipelines to stream data straight from object storage into GPU memory.\n\n“We need to drive utilization up has high as it can go, and in order to do that you’ve got to really think about things differently,” DiFraia said. “We’re doing things on the extreme hot tier with S3 storage and GPU-direct and S3 over RDMA, hot tier, warm tier and even cold tier. Because when we look at the topology for customers, it’s that the life cycle, when you’re talking about tens or hundreds of petabytes or even exabytes, it’s not all going to live in flash.”\n\nThe memory required to process KV cache requirements has occasionally exceeded available GPU memory and system resources, creating a bottleneck for AI inferencing. To address this issue, Samsung Semiconductor has developed scalable [memory expansion](https://semiconductor.samsung.com/news-events/tech-blog/breaking-ai-memory-limits-with-cxl-memory-pooling/) while maintaining GPU performance.\n\n“We’re looking at these different requirements in the KV cache storage tier, and we’ve developed solutions to hit each of these,” Prout explained. “With local on-node performance for GPUs, we’ve developed the PM1723. This is a Gen 6 drive for extreme performance. It’s up to 28.4 gigabytes per second of sequential read throughput and up to 6.6 million random read IOPS.”\n\nMany enterprises face a scenario in which storage for AI inference workloads requires a tradeoff between performance and cost. Addressing this tradeoff requires careful engineering and selection of the optimal storage technology to minimize data movement for each AI-driven need.\n\n“Balancing cost and economics with performance is really tough, especially if you look at these workloads and what they do,” Lembo told theCUBE. “From the GPU server’s perspective, how can I eliminate data movement and land in one place? Operational simplicity is one of the really underrated factors that’s sometimes not considered when you’re trying to optimize infrastructure.”\n\nStay tuned for the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the [Supermicro Open Storage Summit interview series](https://www.thecube.net/events/supermicro/open-storage-summit-2026?utm_source=siliconangle&utm_medium=article&utm_campaign=supermicro_open_storage_summit_2026).\n\n*(* Disclosure: TheCUBE is a paid media partner for the Supermicro Open Storage Summit interview series. Neither Supermicro, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*\n\n##### Photo: SiliconANGLE\n\n# A message from John Furrier, co-founder of SiliconANGLE:\n\nSupport our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.\n\n**15M+ viewers of theCUBE videos**, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.\n\n# Are you AWS customer? Support SiliconANGLE Financially by buying your AWS services from our Marketplace portal page and links.\n\n**About SiliconANGLE Media**\n\n[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),\n\n[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),\n\n[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),\n\n[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),\n\n[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.\n\nFounded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.", "url": "https://wpnews.pro/news/multi-tier-storage-rewrites-the-economics-of-ai-inference", "canonical_source": "https://siliconangle.com/2026/08/11/multi-tier-storage-solutions-optimize-ai-supermicroopenstoragesummit/", "published_at": "2026-08-11 18:36:46+00:00", "updated_at": "2026-08-11 19:05:23.431292+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-products"], "entities": ["Super Micro Computer Inc.", "Intel Corp.", "Samsung Semiconductor Inc.", "Western Digital Corp.", "WekaIO Inc.", "Scality Inc.", "Paul McLeod", "QuickAssist Technology"], "alternates": {"html": "https://wpnews.pro/news/multi-tier-storage-rewrites-the-economics-of-ai-inference", "markdown": "https://wpnews.pro/news/multi-tier-storage-rewrites-the-economics-of-ai-inference.md", "text": "https://wpnews.pro/news/multi-tier-storage-rewrites-the-economics-of-ai-inference.txt", "jsonld": "https://wpnews.pro/news/multi-tier-storage-rewrites-the-economics-of-ai-inference.jsonld"}}