AI is storage’s biggest opportunity - and biggest threat AI is transforming the IT storage industry, creating massive demand for faster data access and protection while introducing risks from AI-driven mishaps and attacks, according to Blocks & Files editor Chris Mellor. The post-ChatGPT era has driven growth for vendors like DDN, Pure Storage, VAST Data, and WEKA, while legacy monolithic arrays from Hitachi Vantara, IBM, and Infinidat are being left behind. Mellor identifies five key intersections of storage and AI, including data provisioning, protection, cyber-resilience, operational use, and defense against AI attacks. AI is storage’s biggest opportunity - and biggest threat Faster access and more secure recoveries are driving business, but mishaps and attacks can endanger data Chris MellorChrisMellorSTORAGE EDITORBlocks & Files editor Published An AI revolution is sweeping through the IT storage world, providing a massively beneficial environment requiring more data to be stored and delivered to AI models and agents, more data to be protected, more data access governance and much better storage operating environments. The downside is that AI can run amok, with data mishaps and deliberate agent-enhanced attacks. We are using AI to make things better and need AI to prevent AI itself from making things worse. Forty-four months ago, when ChatGPT was released, the storage world began an irreversible migration into the AI era. The technology developments needed to provide fast data access to the favorite AI processor, the GPU, revolutionized the NAND and SSD suppliers, the flash array hardware and software vendors and the HPC/supercomputing world. The old and relatively steady, pre-ChatGPT era Dell, HPE, and NetApp-dominated enterprise storage array business met a set of new vendors growing fast, coming from the all-flash array and HPC worlds. Vendors such as DDN, Pure Storage, VAST Data, and WEKA grew rapidly as parallel data access became a key software technology, alongside disaggregated storage array designs, increased use of unstructured data, the rapid growth of analytics, and the emergence of AI-focused data lakes such as Databricks and Snowflake.A whole new public cloud sector emerged: the GPU-as-a-Service neoclouds, such as CoreWeave and Lambda. AI training dominated the early AI storage days but is now being overtaken by AI inference; production AI, with enterprises deploying AI factories to build, tune, deploy, and run their own and third-party AI agents. Model Context Protocol MCP and graph technology enable digital employees to act and reason, access and change data, both structured and unstructured. They can make mistakes, meaning that their activities have to be recorded so that they can be reversed if they take a wrong track. These digital employees have to be governed with Agent Identity Access Management. Stepping back for a moment, we can see storage and AI meet in five places: Storage providing data for AI Protecting AI data and actions Storage cyber-resilience extended to govern AI data access Storage using AI in its own operations Storage being protected from AI-driven attacks. Data for AI The dominant GPU vendor, Nvidia, has eagerly supported storage delivery technologies to keep its GPUs running and not being IO-bound. Feeding them data from parallel file systems using disk drive arrays was not good enough. The disks were replaced by SSDs, with NVMe and PCIe interconnects replacing the disk drive era’s SAS and SATA protocols. The transfer of data from a storage array’s x86 controller and its DRAM to the GPU server’s CPU+DRAM subsystem and then to the GPU’s capacity-limited high-bandwidth memory HBM was too slow. GPUDirect cut out the storage array controller and its memory from the data flow, with RDMA access from the SSDs. This was applied to files first and then to objects with S3 over RDMA-type technologies. Cloudian, MinIO and Scality have been active here. KV caching schemes have been set up to logically extend a GPU’s HBM to SSDs, reduce HBM data load waits and avoid token recomputation. Nvidia has partnered with the main enterprise storage array vendors to make this happen. Disaggregated storage array DASE technology, pioneered by VAST Data, has been adopted by Dell, Everpure, HPE and NetApp. Monolithic, high-end storage arrays from Hitachi Vantara, IBM, and Lenovo-acquired Infinidat have not adopted DASE, nor GPUDirect and not KV caching. Their architecture precludes them from joining in and they are becoming the storage equivalent of mainframes, a left-behind but still necessary niche. AI models process tokens, which are turned into vector embedding data that needs storing and searching. Dedicated vector database suppliers have sprung up, such as Pinecone, Qdrant, Weaviate, and Zilliz, while multi-model OLAP and OLTP databases have added vector support, SingleStore being an example. AI data pipeline technologies are being developed to enable an organization’s entire data estate to be used as an AI data source, but without copying it to a single repository. Apache Iceberg is being used to give data lakes access to external storage and logically bring it into the data lake’s namespace. The monolithic arrays are, of course, data sources for AI pipelines and will contribute. Data management suppliers doubled down on initiatives to map organizations’ data and make it available, via classification, selection, filtering and metadata access so as to protect privileged data, reduce bulk data movements, and accelerate targeted data movement; think Arcitecta, Datadobi, Hammerspace and Komprise. An advantage of DASE array architecture is that storage is decoupled from compute, and the compute can be scaled up to run AI technologies directly on the array. This means that array vendors’ software stacks can be extended upwards into AI data pipelines. They can have the capability to provide services for AI agents and become, in effect, AI operating systems. Indeed VAST Data explicitly calls its extended SW stack its AIOS. Protecting AI data and actions Backup and cyber-resilience vendors, such as Cohesity, Commvault, Druva, Rubrik, and Veeam, have recognized that their backups represent a great data source for AI models and agents. They developed internal AI agents, such as Cohesity’s Gaia, its Gen AI search assistant, to build on this idea. The backup vendors see that their backups contain a temporal record of data changes but they, the changes, have been made by different entities in an organization’s IT estate and are not correlated. AI agents can monitor the backups and find links between events that indicate an existing or developing cyber attack. They can identify the start of an attack, the affected data, the last known good copy of that data, and restore it, helping with cyber-attack response and recovery. Cohesity, Rubrik, and others are promoting cyber-attack simulations where they do this with customer execs and even board members to show them how a cyber-attack needs a cross-business organization response to be effective. Their ability to feed AI data pipelines is growing. AI can be used to search for data relevant to an upstream enquiry, detect and filter out or mask PII data, and then move it to the requesting target. Veeam is going further and starting to combine production data with its backup data for AI purposes. AI cyber resilience The main backup vendors have all become cyber-resilience suppliers and provide some form of identity access management IAM plus related data access monitoring. IAM has been a human-centered activity but is rapidly embracing AI agent data and resource access as well. Consider an AI agent to be a digital employee and the need for agent-focused IAM becomes immediately obvious. Then realize that your digital employees operate inside your IT systems and not outside. Unlike human employees, they operate inside IT systems rather than sitting at a keyboard typing commands or clicking GUI buttons. It gets worse. As well as operating at speeds much, much faster than human employees, organizations will likely have thousands, even tens of thousands of, agents, so the likelihood of accidental data mishaps will be many times higher, and the mishap can take place in seconds. The only way to provide agentic IAM would appear to be to use agents. Cohesity, Druva, HYCU, Rubrik and Veeam are active in this space; using AI agents to provide governance for AI agents. We are also seeing the rise of agent activity recording, with storage of temporary files, so that misplaced agent activity, either accidental or deliberate, can be identified and fixed. This is a fast-developing field and the risks and dangers are still being identified. AI enhancing storage internal operations We have seen storage array vendors using machine learning to monitor array telemetry and identify faults for some time. HPE-acquired Nimble was one of the first vendors to do this and it is now commonplace. Modern AI can do more. Storage and data protection admin can be highly complex. One example is a Veeam upgrade issue. An obvious use of AI chatbots is to have them be the interface between an administrator, using natural language, and the data protection software. The chatbot, or agent, is trained on the data protection software’s features and telemetry and can function as a highly-skilled diagnostic engineer helping admin staff. This, we think, will spread like wildfire, as it will effectively up-skill admin staff. They can operate at a higher level, diagnosing data protection gaps, adding new applications to be protected or applications to be given data access, without having to go down into the weeds and set up each individual action according to policies in a manual. In effect, admin staff can “talk” to their data or arrays, manage a fleet of arrays, institute new policies, optimize operations to balance performance, cost and electricity usage more effectively by using AI. We are going to see AI user interfaces being developed, complementing the existing CLIs and GUIs. The big issue AI agents are intangible, invisible and make decisions at microsecond speeds. We human operators are tangible, visible and make decisions slowly. It’s easy for computers to monitor us but it’s darn near impossible for us to monitor agents. We will have to use agents to monitor agents, and that means we have to trust our monitoring agents, really trust them. They have to be immutable, immune to having their identities stolen, and capable of detecting and stopping rogue agents in their tracks. Imagine a North Korea cyber-hacking group developing their own swarm of attack agents. We need to be ready for this. The wolves will be coming. ®