{"slug": "dell-adds-data-context-prep-and-storage-features-to-its-ai-data-platform", "title": "Dell Adds Data Context, Prep, And Storage Features To Its AI Data Platform", "summary": "Dell added data context, preparation, and storage features to its AI Data Platform, which the company says addresses the top challenge cited by 3,800 enterprise IT decision-makers and AI experts it surveyed: data quality, availability, management, and security. Dell senior vice president of its Infrastructure Solutions Group Varun Chhabra said \"the infrastructure is ready, but the data isn't,\" and described the platform's three layers — a Data Orchestration Engine, Data Engines, and Storage Engines including PowerScale, ObjectScale, and the Lightning File System released in April after its March unveiling at Nvidia's GTC 2026. The platform is supported by Nvidia technology including CUDA libraries, NIM microservices, Nemotron Retriever models, and the open source cuVS vector algorithm library.", "body_md": "# Dell Adds Data Context, Prep, And Storage Features To Its AI Data Platform\n\nEarlier this year, Dell published the results from a survey of 3,800 enterprise IT decision-makers and AI experts from around the world talking about their experiences in adopting and scaling the technology, how they use it and the accelerates they put in place. What the vendor was told was that data – including its quality, availability, management, and security – [is the top challenge](https://www.dell.com/en-gb/blog/five-insights-for-smarter-enterprise-ai-adoption/).\n\nThat didn’t surprise Dell executives, who since launching its AI Data Platform two years ago have been sounding that message that what’s tripping up organizations [bringing AI into their operations](https://www.nextplatform.com/compute/2026/09/07/dell-says-ai-will-drive-75-percent-of-datacenter-demand-by-2030/5294882) is not so much GPUs or similar hardware situations but the data.\n\n“This is the gap that we hear from customers every day,” [Varun Chhabra](https://www.linkedin.com/in/varuncal/), senior vice president of Dell’s Infrastructure Solutions Group, told journalists at a recent media briefing. “The infrastructure is ready, but the data isn't, and without data that's ready for AI, investments turn from isolated pilots and turn into isolated pilots that never really reach the full potential through scale to the rest of the organization. The real bottleneck is not access to compute, it's not often access to models. It's actually access to data. Enterprise data is just not in a place where it's ready to support scaling of AI workloads.”\n\nAI needs a lot of data to scale, but much of that data is housed in different locations, from the cloud and datacenters to file systems, databases, applications, and edge locations, Chhabra said, adding that “much of it is trapped in silos, silos based on workloads, and there is often not a unified way to reach it. Much of is unstructured. It's often dark, which means it's effectively invisible to AI. It is often ungoverned, so the teams are stuck between AI that needs access and policies that require control.”\n\nThrough the AI Data Platform – a foundational element of the [Dell’s AI Factory](https://www.nextplatform.com/ai/2024/05/21/dell-wants-to-help-you-build-your-ai-factory/1637009) – the company is working to put the tools and capabilities in place to address the myriad data challenges. As seen below, the platform is essentially three layers, starting with the Data Orchestration Engine for taking in, preparing, labeling, and enriching data, complete with unified data pipelines, distributed control that decouples compute from storage, and native access to Nvidia NIM [microservices](https://www.nextplatform.com/code/2025/04/23/nvidia-nemo-microservices-for-ai-agents-hits-the-market/1645836), [AI Blueprints](https://www.nextplatform.com/control/2024/08/27/nvidia-rolls-out-blueprints-for-the-next-wave-of-generative-ai/1648263), and templates.\n\nThe Data Engines for analytics, processing, and search helps organizations find the use the right data they need in hours instead of weeks, Chhabra said. [Then come the Storage Engines](https://www.nextplatform.com/store/2026/05/29/data-and-storage-at-the-center-of-the-ai-stack/5247585) that include PowerScale for network attached storage (NAS) for unstructured data and high-throughput AI workloads, ObjectScale that supports S3-over-RDMA and Nvidia CUDA libraries for large object repositories and periodic snapshots of AI models, and the Lightning File System, a fast software-defined parallel system announced at Nvidia’s GTC 2026 show in March and released a month later. It’s aimed at high-scale training and inferencing.\n\nThe platform is supported by Nvidia technology. Along with the CUDA libraries and NIM microservices, the GPU-driven platform includes Nvidia’s Nemotron Retriever models for document parsing, embedding, and reranking, and cuVS, an open source library of GPU-accelerated algorithms for vector indexing and search.\n\nThe goal is to ensure that these unified storage layers provide exabyte-scale capabilities to ensure that GPUs aren’t left waiting for data to arrive and to address what Chhabra calls the “pilot reproduction gap.”\n\n“The use case works in the demo, or there's a small pilot rollout, and users are excited, leadership is excited,” he said. “But then when you start scaling it out, all of a sudden challenges around data governance – not having the right data, having dirty data – impacting the AI challenges or AI outcomes come in. That's really where the real enterprise scale is held back.”\n\nDell this week is adding more capabilities to the platform, with a focus on agentic AI. The new features are designed to give agents common context and meanings to terms that are scattered across structured and unstructured data, an easier path to finding the right context for such data, and then give them a way to use that context.\n\nThe challenge is that while agents on the platform access the same data, documents, and history, they have to reconstruct their understanding and regenerate tokens every time a query is run and more data is analyzed, Chhabra said.\n\n“This costs extra tokens,” he said. “It often creates some churn because there's a back and forth that inevitably ends up happening with the human in the loop, which then slows things down. Often that means the answers that you're getting out of this are not fully trustworthy, so context is really what's important. What all of [the new features drive] is fewer steps and lower costs because agents don't have to keep rebuilding context in every request. They actually have data that can flow into the agent that reduces compute costs [and] the number of tokens that are generated.”\n\nThe Unified Semantic Layer provides the rules and definitions to words so they mean the same thing wherever they appear. In a use case Chhabra outlined, he noted that the word “defect” can have one definition in one manufacturing plant and another definition in a second plant. The Unified Semantic Layer, which includes a searchable glossary, lets agents understand the word means the same thing at each plant. Dell also is using Nvidia’s Auto-Ontology open source library that can create knowledge graphs from enterprise data.\n\nThe Enterprise Knowledge Graph shows how structured and unstructured data is related using metadata, lineage, and query history to continuously tune the graph. While the semantic layer can show agents what a word actually means, the knowledge graph can provide connections between that data and data in other systems.\n\n**“** If you're trying to draw some correlations between how often a defect happens and whether it's connected to a particular supplier or a particular distributor or anything in the supply chain, rather than having the model actually having to do rerun and create tokens every time, the Enterprise Knowledge Graph can actually make those connections for the model, making it much more token-efficient,” he said.\n\nUsers can then take what’s been done with the semantic layer and knowledge graph to create Knowledge Agents that each can work as a trusted agent about a particular topic. Rules can be built around them dictating the guidance it follow, how much it can spend in tokens, and what data it can see. It also uses Nvidia’s Nemotron Retriever models for reasoning and visually understanding the data.\n\nChhabra said Knowledge Agents also give enterprises greater flexibility in their AI environments, such as choosing the size of the AI model they’re using for a specific use case, whether to use general-purpose models in the cloud, or open source models on-premises.\n\nDell also is bringing Nvidia’s cuDF GPU-accelerated library to the Data Processing Engine to accelerate data processing and analytics and Apache Arrow to move data efficiently between data storage and processing so data can be queried where it sits. Organizations can now use GPUs rather than CPUs in data processing, which can mean 20 times faster batch processing of data than what can be done with a CPU, and four times the data processing speed.\n\nThe cuDF-and-Arrow combination complements the cuVS-augmented search and cuDF analytics already in the Dell Processing Engine.\n\nWith the new open source Dell Storage Performance Tool for ObjectScale and PowerScale, organizations can more accurately measure the needed size of their AI infrastructures and keep data flowing to their GPUs.\n\n“Organizations can define the workload they want to benchmark our storage platforms with, and then checkpoints style writes, high-concurrency reads, mixed read-write environments, even high-score queries,” Chhabra said. “They can all run these benchmarks on their own infrastructure with the same tool that Dell Engineering is using internally for benchmarking our products.”", "url": "https://wpnews.pro/news/dell-adds-data-context-prep-and-storage-features-to-its-ai-data-platform", "canonical_source": "https://www.nextplatform.com/store/2026/10/06/dell-adds-data-context-prep-and-storage-features-to-its-ai-data-platform/5301356", "published_at": "2026-10-06 12:59:29+00:00", "updated_at": "2026-10-06 13:17:09.170344+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-products", "ai-tools"], "entities": ["Dell", "Varun Chhabra", "AI Data Platform", "AI Factory", "Nvidia", "PowerScale", "ObjectScale", "Lightning File System"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/dell-adds-data-context-prep-and-storage-features-to-its-ai-data-platform", "markdown": "https://wpnews.pro/news/dell-adds-data-context-prep-and-storage-features-to-its-ai-data-platform.md", "text": "https://wpnews.pro/news/dell-adds-data-context-prep-and-storage-features-to-its-ai-data-platform.txt", "jsonld": "https://wpnews.pro/news/dell-adds-data-context-prep-and-storage-features-to-its-ai-data-platform.jsonld"}}