The hidden tax on enterprise AI: Why data architecture is the ROI problem nobody budgeted for A 2025 report from MIT's NANDA initiative found that 95% of enterprise generative AI pilots produced little or no measurable impact on the profit-and-loss statement, with the report pointing to gaps in how companies integrate AI into operations rather than model quality. The analysis attributes the shortfall to "data gravity" — the tendency of large data accumulations to pull applications toward them — which drives unplanned costs for pipeline maintenance, reconciliation, governance and security that typically surface 12 to 24 months after a platform is procured. Those figures should change how companies build AI business cases, since a successful pilot does not prove an organization can run the same system across its full data estate. The hidden tax on enterprise AI: Why data architecture is the ROI problem nobody budgeted for Many large companies have spent the past few years investing in artificial intelligence infrastructure, software and implementation. Boards approved the plans. Finance built the business cases. Procurement negotiated for computing capacity. One question remained. How would they manage the data underneath those systems? This isn’t the kind of failure that makes headlines yet. There is no major outage, breach or recall. Instead, it is a quiet, compounding drag that appears as unplanned headcount and slipping timelines. Talk to enough infrastructure leaders about what happened and you’ll hear the phrase “data gravity” more than once. The scale of the problem is visible in the numbers. A 2025 report from MIT’s NANDA initiative https://nanda.media.mit.edu/ai report 2025.pdf found that 95% of enterprise generative AI pilots produced little or no measurable impact on the profit-and-loss statement. The report points to gaps in how companies integrate AI into their operations, not simply to model quality. Those numbers should change how companies build AI business cases: A successful pilot does not prove an organization can run the same system across its full data estate. The assumption baked into the budget Most enterprise AI architectures still assume that data will move to wherever the new platform lives. Pull the data into one place and the intelligence layer on top will just work. It’s a reasonable assumption, but it is almost never true at enterprise scale. Enterprise data has gravity. Regulations dictate where some records may live. Sovereignty rules keep other datasets inside a jurisdiction, even when compute is cheaper somewhere else. Business units that spent a decade building governance around a dataset have little incentive to hand it to a new central store and sometimes have no legal path to do so. Applications built around their data can also break in expensive ways when someone tries to separate the two. This is typically hidden in a proof of concept. It runs on a small, pre-cleaned slice of data. The costs appear later, when the program has to deal with the messy majority left out of the demo. “Data gravity” is the idea that large accumulations of data pull applications towards them and not the other way around. The bigger the dataset, the more expensive it is to move. Transfer fees are only the beginning; latency, bandwidth, security controls and operational effort add to the bill before a model produces a useful answer. Where the bill comes due The cost does not usually show up as one line item. It shows up as dozens of small demands on teams that already have too much to manage. They have to build and maintain pipelines to move data out of systems that were never designed for that kind of traffic, reconcile each copy with the original when the source system changes, and apply the same governance controls wherever the data is stored. They also have to keep track of the extra datasets created for development, testing, analytics and training. None of these tasks looks disastrous by itself. Together, they become a tax on the program—one that gets more expensive as the organization tries to scale it. Why the invoice arrives late The most dangerous part of this tax is its timing. It often appears 12 to 24 months after the platform is procured and the team is staffed, when initial success metrics have already been reported upward. By then, contracts are signed, teams are hired and the direction is publicly committed. Unwinding a centralization-first architecture is far more expensive than designing around data gravity from the start. Compute is easy to identify and assign to a budget; pipeline maintenance, reconciliation, governance, security reviews and the staff needed to keep the system running are spread across different organizations. The missing layer is context, not just storage Many postmortems on stalled AI programs stop at “our data wasn’t ready.” That diagnosis is accurate but incomplete. The deeper issue is that organizations are trying to solve a context problem with a storage strategy. A model does not need raw data dumped in front of it. It needs context: the ability to find the right record, cross-reference it against policy, respect access rules and work from information that is current rather than a snapshot from the week the project began. That does not come from a bigger warehouse. It comes from an enterprise-wide context layer: a consistent, governed way to discover, connect and retrieve information across systems and locations without requiring every source to hand its data to a central store. The underlying records, metadata and vectors AI systems use to reason about meaning are not the same thing. Records remain constrained by regulation, sovereignty, ownership and the applications built around them. Metadata and vectors can help AI find and reason about those records without inheriting every constraint that keeps the source data where it lives. The question is not how to get all data into one place. It is how to make the right context available wherever the data already lives, under whatever rules govern it. Enterprises that treat context, not consolidation, as what they are building avoid rebuilding the same tax under a different architecture two years later. The four key questions to ask Before signing off on a major AI infrastructure investment, an organization should be able to answer four basic questions: Which datasets are regulated or contractually restricted from moving? Who owns governance for each dataset today, and what happens when a copy exists elsewhere? How will those copies stay synchronized as source systems evolve? And can intelligence go to where the data already lives, or does the architecture require everything to be centralized? These questions do not choose the vendor for you, but they will shape the shortlist. The right architecture choice and vendor selection will lead you to AI outcomes and automate your business processes faster and at lower cost. These questions enable data architecture to be a first-class line item in the AI business case, not a detail to resolve after the compute deal closes. The real ROI problem The industry’s instinct is to measure AI ROI in tokens, throughput and GPU utilization. Those metrics are useful for gauging compute efficiency, but they tell only half the story. The other constraint on enterprise AI value sits upstream of compute: whether the data feeding the model can be trusted, governed and kept current without an ever-growing tax on the people around it. Token length and GPU efficiency tell you what it costs to run a model. They do not tell you whether the model has the context it needs to make a good decision. That depends on whether the underlying data is accurate, current, governed and connected to the right business process. It is the context curated from the data that determines the accuracy of agentic actions and whether you are on the path towards becoming an AI native organization or drifting away from it. Compute is a line item enterprises know how to budget for. Data gravity is the line item many are still discovering, usually long after the check clears. The companies that get the most from AI will be the ones that design their data architecture to serve the context agents needed without moving or centralizing the entire distributed state. Gaurav Chawla is Dell Technologies Inc.’s Fellow and VP, chief technology officer. He wrote this article for SiliconANGLE. Image: geralt/Pixabay https://pixabay.com/illustrations/binary-code-digitization-binary-6109177/ A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network , where technology leaders connect, share intelligence and create opportunities. - 15M+ viewers of theCUBE videos , powering conversations across AI, cloud, cybersecurity and more - 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/ https://siliconangle.com/aws-marketplace/ About SiliconANGLE Media SiliconANGLE https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552 , theCUBE Network https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da , theCUBE Research https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f , CUBE365 https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6 , theCUBE AI https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683 and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI. Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.