Enterprise AI’s Real Bill Arrives in the Data Layer

Enterprise AI has a data problem. A 2025 report from MIT’s NANDA initiative found that 95% of enterprise generative AI pilots produced little or no measurable impact on the profit-and-loss statement. The missing ingredient is not always a better model; it is the architecture required to give that model useful, current, and properly controlled information.
Most enterprise AI architectures assume data will move to wherever the new platform lives. That assumption turns a technical choice into a long-running bill, because transfer fees, latency, bandwidth, security controls, and operational effort arrive before a model produces a useful answer.
Gaurav Chawla describes data gravity as “the idea that large accumulations of data pull applications towards them and not the other way around.” Records do not relocate because an AI project has appeared on an enterprise roadmap; they remain tied to regulation, sovereignty, ownership, and applications built around them.
The cost of data movement and management rarely arrives as one dramatic invoice. It appears as dozens of small demands on teams, then tends to surface 12 to 24 months after the platform is procured and the team is staffed—well after the launch has earned its celebratory slide.
Context Is Not a Storage Strategy
Moving raw data into a new system does not solve the problem by itself. A model needs context, not just raw data, to operate effectively, yet many organizations are trying to solve a context problem with a storage strategy.
Chawla puts the distinction plainly: “A model does not need raw data dumped in front of it. It needs context: the ability to find the right record, cross-reference it against policy, respect access rules and work from information that is current rather than a snapshot from the week the project began.”
Metadata and vectors can help AI find and reason about records without inheriting every constraint that keeps the source data where it lives. That approach does not erase regulation, sovereignty, ownership, or application dependencies; it gives AI a way to work with information while those constraints remain in place.
Production Exposes the Architecture
A pilot can show that a model works under controlled conditions. It cannot prove that the same system can run across an organization’s full data estate, where data volumes, users, relationships, and context change.
“A successful pilot does not prove an organization can run the same system across its full data estate,” Chawla said. The gap between development and production is where data architecture stops being a diagram and becomes an operating cost.
Shekhar Iyer described the failure mode through an AI support agent: “It works fine in development. But once in production, it encounters new variables and may recommend the wrong resolution because it lacks a customer’s product configuration or prior case history.” The model has not necessarily become worse; the surrounding information is incomplete.
That same issue reaches beyond support tools. Richard Clough said, “Every large organization now has an AI strategy, but almost all of it is aimed at AI that lives on a screen, not the AI that will drive vehicles, operate machinery and move goods through warehouses.” Those systems depend on timely context and reliable relationships between records, not a one-time data dump.
The lesson is unglamorous but costly: enterprise AI needs a data strategy that respects where information lives, how it changes, and who can use it. Without that foundation, the platform may launch on schedule while the real data tax waits 12 to 24 months to make itself known.
Based on




