Skip to main content
D squared Enterprise Advisory

Data & AI

Trusted data as the foundation of AI adoption

AI quality is bounded by data quality. Why ownership, lineage and fitness-for-purpose decide what your AI investments can actually return.

Denys Dyeyev · · 6 min read

Every AI capability an organization deploys is, underneath, a claim about its data: that the data means what we think it means, that it is current enough, complete enough and owned by someone who will notice when it breaks. Most organizations discover the strength of that claim only after an AI initiative starts producing confident answers built on data nobody actually vouches for.

This is why AI adoption is, in large part, a data trust program wearing a more fashionable name. The ceiling on what your AI investments can return is set less by model choice than by the state of the data those models read. Models improve every quarter without your involvement; your data does not.

Trust is fitness for purpose, not perfection

The first trap to avoid is treating data quality as an absolute. Data is not good or bad in the abstract; it is fit or unfit for a specific use. Customer records that are perfectly adequate for monthly invoicing may be dangerously stale for a real-time service conversation. A tolerance that is acceptable in aggregate reporting may be unacceptable when an AI system acts on individual records.

Data is not good or bad in the abstract; it is fit or unfit for a specific use.

This reframing matters because it makes the problem tractable. You do not need perfect enterprise data; no one has that. You need to know, for each AI use case that matters, which data it depends on and whether that data meets the bar that use case requires. Trust becomes a measurable, scoped property instead of an aspiration.

Ownership is the first fix

When data problems surface, the instinct is to reach for tooling: catalogues, quality dashboards, observability platforms. Tools help, but they automate a decision the organization usually has not made: who owns this data? Ownership means a named person accountable for a dataset's definition, quality and appropriate use, someone who can answer questions and authorize fixes.

Most data quality problems persist not because they are hard to fix but because they are nobody's job to fix. Establishing ownership for the data domains your priority AI use cases depend on is unglamorous work, and it moves the needle more than any platform purchase.

Most data quality problems persist not because they are hard to fix but because they are nobody's job to fix.

Lineage: knowing what fed the answer

AI raises the stakes on a capability many organizations have deferred: lineage, knowing where data came from and what transformed it along the way. When an AI system gives a wrong or surprising answer, the first diagnostic question is always what did it read? If answering that takes a forensic investigation, every incident becomes expensive and every assurance to a regulator becomes hand-waving.

Retrieval-based AI architectures make this concrete: the quality of answers tracks the quality and freshness of the sources being retrieved. Curating those sources (deciding what belongs in the knowledge base, who maintains it and how staleness is detected) is a data governance activity, whatever the project plan calls it.

Scope foundations to use cases

The final trap is the enterprise-wide data program that must finish before AI can start. These programs fail by their own weight: three years of foundation-laying with value perpetually one phase away. The alternative is to let prioritized AI use cases pull the data work: for each high-value use case, identify the data it depends on, assess fitness, fix ownership and quality for that slice, and ship. Each use case leaves the data landscape genuinely better, and the improvement compounds.

The trap

value, one phase away

One enterprise-wide data programme that must finish before AI can start.

The alternative

Use case 1

Fix the data it needs, then ship value

Use case 2

Fix the data it needs, then ship value

Use case 3

Fix the data it needs, then ship value

Each use case leaves the data landscape better, and the improvement compounds.

Two ways to build data foundations. The programme that must finish first rarely does; use-case-pulled work ships value and compounds.

Data trust built this way is not a prerequisite for AI adoption; it is a product of doing AI adoption properly. The organizations that understand this stop asking whether their data is ready for AI and start asking, use case by use case, what it would take to make it ready. That question has an answer, a cost and a deadline, which is exactly what a foundation should have.

Key takeaways

  • AI quality is bounded by data quality; models improve on their own, your data does not.
  • Trust is fitness for purpose, not perfection: the same data can be fit for one use and unfit for another.
  • Ownership is the first fix; most quality problems persist because they are nobody's job to fix.
  • Let prioritized use cases pull the data work, so each one ships value and the improvement compounds.

Discuss how this applies to your organization.