Data Management

The CDO’s Next Battleground: Proving Where Your Data Actually Came From

Written by: Gopichand Mannava | Chief Data Architect, State of Connecticut & Independent Researcher

Updated 10:00 AM EDT, September 29, 2026

post detail image
Gopichand Mannava | Chief Data Architect, State of Connecticut & Independent Researcher Gopichand Mannava is an AI governance and public-sector data architecture researcher with experience leading enterprise data initiatives for the State of Connecticut.

For years, Chief Data Officers (CDOs) have been measured by how well they improve data quality: whether information is accurate, complete, consistent, and available when the organization needs it.

Data quality remains essential. But artificial intelligence has introduced deeper questions about whether we can prove:

  • Where data came from
  • How it changed
  • Whether it is appropriate for the AI decision it is being asked to support

Those questions get to the difference between data quality and data provenance.

Data quality focuses on whether a value is reliable for a purpose, while data provenance traces its history. That history may include its origin, how it was collected or generated, ownership or usage rights, and how it has been transformed or enriched as it moves through the enterprise.

A dataset can be accurate and internally consistent yet still be unsuitable for an AI use case if an organization cannot verify its source, permissions, or material changes before it reaches a training set, retrieval system, or decision workflow.

Traditional data lineage remains important, but today’s AI programs bring more contributors into a data asset’s history. Operational data may be combined with third-party enrichment or synthetic records, while machine-generated classifications, embeddings, and documents processed by or generated from large language models can add further layers. The more sources and transformations involved, the harder that history becomes to track and verify.

For CDOs, provenance is becoming a practical control for trust – not merely a documentation exercise.

Why AI makes provenance harder

Established lineage practices typically trace data from a known source system through integrations and transformations into a warehouse, report, or application. That work remains indispensable.

AI introduces new complications. A single customer attribute may begin in an internal system of record, be enriched by a vendor, inferred by an analytical model, and later incorporated into a generative AI workflow. If those stages are not distinguishable, teams cannot tell which parts of an output are observed facts and which are modeled or generated assumptions.

Transformations have also expanded beyond conventional extract, transform, and load (ETL) processes. Data may be redacted, tokenized, labeled, summarized, embedded, augmented, converted into synthetic examples, or loaded into a retrieval index. Each step can change the data’s meaning, sensitivity, risk profile, and permitted use.

Third-party dependency increases the challenge. Vendors may provide data products, AI services, enriched records, or pre-trained models without sufficient visibility into source data, licensing, synthetic-content use, retention, or model updates. A procurement process focused only on security and service levels can leave a serious provenance gap.

Generated content can also re-enter enterprise systems. A generated summary, recommendation, classification, or synthetic record may later become input to another model or business process. Without clear labels and lifecycle controls, generated material can begin to look like an observed fact.

What CDOs should do

A provenance-first approach does not require rebuilding every data platform. It requires treating provenance as a control that travels with data – and becomes more rigorous as data approaches a high-impact AI use case.

Capture origin at ingestion

At ingestion, assign each dataset, data product, or content collection a durable identifier. Capture the source system or organization, acquisition date, owner, sensitivity level, permitted uses, applicable rights or contractual restrictions, and whether the content is observed, inferred, vendor-enriched, or synthetic.

Store this information in a data catalog, governance platform, or comparable enterprise repository – not in isolated project files or individual team notes.

For third-party data, require provider documentation describing source categories, rights to use and share the data, known limitations, and the presence of modeled, generated, or synthetic elements. Perfect record-level visibility may not be possible, but organizations should know what they can and cannot verify before approving use.

Preserve it through transformation

Every material transformation should create an auditable connection between an output and its inputs. This includes conventional ETL processes, as well as feature engineering, model-assisted labeling, embedding creation, retrieval indexing, redaction, summarization, augmentation, and synthetic-data generation.

Teams should be able to identify the source dataset and version, transformation or model used, responsible team or service account, business purpose, material changes to meaning or sensitivity, and resulting data or model artifact.

For high-impact AI use cases, the evidence should extend across training, testing, validation, fine-tuning, retrieval, deployment, and monitoring. NIST’s AI Risk Management Framework emphasizes lifecycle risk management, accountability, transparency, and context-appropriate documentation.

Govern incomplete provenance

Provenance will not always be complete. Legacy systems may lack documentation, acquired data may have unclear lineage, and vendor datasets may provide broad descriptions rather than record-level traceability.

The correct response is neither to ignore incomplete provenance nor automatically reject all imperfect data. Instead, classify incomplete provenance as a known risk and apply controls proportionate to the intended use.

For low-risk internal experimentation, that may mean labeling, restricted access, and human review. For AI that influences eligibility, payment, enforcement, fraud investigation, employment, health, or other consequential outcomes, incomplete provenance should trigger a much higher bar: remediation, independent validation, restrictions on use, or a decision not to use the data.

“Unknown” should be a governed status — not an invisible gap.

Match data to use

A dataset may be appropriate for one purpose and unsuitable for another. De-identified synthetic data may be useful for software testing, demonstration environments, or early model experimentation. That does not automatically make it appropriate for training a production model that influences benefits, investigations, or service delivery.

Before approving data for an AI use case, CDOs should ask:

  • Is the source known and permitted for this purpose?
  • Does the data represent observed events, inferred attributes, vendor enrichment, or synthetic content?
  • Are privacy, legal, contractual, retention, or intellectual-property restrictions understood?
  • Has the data changed in ways that could distort meaning or reduce representativeness?
  • Can the organization explain the data’s role if a model output is challenged?

These answers should be documented in an AI use-case assessment, model card, system card, or comparable governance artifact.

Synthetic data requires policy

Synthetic data is not inherently less trustworthy than observed data. It can support privacy protection, test-data generation, rare-event simulation, and model development. Its fitness for use depends on knowing how it was created, what it represents, and what limitations it carries.

A practical synthetic-data policy should require clear labeling, traceability to the generation method and source version, approved-use restrictions, and separation from observed data across exports, retraining, reporting, and downstream reuse.

Otherwise, generated patterns may become indistinguishable from real-world evidence. NIST’s Generative AI Profile recommends documenting provenance, including source, signatures, versioning, and watermarks where applicable, and highlights risks involving training data and synthetic content. 

A public-sector scenario

Consider a hypothetical public-sector fraud-detection program with limited historical cases for model development. To expand testing while reducing exposure of sensitive information, the team creates synthetic records based on historical patterns.

Initially, the records are held separately in a development environment. Months later, a retraining pipeline combines multiple sources into a shared feature repository. The synthetic records enter the repository without a persistent label.

During a later training cycle, the model learns from the mixed dataset. When investigators challenge a recommendation, the program cannot determine whether the relevant pattern originated in an observed case, vendor-enriched record, or generated example.

This hypothetical scenario illustrates a governance failure: data changed context as it moved through the enterprise, while the information needed to interpret it did not travel with it.

Durable synthetic-data labels, versioning, approved-use restrictions, training-set review, and transformation records would not prohibit synthetic data. They would make its use visible, deliberate, and governable.

Conclusion

The CDO’s role is expanding beyond data quality management. It now includes ensuring that organizations can understand the authenticity, history, permissions, and limitations of the data powering AI systems.

Data lineage remains essential. AI has raised the stakes through synthetic content, vendor enrichment, complex transformations, external models, and generated outputs that can circulate back into enterprise decisions.

For CDOs, “Is this data accurate?” is only part of the question.

They also need to ask: “Where did it come from, what happened to it, what are we allowed to do with it, and can we prove those answers when they matter?”

 

Related Stories

Similar Topics
Artificial Intelligence
Data Management
Diversity
Testimonials
background imagebackground image
Community Network

Join Our Community

starElevate Your Personal Brand

starShape the Data Leadership Agenda

starBuild a Lasting Network

starExchange Knowledge & Experience

starStay Updated & Future-Ready

logo
Social media icon
Social media icon
Social media icon
Social media icon
About