Data Management
Written by: Gopichand Mannava | Chief Data Architect, State of Connecticut & Independent Researcher
Updated 10:00 AM EDT, September 29, 2026

For years, Chief Data Officers (CDOs) have been measured by how well they improve data quality: whether information is accurate, complete, consistent, and available when the organization needs it.
Data quality remains essential. But artificial intelligence has introduced deeper questions about whether we can prove:
Those questions get to the difference between data quality and data provenance.
Data quality focuses on whether a value is reliable for a purpose, while data provenance traces its history. That history may include its origin, how it was collected or generated, ownership or usage rights, and how it has been transformed or enriched as it moves through the enterprise.
A dataset can be accurate and internally consistent yet still be unsuitable for an AI use case if an organization cannot verify its source, permissions, or material changes before it reaches a training set, retrieval system, or decision workflow.
Traditional data lineage remains important, but today’s AI programs bring more contributors into a data asset’s history. Operational data may be combined with third-party enrichment or synthetic records, while machine-generated classifications, embeddings, and documents processed by or generated from large language models can add further layers. The more sources and transformations involved, the harder that history becomes to track and verify.
For CDOs, provenance is becoming a practical control for trust – not merely a documentation exercise.
Established lineage practices typically trace data from a known source system through integrations and transformations into a warehouse, report, or application. That work remains indispensable.
AI introduces new complications. A single customer attribute may begin in an internal system of record, be enriched by a vendor, inferred by an analytical model, and later incorporated into a generative AI workflow. If those stages are not distinguishable, teams cannot tell which parts of an output are observed facts and which are modeled or generated assumptions.
Transformations have also expanded beyond conventional extract, transform, and load (ETL) processes. Data may be redacted, tokenized, labeled, summarized, embedded, augmented, converted into synthetic examples, or loaded into a retrieval index. Each step can change the data’s meaning, sensitivity, risk profile, and permitted use.
Third-party dependency increases the challenge. Vendors may provide data products, AI services, enriched records, or pre-trained models without sufficient visibility into source data, licensing, synthetic-content use, retention, or model updates. A procurement process focused only on security and service levels can leave a serious provenance gap.
Generated content can also re-enter enterprise systems. A generated summary, recommendation, classification, or synthetic record may later become input to another model or business process. Without clear labels and lifecycle controls, generated material can begin to look like an observed fact.
A provenance-first approach does not require rebuilding every data platform. It requires treating provenance as a control that travels with data – and becomes more rigorous as data approaches a high-impact AI use case.
At ingestion, assign each dataset, data product, or content collection a durable identifier. Capture the source system or organization, acquisition date, owner, sensitivity level, permitted uses, applicable rights or contractual restrictions, and whether the content is observed, inferred, vendor-enriched, or synthetic.
Store this information in a data catalog, governance platform, or comparable enterprise repository – not in isolated project files or individual team notes.
For third-party data, require provider documentation describing source categories, rights to use and share the data, known limitations, and the presence of modeled, generated, or synthetic elements. Perfect record-level visibility may not be possible, but organizations should know what they can and cannot verify before approving use.
Every material transformation should create an auditable connection between an output and its inputs. This includes conventional ETL processes, as well as feature engineering, model-assisted labeling, embedding creation, retrieval indexing, redaction, summarization, augmentation, and synthetic-data generation.
Teams should be able to identify the source dataset and version, transformation or model used, responsible team or service account, business purpose, material changes to meaning or sensitivity, and resulting data or model artifact.
For high-impact AI use cases, the evidence should extend across training, testing, validation, fine-tuning, retrieval, deployment, and monitoring. NIST’s AI Risk Management Framework emphasizes lifecycle risk management, accountability, transparency, and context-appropriate documentation.
Provenance will not always be complete. Legacy systems may lack documentation, acquired data may have unclear lineage, and vendor datasets may provide broad descriptions rather than record-level traceability.
The correct response is neither to ignore incomplete provenance nor automatically reject all imperfect data. Instead, classify incomplete provenance as a known risk and apply controls proportionate to the intended use.
For low-risk internal experimentation, that may mean labeling, restricted access, and human review. For AI that influences eligibility, payment, enforcement, fraud investigation, employment, health, or other consequential outcomes, incomplete provenance should trigger a much higher bar: remediation, independent validation, restrictions on use, or a decision not to use the data.
“Unknown” should be a governed status — not an invisible gap.
A dataset may be appropriate for one purpose and unsuitable for another. De-identified synthetic data may be useful for software testing, demonstration environments, or early model experimentation. That does not automatically make it appropriate for training a production model that influences benefits, investigations, or service delivery.
Before approving data for an AI use case, CDOs should ask:
These answers should be documented in an AI use-case assessment, model card, system card, or comparable governance artifact.
Synthetic data is not inherently less trustworthy than observed data. It can support privacy protection, test-data generation, rare-event simulation, and model development. Its fitness for use depends on knowing how it was created, what it represents, and what limitations it carries.
A practical synthetic-data policy should require clear labeling, traceability to the generation method and source version, approved-use restrictions, and separation from observed data across exports, retraining, reporting, and downstream reuse.
Otherwise, generated patterns may become indistinguishable from real-world evidence. NIST’s Generative AI Profile recommends documenting provenance, including source, signatures, versioning, and watermarks where applicable, and highlights risks involving training data and synthetic content.
Consider a hypothetical public-sector fraud-detection program with limited historical cases for model development. To expand testing while reducing exposure of sensitive information, the team creates synthetic records based on historical patterns.
Initially, the records are held separately in a development environment. Months later, a retraining pipeline combines multiple sources into a shared feature repository. The synthetic records enter the repository without a persistent label.
During a later training cycle, the model learns from the mixed dataset. When investigators challenge a recommendation, the program cannot determine whether the relevant pattern originated in an observed case, vendor-enriched record, or generated example.
This hypothetical scenario illustrates a governance failure: data changed context as it moved through the enterprise, while the information needed to interpret it did not travel with it.
Durable synthetic-data labels, versioning, approved-use restrictions, training-set review, and transformation records would not prohibit synthetic data. They would make its use visible, deliberate, and governable.
The CDO’s role is expanding beyond data quality management. It now includes ensuring that organizations can understand the authenticity, history, permissions, and limitations of the data powering AI systems.
Data lineage remains essential. AI has raised the stakes through synthetic content, vendor enrichment, complex transformations, external models, and generated outputs that can circulate back into enterprise decisions.
For CDOs, “Is this data accurate?” is only part of the question.
They also need to ask: “Where did it come from, what happened to it, what are we allowed to do with it, and can we prove those answers when they matter?”