Data Management
Written by: Lana DeMaria | Director of Privacy and IT Compliance, Zillow
Updated 11:00 AM EDT, September 16, 2026

For decades, the discipline of data governance has been a story told in rows and columns. Organizations have become experts at governing structured data: locking down databases, sanitizing spreadsheets, and building rigid controls around structured systems — fortresses around the “knowns.”
Yet most enterprises still ignore their “dark matter”: the 80% to 90% of corporate information that exists as emails, PDFs, Teams messages, video transcripts, design files, and loose contracts. This unstructured data is where the context of the business lives. But it is also where some of the greatest risk hides.
In the era of generative AI and increasing regulatory scrutiny, the “store it and ignore it” strategy is no longer sustainable. Organizations need a practical framework for improving visibility, reducing risk, and governing unstructured data continuously at scale.
Traditional governance tools were built for structured data (SQL databases), which is neat, tidy, and easy to query. Unstructured data is messy, human, and growing exponentially.
The danger lies in the “Visibility Gap.” While the structured 15% and semi-structured 5% of enterprise data are heavily guarded with firewalls, encryption, and access controls, the unstructured 80% often sits in open S3 buckets, over-permissioned SharePoint sites, and forgotten network drives.
(Visualizing the imbalance between structured and unstructured enterprise data)


The rise of generative AI and modern “double extortion” ransomware tactics is making this visibility gap increasingly difficult for organizations to ignore. Two trends, in particular, are driving this shift.
Two forces are colliding to make this visibility gap untenable: AI adoption and modern cyber threats.
Employees are eager to use Large Language Models (LLMs) to summarize reports or generate code. Without governance, they may inadvertently feed proprietary unstructured data — strategy documents, HR records, and intellectual property — into public models.
Conversely, internal Retrieval-Augmented Generation (RAG) systems that lack permission awareness can accidentally surface sensitive executive memos to a junior intern asking a simple question.
Modern ransomware groups have evolved from simple encryption to “double extortion.” They don’t just lock data; they threaten to release it publicly. They specifically target unstructured repositories because that is where the embarrassing, regulated, and valuable secrets are kept.
(How AI adoption and ransomware are increasing governance exposure)

Most organizations attempt to solve this problem with “projects.” They hire consultants to clean up a file share, or they buy a storage optimization tool to archive old files. These projects fail because they treat data governance as an event, not a lifecycle.
These failures usually stem from three outdated assumptions:
Assuming that a sensitive data file (ex: “HR”) is inherently secure. In reality, users save sensitive data to their desktops or “Misc.” folders constantly.
Expecting humans to manually tag files with predefined classification labels (i.e., public, internal, confidential). At the petabyte scale, asking employees to tag their documents is a fantasy. It introduces friction and inconsistency.
Relying on Regular Expressions (pattern matching) to find sensitive data. A 16-digit number might be a credit card, or it might be a part number. Regex creates so much noise (and false positives) that security teams eventually turn the alerts off.
(Three outdated assumptions that prevent governance programs from scaling.)

To break the paralysis, organizations need a cyclical, automated framework. This is not a “clean-up” job; it is a new operational standard.
(A continuous framework for discovery, classification, and automated compliance)

Modern discovery is not about moving data; it is about indexing it in place. The goal is to build a “Metadata Map” of the environment.
We can do this by focusing on three key areas:
“Context is King.”
This is where AI changes the game. Instead of relying on keywords, organizations can use semantic analysis and machine learning to classify information based on meaning and context. (Source: Why Unstructured Data Governance is the Key to Scaling AI.)
Some examples include:
“Policy as Code.”
Governance must be automated. We cannot rely on a user remembering to delete a draft after a project ends.
One of the most compelling arguments for this framework is the economic one. Unstructured data follows a predictable curve: its business value drops rapidly over time, while its risk (liability) grows or remains static.
(The “danger zone” begins when risk outweighs business value.)

The “danger zone” begins when the value of the data drops below the risk of retaining it. This becomes the optimal point for archival or deletion.
We often talk about AI governance, but we must also talk about AI for governance. The scale of unstructured data (often billions of files) exceeds human capacity. AI is the only tool capable of bridging this gap.
Traditional search looks for matching words. AI-driven vector search converts text into numerical coordinates (vectors), allowing the system to understand concepts rather than exact phrasing.
For example, a search for “Termination Agreements” will also identify documents titled “Severance Deal” or “Separation Contract,” ensuring legal holds capture all relevant documents, not just the ones with the correct filename.
Instead of asking data owners to review thousands of files manually, organizations can use generative AI to summarize repositories and recommend actions.
This reduces the cognitive burden on business users and enables governance to actually happen.
So, who owns all of this? Do we need to hire a massive new army of compliance officers?
No. We need a matrixed responsibility model.
(Governance requires clear accountability across strategy, technology, and business teams.)

Role: Sets the organization’s risk appetite.
Action: Defines governance rules and policies.
Examples include:
Role: The mechanics.
Action: Manage the infrastructure, governance platforms, and secure environments. They ensure scans are running, tools are functioning properly, and quarantine zones remain secure. They do not own the data.
Role: The decision-makers.
Action: Determine whether data is still operationally valuable and make retention decisions.
The shift is toward “assisted stewardship,” where AI helps surface insights and recommendations while humans remain responsible for the final decision.
Unstructured data is the memory of the organization. Managed poorly, it becomes a fog that obscures vision, hides predators, and leaks secrets. Managed well, it becomes a strategic enterprise asset: a rich source of institutional knowledge capable of supporting the next generation of enterprise AI.
The path forward is not about hiring more archivists or buying more storage. It is about building a continuous, AI-driven ecosystem that automatically separates signal from noise.
By implementing a discipline of Discovery, Classification, and Continuous Compliance, leaders can finally turn the lights on in the dark corners of their enterprise.