Data Management

A Framework for Taming Unstructured Data at Scale

Written by: Lana DeMaria | Director of Privacy and IT Compliance, Zillow

Updated 11:00 AM EDT, September 16, 2026

post detail image
Lana DeMaria | Director of Privacy and IT Compliance, Zillow Lana DeMaria is Director of Privacy and IT Compliance at Zillow, with 20+ years of experience in data governance, privacy, compliance, and AI risk.
(This article originally appeared in CDO Magazine’s AI and Data Governance in the Enterprise Trend Report.)

For decades, the discipline of data governance has been a story told in rows and columns. Organizations have become experts at governing structured data: locking down databases, sanitizing spreadsheets, and building rigid controls around structured systems — fortresses around the “knowns.”

Yet most enterprises still ignore their “dark matter”: the 80% to 90% of corporate information that exists as emails, PDFs, Teams messages, video transcripts, design files, and loose contracts. This unstructured data is where the context of the business lives. But it is also where some of the greatest risk hides.

In the era of generative AI and increasing regulatory scrutiny, the “store it and ignore it” strategy is no longer sustainable. Organizations need a practical framework for improving visibility, reducing risk, and governing unstructured data continuously at scale.

What is unstructured data governance?

Traditional governance tools were built for structured data (SQL databases), which is neat, tidy, and easy to query. Unstructured data is messy, human, and growing exponentially.

The danger lies in the “Visibility Gap.” While the structured 15% and semi-structured 5% of enterprise data are heavily guarded with firewalls, encryption, and access controls, the unstructured 80% often sits in open S3 buckets, over-permissioned SharePoint sites, and forgotten network drives.

Figure 1: The Enterprise Data Iceberg and the 80% Visibility Gap

(Visualizing the imbalance between structured and unstructured enterprise data)

Story Image
Story Image
Image source: https://www.komprise.com/glossary_terms/unstructured-data-management/ 

The rise of generative AI and modern “double extortion” ransomware tactics is making this visibility gap increasingly difficult for organizations to ignore. Two trends, in particular, are driving this shift.

The twin accelerants: Why the risk is accelerating

Two forces are colliding to make this visibility gap untenable: AI adoption and modern cyber threats.

1. The rise of “shadow AI”

Employees are eager to use Large Language Models (LLMs) to summarize reports or generate code. Without governance, they may inadvertently feed proprietary unstructured data — strategy documents, HR records, and intellectual property — into public models.

Conversely, internal Retrieval-Augmented Generation (RAG) systems that lack permission awareness can accidentally surface sensitive executive memos to a junior intern asking a simple question.

2. The ransomware pivot

Modern ransomware groups have evolved from simple encryption to “double extortion.” They don’t just lock data; they threaten to release it publicly. They specifically target unstructured repositories because that is where the embarrassing, regulated, and valuable secrets are kept.

Figure 2: The twin accelerants of risk

(How AI adoption and ransomware are increasing governance exposure)

Story Image

Why traditional governance approaches fail 

Most organizations attempt to solve this problem with “projects.” They hire consultants to clean up a file share, or they buy a storage optimization tool to archive old files. These projects fail because they treat data governance as an event, not a lifecycle.

These failures usually stem from three outdated assumptions:

1. The folder fallacy

Assuming that a sensitive data file (ex: “HR”) is inherently secure. In reality, users save sensitive data to their desktops or “Misc.” folders constantly. 

2. Manual classification

Expecting humans to manually tag files with predefined classification labels (i.e., public, internal, confidential). At the petabyte scale, asking employees to tag their documents is a fantasy. It introduces friction and inconsistency. 

3. Regex reliance

Relying on Regular Expressions (pattern matching) to find sensitive data. A 16-digit number might be a credit card, or it might be a part number. Regex creates so much noise (and false positives) that security teams eventually turn the alerts off. 

Figure 3: Why traditional governance fails

(Three outdated assumptions that prevent governance programs from scaling.)

Story Image

A practical framework for governing unstructured data

To break the paralysis, organizations need a cyclical, automated framework. This is not a “clean-up” job; it is a new operational standard. 

Figure 4: The continuous governance cycle

(A continuous framework for discovery, classification, and automated compliance)

Story Image

1. Discovery (The visibility layer)

Modern discovery is not about moving data; it is about indexing it in place. The goal is to build a “Metadata Map” of the environment.

We can do this by focusing on three key areas:

  • Scan Everything: Connect to on-premise NAS, cloud object storage (AWS, Azure, GCP), and SaaS platforms such as OneDrive, Google Drive, and Slack.
  • Calculate ROT (redundant, obsolete, and trivial) Data Immediately: Typically, around 40% of unstructured data has not been accessed in over three years
  • Identify Permissions: Map out who has access to what. One of the most common vulnerabilities is the “Everyone” group applied to folders containing sensitive data. 

2. Classification (The context layer)

“Context is King.”

This is where AI changes the game. Instead of relying on keywords, organizations can use semantic analysis and machine learning to classify information based on meaning and context. (Source: Why Unstructured Data Governance is the Key to Scaling AI.)

Some examples include:

  • Named Entity Recognition (NER): AI models can distinguish between “Apple” the fruit and “Apple” the company. They can identify Personally Identifiable Information (PII) with much higher precision by reading the surrounding context.
  • Document Clustering: Unsupervised learning can group similar documents together. If the system finds one legal contract, it can often identify thousands of others that look similar, even if they are named “Scan_001.pdf.”
  • Sensitivity Scoring: Every file receives a risk score based on its content, age, and permission settings. For example, a file containing PII that is open to the internet would be considered a critical risk.

3. Continuous compliance (The action layer)

“Policy as Code.” 

Governance must be automated. We cannot rely on a user remembering to delete a draft after a project ends. 

  • Automated Lifecycle Management: Policies should trigger actions automatically. For example: 
    • Rule: “If a Resume is older than 3 years, delete it.” 
    • Rule: “If a file contains PCI data, move it to the Secure Vault.” 
  • The “Defensible Deletion” Imperative: Deleting data is the ultimate security control. If you don’t have it, it can’t be stolen. By enabling automatic deletion based on outlined policies, legal teams can defend the action in court as a standard business practice, rather than spoliation of evidence. 

The value vs. risk curve

One of the most compelling arguments for this framework is the economic one. Unstructured data follows a predictable curve: its business value drops rapidly over time, while its risk (liability) grows or remains static. 

Figure 5: Data value vs. risk curve

(The “danger zone” begins when risk outweighs business value.)

Story Image

The “danger zone” begins when the value of the data drops below the risk of retaining it. This becomes the optimal point for archival or deletion.

AI as a force multiplier for governance

We often talk about AI governance, but we must also talk about AI for governance. The scale of unstructured data (often billions of files) exceeds human capacity. AI is the only tool capable of bridging this gap. 

Vector embeddings for search

Traditional search looks for matching words. AI-driven vector search converts text into numerical coordinates (vectors), allowing the system to understand concepts rather than exact phrasing.

For example, a search for “Termination Agreements” will also identify documents titled “Severance Deal” or “Separation Contract,” ensuring legal holds capture all relevant documents, not just the ones with the correct filename.

The “auto-steward”

Instead of asking data owners to review thousands of files manually, organizations can use generative AI to summarize repositories and recommend actions.

  • Old Way: Send a list of 10,000 file paths to a department head. They ignore it.
  • New Way: AI analyzes the files and sends a digest:
    “I found 5,000 files in the Marketing folder. 98% appear to be assets from the 2019 Winter Campaign. They have not been accessed in four years. Recommendation: Archive to Cold Storage. Approve?”

This reduces the cognitive burden on business users and enables governance to actually happen.

An operational model for leaders

So, who owns all of this? Do we need to hire a massive new army of compliance officers?

No. We need a matrixed responsibility model.

Figure 6: An operational model for leaders

(Governance requires clear accountability across strategy, technology, and business teams.)

Story Image

1. The governance council (strategy)

Role: Sets the organization’s risk appetite.

Action: Defines governance rules and policies.

Examples include:

  1. “We will not retain customer chat logs older than 90 days.”
  2. “We prohibit the storage of unencrypted PII on laptops.”

2. The data custodians (IT and security)

Role: The mechanics.

Action: Manage the infrastructure, governance platforms, and secure environments. They ensure scans are running, tools are functioning properly, and quarantine zones remain secure. They do not own the data.

3. The data owners (Business units)

Role: The decision-makers.

Action: Determine whether data is still operationally valuable and make retention decisions.

The shift is toward “assisted stewardship,” where AI helps surface insights and recommendations while humans remain responsible for the final decision.

What leaders should do next

Unstructured data is the memory of the organization. Managed poorly, it becomes a fog that obscures vision, hides predators, and leaks secrets. Managed well, it becomes a strategic enterprise asset: a rich source of institutional knowledge capable of supporting the next generation of enterprise AI.

The path forward is not about hiring more archivists or buying more storage. It is about building a continuous, AI-driven ecosystem that automatically separates signal from noise.

By implementing a discipline of Discovery, Classification, and Continuous Compliance, leaders can finally turn the lights on in the dark corners of their enterprise.

Related Stories

September 17, 2026  |  In Person

Chicago Leadership Summit

Renaissance Chicago Downtown Hotel

Similar Topics
Artificial Intelligence
Data Management
Diversity
Testimonials
background imagebackground image
Community Network

Join Our Community

starElevate Your Personal Brand

starShape the Data Leadership Agenda

starBuild a Lasting Network

starExchange Knowledge & Experience

starStay Updated & Future-Ready

logo
Social media icon
Social media icon
Social media icon
Social media icon
About