Recognized as a 2026 Gartner® Peer Insights™ Customers’ Choice for DSPM
Get the Report

What Is Dark Data and Can It Live on Public AI Servers?

July 8, 2026Reading time: 11 mins
Mark Stone
Senior Technical Writer
banner-bg-dawn

Dark data used to be a storage problem. It lived in forgotten files and stale archives across your systems, racking up cost and creating the occasional audit headache. But the dark data that keeps people up at night now is a completely different beast. Every time someone drops a contract into ChatGPT or shares a spreadsheet with Claude, a copy of your most sensitive information lands on infrastructure you don’t control.

Your security team can’t see the data or search for it. Nor can they delete it, because as far as your own systems are concerned, it was never there. 

That’s dark data in its newest form, and the copies piling up on public AI servers are the fastest-growing and least understood part of a problem that many security teams assumed they had already handled.

What does dark data mean?

Dark data is information your organization collects and stores during normal operations but never uses for analysis, decisions, or to create any business value. Gartner coined the term years ago, and the picture it painted was a familiar one: stale SharePoint libraries, archived inboxes, abandoned shared drives, log files nobody reads, scanned documents from a project that wrapped in 2019.

The volume is staggering. In Splunk’s State of Dark Data survey of more than 1,300 business and IT leaders, respondents reported that 55% of their organization’s data is dark. 

Dark data is a strange phenomenon: you pay to store it, and you inherit its risk, yet you have almost no idea what it contains.

For most of its history, dark data was a cost-and-clutter story: too much information, too little visibility, and too much money spent storing files nobody’s opening. 

But now public AI tools have taken their seat at the dark data poker table—and they’ve gone all in. And the pot is your data. 

When dark data runs into the AI juggernaut

The classic definition assumes the data is still yours. It lives in your systems, behind your firewall, governed by your policies, even if no one remembers it exists. You could, in theory, go find it.

The newer form breaks that assumption entirely. Every time an employee pastes a customer list into a consumer chatbot, uploads a financial model for “quick analysis,” or runs a confidential deck through a free summarizer, a copy of that corporate data lands on infrastructure you don’t own and can’t inspect. It becomes dark in a far more literal sense: gone from your view the moment you hit Enter.

This is corporate data sitting on public AI servers, and it carries every weakness of traditional dark data plus a few that are genuinely new. You didn’t catalog it; you can’t search for it; and you have no record that it ever left. And depending on the tool and the plan, you may have handed it to an AI tool whose retention practices and training policies you probably skipped over.

How your data ends up there

None of this requires a malicious insider or a sophisticated breach. It happens through ordinary people trying to do their jobs faster. And this is exactly what makes it so hard to stop.

A few of the most common paths:

  • The copy-paste. Someone drops sensitive text into a chatbot for a summary, a rewrite, or a translation. This is the most direct and most invisible route.
  • The file upload. Spreadsheets, contracts, and slide decks are fed into a tool that promises to “analyze your document.” The whole file goes, not just the question.
  • The browser extension. AI assistants read whatever tab you have open, including internal dashboards, CRM records, and draft emails.
  • The free tier. Consumer-grade plans often carry different data-handling terms than enterprise agreements, and employees rarely check which one they signed up for.
  • The helpful agent. AI agents wired into a workflow send context to an external model on every run, multiplying a single risky pattern across thousands of automated interactions.

This is essentially shadow AI in practice: unsanctioned tools that are adopted from the bottom up, faster than any security team can write a policy. And the data leaving your environment is often your most sensitive, because the most difficult problems are usually the ones people want help solving.

Why this version is the dangerous one

Traditional dark data is a sleeping liability. Data on public AI servers is an active risk.

What’s the difference between the two?

You can’t inventory it. A discovery scan of your own environment will never find a contract that’s already been uploaded to a public AI tool. The data is real, the exposure is real, and your tooling reports nothing.

You don’t control retention. With files you forgot in a SharePoint folder, you can still set a deletion policy. Once data is sent to a third-party AI service, retention follows that vendor’s terms, which can range from a few days to indefinite storage, and can change without your knowledge.

The question of language model training looms large. Some tools use submitted content to improve their models. Even when providers say they don’t, you’re still relying on a policy you can’t audit and a setting an employee may have left on default.

Compliance gets very uncomfortable. GDPR, HIPAA, PIPEDA, and similar compliance regulations expect you to know where personal data lives and to honor deletion requests. “It’s on a chatbot vendor’s server and we have no way to retrieve or delete it” is not the answer a regulator wants to hear. A single customer record pasted into the wrong tool can put you out of step with your own privacy commitments.

You learn about the exposure after the fact, if at all. Most organizations discover these risks during an incident, an audit, or when a customer asks a question they can’t answer. By then, the data may have been gone for months.

Dark data, shadow data, and shadow AI: understanding the differences 

These terms get used interchangeably in vendor content, which causes real confusion when a security team tries to scope a project. Each describes a different problem and requires different controls.

Term What it is Where it lives Primary risk
Dark data Data that’s collected and stored but never used or accounted for Stale sites, archived mail, old logs, forgotten drives Compliance exposure, AI retrievability, storage cost
Shadow data Data that’s created or copied outside sanctioned systems and IT visibility Personal cloud accounts, unmanaged apps, exported files Lost visibility, ungoverned data flows
Shadow AI Public AI tools adopted by employees without approval Consumer chatbots, free assistants, unvetted agents Sensitive data being sent to third-party servers

The easiest way to think about these terms is this: shadow AI describes the behavior; the corporate data that ends up on a public AI server becomes a form of exposed dark data; and shadow data is the wider category of any information that slipped outside your governed systems. 

Most enterprises are dealing with all three at once, and a copy-paste into a chatbot can trigger them all in a single afternoon.

What you can do about your dark data problem

You just can’t log onto a vendor’s servers and pull back data that’s been submitted; that ship has sailed. So, the work starts with protecting the data that hasn’t left your environment yet, and there’s plenty you can do on this front. 

Four moves cover most of that ground.

Get visibility into AI tool usage. You can’t manage a behavior you cannot observe. That means understanding which AI tools your people use and what types of data are being shared with them, rather than assuming the acceptable-use policy did its job.

Know where your sensitive data lives first. Data leaves through public AI tools because someone has access and a reason to use it. If you don’t know where your confidential contracts, customer records, and financial models are in your environment, you have little hope of catching them before they leave. Prevent the high-risk data flows. Once you can identify sensitive content based on what it actually is, rather than on a label someone forgot to apply, you can intercept the instances that matter most: PII, source code, regulated records, etc., before they reach an external model.

Offer a sanctioned alternative. People use consumer chatbots because they work while approved options may be slower or don’t exist. Give them a governed way to get the same job done, and the shadow AI pull weakens on its own. Simply banning the tools without replacing them only drives the behavior further underground.

How Concentric AI tackles dark data in AI

Concentric AI helps you understand and manage dark data before your most sensitive information can make its way into public AI applications.

Because our Semantic Intelligence platform uses patented deep learning to discover and classify data by meaning and context, we can determine with exceptionally high accuracy exactly what a data record is. For example, we can tell whether it’s a vendor contract, a purchase order, or another business document.

And once we’ve discovered your data, we’ll help you establish guardrails to prevent sensitive content from being shared with AI—we’ll help you apply classification labels, create and enforce access policies, and block or mask sensitive data from being shared where it shouldn’t.

Frequently asked questions

What is dark data, in simple terms?
Dark data is information your organization collected and stored but never uses. Classic examples include old emails, stale files, and system logs. The newer and riskier form is corporate data that employees have fed into public AI tools, which now lives on servers you cannot see, search, or delete.
Can data I put into ChatGPT or Claude really be dark data?
Yes, and it’s arguably more dangerous than the old version. You collected the data; it now sits in storage you don’t manage, and you have no practical way to find or retrieve it. It carries every weakness of traditional dark data plus the retention and training policy questions that come with a third-party AI vendor.
Does a paid or enterprise AI plan solve this?
It helps significantly because enterprise agreements often carry stronger data-handling terms than free consumer tiers. It doesn’t solve the underlying problem, though. The risk comes from employees using whichever tool is fastest, which is frequently the free consumer version they signed up for on their own. Visibility into actual usage matters more than the terms of the plan or tool you officially sanctioned.
How is dark data different from shadow AI?
Shadow AI is the behavior of people using unsanctioned AI tools. The corporate data that those tools end up holding is the dark data that behavior produces. One is the action, the other is the lasting exposure.
Can we get our data back from a public AI tool?
Usually not in any meaningful sense. Retention follows the vendor's terms, and you have no way to confirm what was kept, copied, or used. This is why prevention and visibility matter so much: the realistic play is to stop sensitive data before it leaves, since you can’t reliably claw it back after it’s gone.
How do we find corporate data that has already gone to public AI tools?
You cannot scan a vendor's servers, so the practical approach is to identify which AI tools your people use and what categories of data are flowing toward them, while building an accurate map of where your sensitive data lives internally. That combination tells you what is most at risk and lets you intercept the highest-priority flows going forward.

The latest from Concentric AI