Shadow AI: Why Employees Feed Company Data Into ChatGPT 

2026-07-14
Marcel Piekarski
Shadow AI: Why Employees Feed Company Data Into ChatGPT 
8 min.

Shadow AI is a growing headache for companies trying to keep track of their own data. This article shows how it happens inside organizations and what to do about it, short of banning AI altogether.

TL;DR

Shadow AI is what happens when employees use public AI tools at work without company oversight. It’s now everywhere, and it exposes companies to data leaks, GDPR trouble, and expensive breach response. Bans don’t fix it, because employees just keep using AI under the radar. What works is a clear AI usage policy, backed by a secure tool employees can actually use, one that runs on the company’s own infrastructure and anonymizes sensitive data before anything leaves the network. 

Do you know where your company’s data goes when someone opens ChatGPT? As GenAI tools spread across the workplace, few stop to ask where that data ends up. At work, more and more people use ChatGPT and similar tools to work faster, usually without telling their employer. That’s Shadow AI: employees using AI at work without the company knowing. 

The reason is simple. An employee reaches for a public language model because it makes their job easier and faster. If the company hasn’t put a safe option in front of them, it’s no surprise they use what’s already at their fingertips. 

This article explains what Shadow AI is, why it puts company data at risk, and how to address it without banning AI outright. 

What Shadow AI Really Is 

IT teams already know a similar problem: Shadow IT. That’s when employees use external apps or services without approval, like sending files through a personal Dropbox account because the company hasn’t given them a better one. Shadow AI is the same idea, but with generative AI. 

In practice, Shadow AI means employees pasting company content into public tools like ChatGPT, Claude, Gemini, or Copilot, without any risk assessment and without the security team knowing. From a network perspective, it’s almost invisible. An AI query looks like any other web traffic, no different from opening a regular page. Nothing gets flagged; nothing gets caught, because the data leaves through the chat window. 

The scale of shadow AI in numbers
The scale of shadow AI in numbers

The scale is already well documented. A KPMG study surveying more than 48,000 people across 47 countries found that most employees using AI at work rely on public tools instead of vetted, company-approved alternatives. Companies, meanwhile, often lack the oversight to catch it. An IBM report shows that 63% of companies that suffered a breach had no AI governance policy and no way to detect Shadow AI. Employees are using AI every day, and their employers usually don’t know what data is going in. 

Why Shadow AI Slips Out of Control 

When companies first hear about Shadow AI, the instinct is often to blame employees. That reaction is understandable, but it misses what’s actually going on. The problem rarely comes down to bad intent. It’s time pressure, and the fact that a public model is right there while the company alternative either doesn’t exist or isn’t as good. 

The KPMG data tells the same story. A majority of employees using AI don’t tell their employer about it, and many pass off AI-generated content as their own work. Most of them don’t know whether laws or company rules on AI apply to what they’re doing, if such rules exist at all. To them, AI is just another tool, and it doesn’t occur to them to report it. 

When employees don’t disclose their AI use and the organization has no way to detect it, Shadow AI grows unnoticed. Companies typically find out only after an incident. 

What’s Actually Leaking, and Why It Matters 

Before we get to the numbers, it helps to know what actually leaves the company through the chat window. We’re talking about the assets a business runs on. 

The most common ones are customer data, snippets of source code, internal documents, financial figures, and materials that haven’t been published yet. Developers are a big share of GenAI users, which means code and proprietary solutions are especially exposed. Every paste is a potential data leak, and the company loses control the moment someone hits enter. 

There’s a legal side to this as well. When an employee pastes customer personal data into a public model, that data usually ends up on servers outside the European Economic Area. It happens without a legal basis, the company loses track of where the data lands, and there’s no realistic way to have it deleted. That’s a direct breach of GDPR principles around data minimization, purpose limitation, and confidentiality. If regulators come knocking, it’s the company, as data controller, that has to answer to the national data protection authority for losing control of the data. Picture a bank employee pasting a customer file, national ID number and all, into a public tool to draft a faster reply. One paste is enough to create a breach that must be reported and explained. 

On top of GDPR, there’s a newer layer to deal with. The EU’s AI Act has been rolling out since 2024, and the rules covering general-purpose models, including large language models, have applied since August 2025. The AI Act runs alongside GDPR, not in place of it, so companies in regulated sectors now have two overlapping frameworks to manage at once. 

The table below breaks down how different types of data translate into concrete risk and consequences. 

Type of Data Real Risk Consequence 
Customer personal data (including national ID numbers) Transfer outside the EEA with no legal basis GDPR breach, mandatory notification, administrative fines 
Source code and proprietary solutions Retention outside the company, possible use for model training Loss of competitive edge and intellectual property 
Internal documents and financial data Confidential information exposed ahead of time Loss of trade secrets, reputational damage 
Data in regulated sectors (finance, healthcare) Processing outside oversight and outside compliance Exposure to audits and industry sanctions 

What Shadow AI Actually Costs 

Legal and business risk translates into real, countable costs from breaches that have already happened. 

According to IBM’s 2025 Cost of a Data Breach Report, incidents involving Shadow AI cost an average of about $670,000 more than other breaches. One in five organizations surveyed had a breach tied to unsanctioned AI use, and in the vast majority of those cases, roughly 97%, there were no AI access controls in place at all. Two-thirds of those incidents ended with customer personal data getting exposed. Shadow AI leaks follow the same pattern over and over, and they carry a price you can put a number on. 

In regulated sectors, the stakes climb even higher, because the cost of the breach itself is only part of it. There’s also exposure to regulatory fines and audits. A financial institution, insurer, or healthcare provider that can’t say where its customer data ends up is looking at more than a financial hit. It’s also a regulatory investigation waiting to happen. 

The Samsung Case, and What the Biggest Data Leak Teaches Us 

The most famous Shadow AI incident hit a company nobody could accuse of being weak on tech. It’s a story that shows how fast things can go wrong, and it also points to some concrete conclusions. 

In 2023, Samsung allowed employees in one division to use ChatGPT. Within less than three weeks, there were three separate incidents in which employees pasted sensitive source code and a transcript of an internal meeting into the chat, trying to speed up their work. As a result, confidential data ended up on OpenAI’s servers, and Samsung lost control of it. The company’s first response was to ban public language models, and then it went further and started building its own AI tools running on its own infrastructure. 

Even a company with serious technology muscle didn’t manage to stay ahead of Shadow AI. Its first move was a ban, but it eventually landed on the conclusion that the only real fix was to run its own tool in-house. That decision came after the damage was already done, once the data had leaked. The same path is available before an incident, not just after one. 

How to Limit Shadow AI Without Banning It 

Bans mostly cover the problem up rather than solve it, so a different approach is needed. To actually limit Shadow AI, start from a simple observation: employees aren’t trying to break the rules; they’re trying to get their work done. If you give them a safe tool that works as smoothly as a public model, they have no reason to go looking elsewhere. The starting point, then, is deliberate management of AI inside the organization, or AI governance. 

Banning specific tools doesn’t take away the need to use AI. Employees will still find a way to meet that need, only now they’ll do it beyond the company’s reach. That leaves the organization losing on both fronts: it doesn’t get its control back, and it also misses out on the gains that a properly deployed AI tool would deliver. 

AI Usage Policy and Data Classification 

A clear AI usage policy backed by proper data classification is the foundation for anything AI-related in a company. Employees need to know which information they can work with freely and which should never leave the company. A general “don’t do it” without examples or explanations doesn’t cut it. Public data is one thing; customer personal data and trade secrets are another. An AI policy works best when employees have a tool that makes it easy to follow it in the first place. 

Secure AI Tools for Business 

Secure AI tools for business are designed so that data stays under the organization’s control. The company gives employees a single application that logs every query and ties it to a specific user. In a well-built tool, employees can work with multiple models in one place, so they don’t need to go back to public options. 

On-Premises Deployment and Sensitive Data Anonymization 

The way you deploy decides whether the data really stays inside the company. A model running on the organization’s own infrastructure (on-premises) or in an isolated private cloud environment means no query ever leaves the corporate network. For companies in regulated sectors, where data isn’t allowed to touch an external cloud in the first place, this is often the only acceptable path. 

When a company does use external models, the protection comes from sensitive data anonymization. The Extentum.AI platform automatically detects sensitive data in a query and replaces it with fictitious substitutes before the query goes out to the model, then restores the original values once the response comes back. The anonymization engine is trained to recognize local data formats, including things like national ID and tax identification numbers. The employee sees a natural, complete response, and sensitive data never leaves the company in a readable form because it’s replaced with fictitious values on the way out. 

Mature GenAI solutions also let you build multi-step processes with human oversight and auditability built into every step. This kind of setup, based on AI agents that run inside a designed workflow, gives the company visibility into every decision along the way and control over how tasks unfold. 

Shadow AI shows up wherever there’s no safe alternative. When the company puts one on the table, employees will simply use it. 

Data security with Extentum AI
Data security with Extentum AI

FAQ

Is Using ChatGPT at Work Legal? 

Using ChatGPT or similar tools isn’t banned. The problem shows up the moment personal or confidential data enters a public model. That’s when GDPR obligations kick in, and the responsibility sits with the company as data controller, not with the individual employee. Legality comes down to what data you put in and on what terms, not to whether you use the tool at all. 

How Do You Spot Shadow AI in Your Company? 

Shadow AI usually doesn’t stand out, because AI traffic looks just like any other encrypted traffic. Common signals are a gap between the official absence of AI tools and teams that suddenly work much faster, along with the use of personal accounts on company devices. You only get the full picture once you review which applications employees are actually using. 

What Happens to Data Pasted Into ChatGPT?

The content of the query goes to the model provider’s servers, usually outside the European Economic Area. Depending on the version of the tool and account settings, that data may be used for further model training. Once the query is sent, the company loses control over exactly where the information sits and who can access it. Removing data from a model that has been trained on it is generally impossible. 

Sources 

1. KPMG, Trust, Attitudes and Use of Artificial Intelligence: A Global Study 2025, 2025 – https://kpmg.com/xx/en/our-insights/ai-and-technology/trust-attitudes-and-use-of-ai.html 

2. IBM Newsroom, IBM Report: 13% of Organizations Reported Breaches of AI Models or Applications, 2025 – https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls 

3. European Commission, Regulatory Framework Proposal on Artificial Intelligence (AI Act), 2024 – https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai 

4. IBM Think Insights, Shadow AI Is Putting Your Data at Risk, 2025 – https://www.ibm.com/think/insights/shadow-ai-data-breach 

5. Dark Reading, Samsung Engineers Feed Sensitive Data to ChatGPT, Sparking Warnings About AI Use in Workplace, 2023 – https://www.darkreading.com/vulnerabilities-threats/samsung-engineers-sensitive-data-chatgpt-warnings-ai-use-workplace 

6. TechCrunch, Samsung Unveils ChatGPT Alternative Samsung Gauss That Can Generate Text, Code and Images, 2023 – https://techcrunch.com/2023/11/08/samsung-unveils-chatgpt-alternative-samsung-gauss-that-can-generate-text-code-and-images/ 

7. Extentum AI, AI Assistant in Business – How Does It Differ from an AI Agent, and Which Should You Choose in 2026?, 2026 – https://extentum.ai/extentum-ai-ai-assistant-vs-ai-agent-2026/