research / shadow-ai-enterprise-data-leakage
Shadow AI at work: what employees paste into AI tools
Banning AI tools pushes usage out of sight. A practical approach to visibility, data protection, and policy that lets teams use AI safely.
Every organisation now has an AI policy, whether it wrote one or not. If there's no approved tool, employees use their personal accounts. If there's an approved tool but it's slow or limited, they use something else. This is shadow AI, and it's the most common route for sensitive data to leave a company through an AI system.
What actually leaks
The risky moments are ordinary ones:
- A developer pastes a stack trace that contains a database connection string.
- Someone in sales uploads a customer list to "clean up the formatting."
- Legal asks a chatbot to summarise a contract under NDA.
- Support pastes a ticket thread full of personal data to draft a reply.
None of this is malicious. It's people trying to work faster. But depending on the tool, the account type, and the provider's terms, that data may be retained, reviewed, or used for training. It also sits outside your logging, retention, and incident-response processes.
Why blocking alone fails
Blocking well-known AI domains at the proxy feels decisive, but it has three problems:
- The list never ends. New AI tools, browser extensions, and AI features inside existing SaaS products appear every week.
- It moves usage to personal devices, where you have no visibility at all.
- It gives up real productivity gains, and the business will push back, usually with good reason.
Start with visibility, then make the safe path easy
First, find out which AI tools people use, how often, and from which teams. You can't write a sensible policy without that picture, and it's usually bigger than anyone expects.
Next, give people a good sanctioned option: an enterprise plan with clear data-handling terms, single sign-on, and admin controls. Most shadow AI disappears when the approved path is also the easiest one.
Inspect data where it enters the AI tool
The point of use is where protection works. That means the browser for chat tools, and the endpoint for desktop apps, coding agents, and scripts that call model APIs directly. It's where you can catch source code, credentials, personal data, and classified documents before they're submitted. Then match the response to the risk:
- Coach for low-risk cases ("This looks like customer data. Did you mean to use the company workspace?").
- Redact for medium-risk cases, replacing secrets or identifiers with placeholders.
- Block for high-risk content such as credentials, regulated data, or material under legal hold.
Finally, write a short policy. One page is enough: which tools are approved, which data types never go into any AI tool, and who to ask. Nobody reads a long one.
Questions for your next security review
- Do we know which AI tools are used here, and by how many people?
- Is there a sanctioned option that's genuinely good enough?
- Could we detect an API key pasted into a chatbot today?
- Are AI-assisted features inside our existing SaaS covered by our vendor reviews?
AI use at work is not going away. The goal is to make the safe path the easy path, and to have visibility when it isn't.