customer data PII risk quantification

Customer Data in AI Prompts: Quantifying the Risk You Cannot See

Customer Data in AI Prompts: Quantifying the Risk You Cannot See

Customer data flows into AI prompts through two distinct channels, and most companies have visibility into neither. The first is the formal channel: employees using AI tools deliberately to work with customer data, drafting communications, summarizing records, generating analysis. The second is the contextual channel: customer data included in prompts not because the employee wants the model to process it, but because it is present in the context they are pasting for other purposes.

The contextual channel is the larger one by volume, and it is the harder one to address because employees often do not recognize when it is happening. A support rep pasting a ticket thread to ask Claude to draft a reply is focused on the drafting task, not on the fact that the ticket thread contains the customer's email address, account tier, and a description of a billing dispute. Those elements go along with the context. The employee is not deciding to share them. They are there.

The Channels and How They Generate Exposure

Customer support workflows are the highest-volume source of customer data in AI prompts at companies that handle consumer customers. The workflow pattern is consistent: agent receives ticket, reads context, formulates response, uses AI to draft the response faster, pastes ticket content to provide context for the draft. The ticket content contains the customer's contact information, account details, the specifics of their issue, and often prior communication history. All of that transmits with the paste.

Under GDPR, an AI inference endpoint that processes personal data on behalf of the data controller is a data processor under Article 28. That creates a contractual obligation: you need a data processing agreement with the vendor before personal data can be transmitted to their systems. If your customer support team is using ChatGPT with no DPA in place, every ticket that includes personal data is potentially a compliance gap. If you have a DPA in place, you are covered contractually, but you still need to be able to demonstrate to a regulator that you have controls on what personal data is being transmitted and how. "We have a DPA" is not the same as "we have controls."

Under CCPA and CPRA, the equivalent obligation is a service provider agreement with a data processing limitation clause. Similar structure, similar compliance gap when the AI tool is being used without a formal contract that covers personal data processing.

Sales and CRM Data: Volume Underestimated

Sales workflows generate a category of AI prompt exposure that tends to get less attention than customer support but operates at high volume. The workflow: sales rep has a call coming up, pulls up the CRM account record, pastes some or all of it into an AI chat interface for call prep. The CRM record contains: contact name, title, company, email, phone, deal stage, revenue figure, prior meeting notes, integration data from email and calendar.

That is a dense personal data payload per paste. At a 150-person company with an active outbound sales team, the volume of CRM data moving through AI prompts during call prep and follow-up activities is substantial. The individual incidents are small. The aggregate is not.

The GDPR concern here is compound. The contact-level personal data (name, email, phone) triggers basic Article 4 obligations. The content of meeting notes and prior conversation summaries may include additional personal information about the contact that was gathered during earlier stages of the relationship. Deal stage and revenue figures are not personal data, but they may be subject to confidentiality obligations in vendor agreements or investor agreements. The same paste that raises a GDPR flag may also raise a contractual confidentiality flag depending on the content.

Developer Workflows and Customer-Adjacent Data

Developer AI tool usage generates a less obvious but significant customer data exposure pathway. When engineers work with production data in debugging or optimization contexts, and use AI tools to analyze query results, optimize data pipelines, or debug customer-reported issues, customer data can appear in prompts without ever being intentionally included.

Consider a developer debugging a slow database query for a customer reporting slow load times. The developer pulls a sample query result set to analyze the performance issue. The result set includes customer identifiers, account fields, and timestamps. The developer pastes the result set into ChatGPT to ask for query optimization advice. The optimization question is about the query structure. The customer data in the result set is context that went along for the ride.

This pattern is not malicious and is often not recognized as a data governance event. From the developer's perspective, it is a debugging session. From a GDPR or CCPA perspective, it is personal data being transmitted to a third-party processor without a clear lawful basis for that specific processing activity. The gap between those perspectives is where the compliance exposure lives.

Quantifying Without a Content Inspection Layer

Without prompt content inspection, quantifying how much customer data is leaving through AI prompts requires estimation from indirect signals, and the estimates tend to be conservative.

The estimation approach we recommend is workflow-based: identify the employee roles and workflows that involve regular customer data access, estimate the AI tool usage rate for those workflows based on traffic volume to AI endpoints from those organizational units, and apply a conservative estimate of the proportion of those prompts that likely include customer data based on the workflow characteristics.

For a support-heavy operation, the fraction of support-team AI prompts that include some form of customer data is high, likely above 60% for teams that primarily use AI for response drafting, based on the workflow structure. For developer teams, the fraction is lower but nonzero. For sales teams, it depends heavily on whether the primary use case is call prep (high customer data content) or internal communications (lower).

These estimates are not precise. They are directional. The point of the exercise is not to produce a number you can put in an audit report. It is to establish that the volume is material enough to warrant investment in actual measurement rather than continued estimation.

The Measurement Problem Is Not Unsolvable

Measurement requires a layer that can see prompt content. Traffic monitoring at the destination level tells you which AI endpoints are receiving traffic. It does not tell you what is in the requests. For organizations that want a concrete rather than estimated picture of customer data volumes in AI prompts, the measurement layer needs to operate at the content level.

The measurement phase does not need to begin as enforcement. Deploying prompt classification in logging-only mode for 30 to 60 days, without any blocking or alerting, produces a content profile of what is flowing. That profile tells you which employee populations are generating customer data in prompts, which data categories are most prevalent, and which AI tool destinations are receiving the highest-sensitivity traffic. That data is the foundation for designing a proportionate governance program rather than one based on assumptions about what employees are doing.

We are not suggesting that measurement solves the compliance problem. If you discover that your support team has been transmitting GDPR-regulated personal data to an AI tool without a DPA in place, the measurement tells you the extent of the gap but does not close it retroactively. It does, however, give you the information needed to prioritize the gap remediation and design controls that address the actual exposure rather than hypothetical scenarios.

The Risk You Cannot See Is Still the Risk You Own

There is a category of security posture that could be described as "we have not seen an incident, therefore our controls are sufficient." For AI prompt data egress, this posture is particularly fragile because the absence of visible incidents does not indicate the absence of exposure. If you have no mechanism to see prompt content, you cannot observe whether incidents are occurring. The absence of observation is not evidence of absence.

Regulators are increasingly aware of this gap. Data protection authorities in several EU member states have begun asking specifically about AI tool data processing in audit contexts. The question is not "have you had an incident" but "do you have controls and can you demonstrate what data is flowing to AI systems." Organizations that have been relying on the absence of incidents as evidence of adequate controls are encountering that question with limited answers.

The same dynamic applies to contractual obligations. Many customer agreements include data processing limitations that restrict how customer data can be shared with third parties. AI inference endpoints are third parties. Whether a specific AI vendor counts as a third party for contract purposes depends on the agreement language, but the risk that customer data in AI prompts creates a contractual breach is not theoretical. It is a live question for any company that handles customer data under contracts with data processing restrictions.

See Unbound in action on your AI stack.

30-minute live session. We deploy, run detection, and walk through findings with your security team.

Request Demo