The privacy baseline for putting customer data in AI tools
Decide once which customer data can go into an AI tool, pick the account tier that matches, and check the settings that actually change your exposure.
on this page · 0 / 0 checked
The decision usually gets made by whoever is in a hurry. Someone pastes a customer’s email into a chat window to draft a reply, or wires a support inbox into an automation that summarises tickets overnight, and whether that was allowed gets settled after the fact by the fact that it happened. A privacy baseline is that decision made once, in advance, and made cheap enough that the person in a hurry can still follow it.
This is orientation, not legal advice. It is written for a solo operator or a team small enough that nobody owns compliance. If you hold health records, payment card data, or anything about children, or you are the controller for a large EU customer base, an hour with a lawyer beats this page. Everyone else needs four things: a rule for what goes in, an account tier that matches the rule, a contract wherever other people’s personal data is involved, and a list of the automations doing all of this while you sleep.
The account you are logged into decides most of it
The same sentence pasted into two accounts gets two different fates, and the difference is the plan, not the wording.
On Claude, Anthropic trains new models on data from Free, Pro and Max accounts when the training setting is on, including Claude Code used from those accounts, and told users they had until October 8, 2025 to make a choice [1]. With that setting on, data is retained in de-identified form for up to 5 years in model training pipelines [2]. With it off, the 30-day retention period applies [1]. Those consumer terms explicitly do not cover Claude for Work, Claude for Government, Claude for Education, or API use including via Amazon Bedrock and Google Cloud’s Vertex AI [1]. Deleting a conversation removes it from your chat history immediately and from back-end storage within 30 days, but conversations flagged for a usage policy review are kept for up to 2 years, and the safety classification scores for up to 7 [2].
OpenAI draws the same line in the same place. It lists ChatGPT Business, Enterprise, Healthcare, Edu, Teachers and the API Platform as the products whose data is not used to train models by default; on Business, admins control retention and deleted conversations are removed within 30 days, and API data is kept up to 30 days to provide the service and identify abuse [4]. The consumer plans are not on that list.
Google is the bluntest of the three. Its Gemini Apps notice tells you plainly not to enter confidential information you would not want a reviewer to see or Google to use to improve its services, because human reviewers read some of the data [6]. With Keep Activity on, activity auto-deletes after 18 months by default, adjustable to 3 or 36 months. With it off, chats are retained for 72 hours. And the part people miss: chats that have been through human review, plus related data such as your language and device type, are not deleted when you delete your activity, and are kept for up to 3 years [6].
Sort by the sentence you would have to write afterwards
The useful sort is not “sensitive” versus “not sensitive”, because everything feels a bit sensitive at 11pm. Sort by the sentence you would have to write if the text turned up somewhere public with your name on it.
Pile one is material where that sentence is an apology for an inconvenience: your own writing, published information, product copy, internal drafts with no personal details in them, examples you invented. Paste it. The judgment cost of thinking about pile one is higher than the risk.
Pile two is everything with a name attached that is not yours to give away, where the sentence would be a confession: client emails, customer lists, employee records, anything under an NDA you signed, credentials, invoices with amounts and parties on them. Pile two never goes in raw. It gets stripped first, or it stays out, or it goes into an account and tier that you chose specifically for it.
Notice that pile two is defined by ownership, not by how private the content feels. A dull contract you are not allowed to share is pile two. A distressing story you made up is pile one. People get this backwards because they sort by emotional weight, and emotional weight is not what a breach notice is about.
De-identification has to remove the shape, not the name
Replacing “Sarah Mitchell” with “Client A” is where de-identification starts, and it is not where it can stop. The details around the name do the identifying. A dental practice in a town of 9,000, a disputed amount, a treatment date and a solicitor’s involvement point at exactly one person with the name already gone, and they point harder because you were careful enough to signal that the case is real.
The pattern that works is to change the name, change the place, change the industry when it is not load-bearing, and round every number. “A service business is owed around five thousand by a long-standing client who has stopped replying” carries everything a model needs to draft the difficult letter. The original specifics added nothing to the output. They were only ever risk.
Two rules keep this honest. If de-identifying a document takes longer than the model saves you, the document belongs in pile two untouched and you write the thing yourself. And if you find yourself putting the real details back in “so it understands the situation properly”, you have not de-identified anything. You have delayed the paste.
Your retention window is a promise, not a law of nature
The retention number a vendor publishes is a promise about ordinary operation. It is not a guarantee that the data is gone, because the vendor is also subject to courts.
In the New York Times copyright litigation, OpenAI was ordered to retain all user content indefinitely going forward, including deleted ChatGPT chats and API content, overriding its normal deletion. ChatGPT Enterprise and Edu subscribers were excluded, as were API customers using zero data retention agreements [5]. The preservation obligations ended on September 26, 2025, and by October 22, 2025 OpenAI said it was no longer under a legal order to retain consumer content indefinitely, while limited April to September 2025 data is still stored securely in response to ongoing demands [5].
The durable lesson is not about one lawsuit. It is that “deleted after 30 days” describes what happens when nothing unusual is going on, and unusual things happen to large vendors constantly. Notice which tiers were carved out. Enterprise and education tiers and zero-retention agreements were not swept in, because those customers had contracts that made the data awkward to hold. Paying for the tier bought a narrower blast radius, which is a better reason to pay than the feature list.
The processing agreement is the entry requirement, not the upgrade
If you handle other people’s personal data as part of your work and you put it through an AI tool, you are the controller and the tool is your processor. Under GDPR Article 28, a controller shall use only processors providing sufficient guarantees to implement appropriate technical and organisational measures, and the processing must be governed by a contract that sets out the subject matter and duration, the nature and purpose, the type of personal data and categories of data subjects, and the obligations and rights of the controller [8]. That contract is the data processing agreement.
This is why the business tier is not a nicer version of the same thing. It is the version that comes with the paperwork. OpenAI says it can execute a DPA with customers for ChatGPT Business, ChatGPT Enterprise and the API in support of GDPR compliance, along with SOC 2 Type 2 audits and AES-256 encryption at rest [4]. The consumer plans are not named in that sentence, which is the whole point. Anthropic’s commercial terms cover Claude for Work and the API rather than the consumer plans [1]. Claude Team seats are $20 per seat per month billed annually, or $25 monthly, for teams of 2 to 150 [3], against $17 to $20 for an individual Pro subscription [3]. The gap between a consumer subscription and a covered seat is small enough that “we could not afford it” is rarely the real reason.
The automations leak more than the chat window
Most people picture a chat window when they think about this. The larger surface is everything running on a schedule. An automation in Zapier or n8n that reads new rows from your CRM and sends them to a model is pasting customer data hundreds of times a week under whichever account its connector authenticated as, and nobody watches it do that.
The same applies to AI built into tools you already pay for. Notion states that by default Notion and its AI subprocessors do not use customer data to train any models, that Enterprise plans get zero data retention with the LLM providers by default, and that non-Enterprise plans have a 30-day retention maximum before deletion [7]. That is a good posture, and it is also a good illustration of the chain: your document sits with the tool, the tool sends it to a model vendor, and the terms you get depend on which of the tool’s plans you are on.
So the audit is three questions per automation, and they are the same three every time. Which account does it run as, personal or company. What is the retention on that account’s tier. Is there a DPA covering it. Write the answers next to the automation’s name. The ones you cannot answer are the ones to turn off this week.
One page, three rules, and an answer within the hour
A policy nobody reads is decoration. What people follow is one page, kept where they work, with three rules on it.
The first rule is the piles. Pile one goes in freely and pile two gets stripped or stays out. The second rule is the account. Customer-identifying material only goes into the company workspace on the tier you chose, never a personal login, never someone’s phone. The third rule is the one most policies leave out: when unsure, ask, and the answer arrives within the hour.
That third rule is doing the real work. People route around friction, and every hour a question sits unanswered teaches somebody that the safe path is slow. A slow safe path guarantees a fast unsafe one. If you cannot answer within the hour, name someone who can, and put their name on the page.
Default is the Claude Team seat price billed annually [3]; put your own vendor's current price in. Computed in the page; nothing is sent anywhere.
What still goes wrong
Settings move. Anthropic changed its consumer terms and gave users a dated deadline to choose [1], and OpenAI’s ordinary deletion was overridden by a court order and then released from it, both inside 2025 [5]. A screenshot of a toggle from last year proves what was true last year. Put a recurring reminder in the calendar to re-read the retention page for each tool you rely on, and treat the date you checked as part of the fact.
Deletion is also less final than the button suggests. Google keeps human-reviewed chats for up to 3 years and does not remove them when you delete your activity [6], and Anthropic keeps policy-flagged inputs and outputs for up to 2 years [2]. Neither is a scandal. Both mean that “I deleted it” is a weaker sentence than it sounds, and that the decision not to paste is the only one that fully reverses.
Finally, this baseline does not touch the ordinary failures, which are still the likeliest ones. A phished account, a screen share with the sidebar open, a contractor working from a personal login, a laptop left unlocked. AI tools are third-party services that store what you send them, and they inherit every weakness your other accounts already have. The baseline shrinks what is in there to lose. It does not lock the door.
- 01Anthropic — Updates to Consumer Terms and Privacy Policyanthropic.com
- 02Anthropic Privacy Center — How long do you store my data?privacy.claude.com
- 03Claude — Pricingclaude.com
- 04OpenAI — Enterprise privacyopenai.com
- 05OpenAI — How we're responding to The New York Times' data demandsopenai.com
- 06Google — Gemini Apps and your datasupport.google.com
- 07Notion — Notion AI security practicesnotion.com
- 08GDPR Article 28 — Processor (text of Regulation (EU) 2016/679)gdpr-info.eu