The AI data shortage, and who gets to train on your text
Clean human text is the scarce input in AI now, which makes your chats and documents worth something. Here is how to decide who keeps them.
on this page · 0 / 0 checked
404 Media placed a tracking device in a rare book and followed it to an Amazon facility in Las Vegas known as VGT3, which identifies itself with a symbol of a dinosaur holding a book in its claws. Books that arrive there have their spines cut off so they can be scanned, which destroys them. Amazon told 404 Media that it “purchases books through commercial channels to improve the products and services customers use” [8]. Strip away the strangeness and what is left is a procurement decision. Text that no model has read before is now worth buying a book, shipping it across the country and taking it apart.
You produce that kind of text every working day, in smaller quantities and for free. Proposals, client emails, meeting notes, half-finished drafts, the long explanation you typed into a chat window at 11pm because it was faster than writing a brief. None of it was ever on the open web. All of it is exactly the material that has become expensive. This guide is not about whether the shortage is real, and it is not legal advice for anyone with a copyright claim to bring or a signed data processing agreement to enforce. It is for the solo operator or small-team owner who has been pasting real client work into a free account for two years and has never opened the privacy tab. The useful question is narrow: which of your accounts treats your text as raw material, and which does not.
The shortage is about clean human text, not compute
Chips get built and data centres get financed. The supply of human writing does not respond to capital the same way. Epoch AI put the effective stock of quality- and repetition-adjusted human-generated public text at roughly 300 trillion tokens, and projected that language models would fully use that stock between 2026 and 2032, sooner if models are heavily overtrained [6]. That estimate is from June 2024 and it is a projection rather than a measurement, but it fits the behaviour you can observe now: a company buying rare books nobody has digitised and cutting them apart to get the text out [8].
Two properties make old text valuable. It is human-written, and it predates the point where models started filling the internet with their own output. That is the reason the reporting gives for targeting old books specifically. Nothing published before 2022 could have been written by a language model, and models trained on AI-generated text risk “model collapse”, where the quality of their output degrades [8]. It also explains why your archive matters more than its word count suggests. Business writing is human, unpublished, specific, and attached to real decisions with real outcomes, which is the profile of text that is hardest to synthesise and most expensive to buy.
Your chat history is the cheapest fresh text available
Nobody has to buy your writing, because you type it into a text box voluntarily. Whether that ends up in a training pipeline is decided by one setting per account, and the setting is not in the same place or under the same name in any two products.
In Claude, the toggle is called “Help Improve our AI models” and it sits under Settings then Privacy. When it is on, Anthropic uses your new chats and coding sessions to improve its models. When it is off, the company states it “will not use any new chats and coding sessions you have with Claude for future model training”, with the standing exception that conversations flagged by safety classifiers may still be used for trust and safety work [1].
In ChatGPT, the equivalent lives under Settings then Data Controls and is called “Improve the model for everyone”. Turning it off does not hide your history from you. As OpenAI puts it, “Your conversations will still appear in your chat history but won’t be used to train ChatGPT” [3]. Temporary Chats are a separate mechanism and are deleted from OpenAI’s systems after 30 days [3].
In Gemini, the control is Keep Activity. With it on, Google saves your chats and uses that activity to “provide, develop, and improve its services (including training generative AI models)”, with auto-delete set to 18 months by default, changeable to 3 or 36 months, or to no auto-delete at all [5]. With it off, your future chats “won’t appear in your Activity, and won’t be used to train our AI models, unless you choose to send Google feedback”, though they are still held for 72 hours to run the service [5]. Google also states that a subset of chats are read by human reviewers, that those chats are disconnected from your account first, and that reviewed chats “are not deleted when you delete your activity” and are kept for up to three years [5]. That last detail is the one people mis-model. Deleting a conversation from your side is not a recall order.
Business tiers start from the opposite default
The paid tiers are built for buyers who have to answer a client’s security questionnaire, so they invert the assumption. OpenAI states plainly that “By default, we do not train on any inputs or outputs from our products for business users, including ChatGPT Business, ChatGPT Enterprise, and the API” [4]. Anthropic’s Commercial Terms of Service, effective 17 June 2025, put it as an obligation rather than a preference: “Anthropic may not train models on Customer Content from Services” [2].
The practical consequence is about routing, not about which product is better. If the same person holds a personal Claude or ChatGPT login and a business seat, the client work belongs in the business seat, and it stays there even when the personal account is the one already open in the other tab. That single habit does more than any settings audit, because it survives the vendor changing a default. It is also the answer to give when a client asks where their material goes. You can point at a terms page with an effective date on it [2][4], which is a different class of answer from describing a toggle you remember flipping.
Two cautions. A business plan protects what goes through that account, not what you pasted into a free account last year. And the assistant is not the only tool holding your text. Notes apps, transcription services, CRM add-ons and automation platforms all sit on the same material, and each has its own clause. The tools with the least scrutiny are usually the ones that arrived as a free integration.
For those tools there is a method that survives any product rename. Open the terms or the privacy page, search the page for “train”, and read the sentence around every hit. You are looking for three things: whether training is described as something the vendor may do or may not do, whether the paid tier is named separately from the free one, and what date the page carries. A page with no occurrence of the word at all is a finding rather than a relief, because it usually means the answer lives in a linked sub-processor list or a policy that changed since the marketing copy was written. Ten minutes per tool, once a year, done in the order of how sensitive the material is.
One more limit worth stating out loud. Turning a toggle off protects you, not your client. If you are handling someone else’s confidential material, the account it passes through is a term of your engagement with them, and the honest version of that conversation happens before the work starts rather than after a breach notice.
Your website is the other door into your work
Everything published under your own domain sits outside the account settings entirely, and that is where the market has moved fastest. On 1 July 2025 Cloudflare announced it was “changing the default to block AI crawlers unless they pay creators for their content” [7]. The stated reason was traffic. The announcement put the ratios bluntly: with OpenAI, “it’s 750 times more difficult to get traffic than it was with the Google of old”, and with Anthropic, “it’s 30,000 times more difficult” [7].
For most small operators the correct setting is not obviously “block everything”. If your site exists to be found, being absorbed into an assistant’s answers can still send work your way, and blocking removes that possibility along with the extraction. If your site is a body of expertise that competitors would pay for, the calculus flips. The point is to make it a decision with a date attached rather than a default you inherited from whoever set up hosting. Check what your CDN or host currently does with AI crawlers, choose, and write down why.
Buy on what a model does today, not on what more data might do
The last implication is about spending. When better models came mostly from more text, it was reasonable to assume next year’s version would be meaningfully stronger and to buy on that assumption. A shrinking pool of untouched human writing weakens that assumption without disproving it. Epoch’s projection is about one input running out [6], not a forecast that capability stops improving, and gains can still come from curation, from training longer on the same material, and from work done after pre-training. The honest position is that the cheap, obvious source of improvement is closing and what replaces it is not settled.
So treat a vendor’s data story as a real signal alongside price and speed. A company willing to buy rare books and destroy them to scan the pages is building a supply its competitors cannot copy for free [8]. And when you sign anything annual, price it against the model you can test this week. If the tool is worth the money at today’s capability, the contract is safe. If it only makes sense because next year’s release will be better, you are buying a projection, and the input that projection depends on is the thing this whole guide is about.
chats × words × working days. Your own numbers, not an estimate of anyone's average. Computed in the page; nothing is sent anywhere.
What still goes wrong
The settings only bind the future. Anthropic’s control applies to new chats and coding sessions from the point you change it [1], and Google keeps human-reviewed chats for up to three years whatever you delete [5]. Material already inside a completed training run does not come back out, and no vendor offers a mechanism that would make it come back out. Everything in this guide is a decision about what happens next, not a way to undo what happened before you read it.
Policies also move faster than guides do. Every name and number here was checked against the vendors’ own pages on 4 September 2026, and each of those pages carries its own last-updated date that will keep advancing [1][3][5]. A toggle can be renamed, a default can flip during a terms update, and the notification that tells you arrives as a product email you will not read. Re-check the three settings when you renew a plan, and treat any prompt to accept new terms as the moment to look rather than the moment to click through.
The larger limit is that none of this touches the leverage question. A shortage of human text means your writing has value to somebody, but nothing here converts that into money for a business your size. Cloudflare said its next step would be “a marketplace where content creators and AI companies, large and small, can come together” [7], which is a plan rather than a market you can sell into today, and in any case it covers published web content rather than the unpublished material sitting behind your login. What you get for now is control rather than income, which is a smaller prize, and still worth the 20 minutes.
- 01Anthropic — How do I change my model improvement privacy settings?privacy.claude.com
- 02Anthropic — Commercial Terms of Serviceanthropic.com
- 03OpenAI — Data Controls FAQhelp.openai.com
- 04OpenAI — How your data is used to improve model performancehelp.openai.com
- 05Google — Gemini Apps Privacy Hubsupport.google.com
- 06Epoch AI — Will we run out of data? Limits of LLM scaling based on human-generated dataepoch.ai
- 07Cloudflare — Content Independence Day: no AI crawl without compensationblog.cloudflare.com
- 08TechCrunch — Amazon, which started off selling books, is destroying rare texts to train AItechcrunch.com