The copyright risk you actually carry when you use AI
Training-data lawsuits target the model makers, not you. This sorts out the risk that is genuinely yours and shows the settings and clauses that shrink it.
on this page · 0 / 0 checked
A publisher sues a model maker over training data, the coverage arrives with a number attached that has a lot of zeroes in it, and the question reaches your desk in a distorted shape. If you are a freelancer who drafts client copy in ChatGPT, or a two-person shop that runs research through Gemini and writes in Claude, the honest reaction sits somewhere between “this has nothing to do with me” and “should I be worried about the thing I shipped last week.” Both reactions are half right. The half that matters is not the one in the headline.
You are not going to be a defendant in a training-data case. Those cases are about what a lab copied to build a model, and you did not copy anything to build a model. What you did do is generate output, publish it, invoice someone for it, and probably sign something promising the work was yours to sell. That is a different question with a different answer, and it turns less on litigation news than on which plan you are paying for and what you agreed to when you opened the account. This guide is for people in that position. It is not legal advice, and it is not for anyone who has already received a demand letter. At that point you want a lawyer, not a guide.
The lawsuits are upstream of you
Take a concrete one. On 10 July 2026, Hachette Book Group, Cengage Learning, Elsevier, the author Scott Turow and S.C.R.I.B.E., Inc. filed a proposed class action against Google in the Southern District of New York, case number 26-cv-5870 [1]. The complaint pleads four counts. Three are ordinary copyright claims under 17 U.S.C. §§ 106(1) and 501, covering unauthorized reproduction through Google Books and affiliated services, unauthorized reproduction through web scraping, and unauthorized reproduction in the training of Google’s AI models. The fourth is a claim under § 1202(b) of the DMCA for the removal or alteration of copyright management information, the identifying information attached to the works themselves [1].
Read the counts and notice what is missing. Every one of them describes something Google allegedly did to copies of books and articles, before any customer typed a prompt. Nobody in that filing is complaining about a small business that used Gemini to draft a proposal. The defendant is the company that assembled the corpus.
That does not make the litigation irrelevant to you, but it narrows the relevance to one thing. A damages award against a lab is a problem for the lab’s balance sheet. The remedy that would actually reach your desk is one that changes the product: an order restricting what a model may keep training on, or requiring a version to be retrained or withdrawn. That is a continuity risk, the same category as a vendor raising prices or deprecating an endpoint, and you manage it the same way, by not building a business that only works if one specific model stays available on today’s terms.
Your exposure runs through the output and the promises you sign
The claim that could name you looks nothing like the Hachette complaint. It looks like this: something you published under your own name, or your client’s name, reproduces enough of someone’s protected work that they want money for it. Whether a model was involved is almost beside the point. You published it.
The number that decides whether such a claim is worth bringing is set per work. Under 17 U.S.C. § 504(c), a copyright owner may elect statutory damages instead of proving actual losses, and the court awards not less than $750 and not more than $30,000 for each work infringed. If the infringement is willful the ceiling rises to $150,000 per work, and if the infringer proves it did not know and had no reason to believe its acts were infringing, the floor can drop to $200 [8]. Those are the multipliers. A single borrowed paragraph is a nuisance. The same mistake repeated across 40 published pieces is a different conversation.
17 U.S.C. § 504(c) sets $750 to $30,000 per work, rising to $150,000 for willful infringement and falling as low as $200 for innocent infringement [8]. Computed in the page; nothing is sent anywhere.
The second half of your exposure is contractual and entirely self-inflicted. If the agreement you signed with a client warrants that the deliverable is original and does not infringe anyone’s rights, you have taken the risk onto yourself by signature. No vendor term overrides that. It is worth knowing which of your contracts contain that clause before you find out the expensive way.
Indemnity is a feature of the plan you pay for
OpenAI, Anthropic and Google all undertake to defend you against a third-party IP claim, and all three attach that promise to commercial terms rather than consumer ones [2][4][5].
OpenAI’s Service Terms, effective 12 June 2026, cover API customers and the enterprise products, listed as ChatGPT Enterprise, Edu, Healthcare and Business. The commitment is a defence against “any third party claim that Customer’s use or distribution of Output infringes a third party’s intellectual property right” [2]. Anthropic’s Commercial Terms of Service, effective 17 June 2025, say the customer owns its Outputs and that Anthropic assigns whatever right, title and interest it has in them, then undertake to defend claims alleging that “Customer’s paid use of the Services (which includes data Anthropic has used to train a model that is part of the Services) in accordance with these Terms or Outputs generated through such authorized use violates any third-party intellectual property right” [4]. That parenthetical is doing real work. It pulls training-data allegations inside the definition of a claim Anthropic has agreed to defend.
Google publishes an explicit list. The Generative AI Indemnified Services page, last updated 20 July 2026, names Gemini for Google Cloud, the Gemini Enterprise Agent Platform API with the Gemini, Imagen, Veo, Codey and PaLM model families, Gemini Enterprise, NotebookLM Enterprise, Gemini in Workspace and Google Vids, among others [5]. Google splits the promise into two named indemnities rather than one. The training data indemnity “covers any allegations that Google’s use of training data to create any of our generative models utilized by a generative AI service, infringes a third party’s intellectual property right”, and a separate generated output indemnity covers what customers produce with the tools [6]. That first prong is the one that maps directly onto the lawsuits in the headlines.
Now the part people miss. OpenAI’s consumer Terms of Use, effective 1 January 2026, assign you the output in the usual way: “you (a) retain your ownership rights in Input and (b) own the Output” [3]. They contain no matching promise to defend you. They contain the reverse. If you are a business or organisation, you agree to “indemnify and hold harmless us, our affiliates, and our personnel, from and against any costs, losses, liabilities, and expenses (including attorneys’ fees) from third party claims arising out of or relating to your use of the Services and Content” [3]. On the consumer tier, the protection runs from you to the vendor. Google’s indemnified list works the same way by omission: the free Gemini app is not on it [5].
The exclusions are where the indemnity stops
An indemnity is a promise to defend a claim, not a promise that no claim exists, and each one is fenced. OpenAI’s does not apply where the customer knew or should have known the output was infringing, where safety features were disabled or ignored, where the output was modified or combined with non-OpenAI products, where the customer lacked rights to the input materials, where the claim is a trademark claim arising from commercial use, or where the infringing content came from a third-party offering. Beta services are offered “as-is” and “are excluded from any indemnification obligations OpenAI may have to you” [2].
Anthropic’s carve-outs are close to identical in shape: modifications the customer made to the Services or Outputs, combination of the Services or Outputs with technology or content Anthropic did not provide, use “in a manner that Customer knows or reasonably should know violates or infringes the rights of others”, the practice of a patented invention contained in an Output, and trademark-based use of an Output in trade or commerce [4]. Google states its own condition plainly: the generated output indemnity “only applies if you didn’t try to intentionally create or use generated output to infringe the rights of others, and similarly, are using existing and emerging tools, for example to cite sources to help use generated output responsibly” [6].
Two exclusions appear in all three. The first is knowledge. Prompting a model for something in the style of a named living author and then selling the result is the textbook way to land outside every one of these clauses. The second is modification and combination, and that one deserves a slow read, because editing a draft and dropping it into a page you built is what everybody does with AI output all day.
AI-only output is not yours to protect
The other side of the ledger is what you own, and here the position has been stable for years. The U.S. Copyright Office’s registration guidance holds that “copyright can protect only material that is the product of human creativity”. Where an AI technology “receives solely a prompt from a human and produces complex written, visual, or musical works in response”, the guidance treats the traditional elements of authorship as determined and executed by the technology rather than the human user, and where the technology determines those expressive elements, “the generated material is not the product of human authorship” [7].
That has two consequences worth planning around. First, the ebook, course outline or template library you generated wholesale is not an asset you can stop a competitor from lifting. Human-authored modifications and creative selection or arrangement of AI-generated material can be protectable, but only the human contribution gets protection [7]. If a body of work is meant to be defensible, the human work has to be real and you have to be able to describe it.
Second, if you ever register, disclosure is not optional. Applicants “have a duty to disclose the inclusion of AI-generated content in a work submitted for registration and to provide a brief explanation of the human author’s contributions”, and AI-generated content should be named in the “Material Excluded” field [7]. Skipping that is worse than not registering at all. The Office may cancel a registration where information essential to its evaluation of registrability was omitted, and a court may disregard the registration in an infringement action if the applicant knowingly provided inaccurate information [7].
Habits that keep you inside the coverage
None of this requires a compliance function. It requires that the tier you work on matches the work you do. Anything you publish, sell or hand to a client should be generated on a paid business plan whose terms name an IP indemnity, and you should check that the specific product you are using appears on the vendor’s covered list rather than assuming the brand covers everything wearing its name [5].
Keep the prompt and the raw draft next to the finished file. If a claim ever arrives, the whole fight over the knowledge exclusion turns on what you can show about what you asked for and what you did with the answer, and a folder of prompts is cheap insurance against an argument you would otherwise have to make from memory. Turn on citation and grounding features where the tool offers them, because at least one vendor has written the use of those tools into the condition of its indemnity [6]. Before anything distinctive goes out under someone else’s name, put a couple of its more unusual phrases through a search engine. And read the warranty clause in the next client contract you sign, because that clause, not the model’s terms of service, is the one that decides who pays.
What still goes wrong
The modification and combination exclusion is the loose thread. OpenAI’s terms exclude output that was “modified, transformed, or used in combination with products or services not provided by or on behalf of OpenAI”, and Anthropic’s exclude modifications the customer made and combination with technology or content Anthropic did not provide [2][4]. Taken at face value, editing a draft and publishing it inside your own site could sit on the wrong side of that line, which would make the indemnity narrower in practice than it reads in a press release. No court has told anyone how wide those carve-outs really are. Treat the indemnity as a useful backstop rather than a licence to stop paying attention.
Everything above carries an effective date, and the dates move. OpenAI’s Service Terms are dated 12 June 2026, its consumer Terms of Use 1 January 2026, Anthropic’s Commercial Terms 17 June 2025, and Google’s indemnified services list was last updated 20 July 2026 [2][3][4][5]. Those documents get revised, sometimes narrowed, usually without an announcement. If your business depends on a clause, put a calendar reminder on it and re-read the clause rather than a summary of it.
Finally, the scope here is narrow on purpose. This is U.S. copyright, and it says nothing about trademark, right of publicity, trade secrets, your obligations under a client’s own AI policy, or how any of this works outside the United States. It is also no use to anyone training a model on scraped data, which is a materially different risk with materially different defendants. If you have an actual claim in hand, stop reading guides.
- 01Hachette Book Group et al. v. Google LLC — Complaint (S.D.N.Y., No. 26-cv-5870)publishers.org
- 02OpenAI — Service Termsopenai.com
- 03OpenAI — Terms of Useopenai.com
- 04Anthropic — Commercial Terms of Serviceanthropic.com
- 05Google Cloud — Generative AI Indemnified Servicescloud.google.com
- 06Google Cloud — Protecting customers with generative AI indemnificationcloud.google.com
- 07U.S. Copyright Office — Copyright Registration Guidance: Works Containing Material Generated by Artificial Intelligencecopyright.gov
- 0817 U.S. Code § 504 — Remedies for infringement: Damages and profitslaw.cornell.edu