AI copyright lawsuits are not one legal dispute playing out in different courtrooms; they split into at least two genuinely separate questions. The first is whether training a model on copyrighted material without permission is itself infringement, or whether it qualifies as fair use because the model does not store or reproduce the work directly. The second is whether a specific generated output improperly reproduces or closely imitates a protected work, which is a narrower and more fact-specific claim. Marquee cases like The New York Times v. OpenAI and Microsoft, Authors Guild v. OpenAI, and Getty Images v. Stability AI have kept both questions in front of courts at once. Rulings have differed depending on which question was actually being decided, and depending on how the training data was obtained in the first place, which is why there is no single answer to "is AI training legal."
What changed in 2026
- The training-data question got a real split, not a clean answer. Multiple courts have found that training on lawfully acquired copyrighted works can qualify as transformative fair use, while separately finding that training on pirated copies of the same works does not get the same protection, regardless of how the resulting model is used.
- At least one headline case resolved through settlement rather than a final verdict. Anthropic agreed to settle Bartz v. Anthropic, reported in the range of a billion dollars or more with a class of authors over claims that pirated book copies from shadow-library sites were used to build training data, one of the largest copyright settlements in history and a signal that statutory damages exposure at scale is a real business risk, not a theoretical one.
- Output-side cases against image generators kept moving separately from text cases. Suits from stock-image and entertainment companies against image and video generators have focused more narrowly on specific outputs resembling copyrighted characters or photos, a distinct legal theory from the training-data fair use fights.
- Licensing deals accelerated as a parallel track to litigation. Several major publishers and rights holders signed direct licensing agreements with AI companies, sometimes while litigation with other companies continued, suggesting the industry is settling into a mixed model of licensing plus litigation rather than one clean legal resolution.
The major case categories
| Category |
What is being argued |
Roughly how it has trended |
| Training-data fair use (lawfully acquired data) |
Whether training on legally obtained copyrighted works is transformative fair use |
Leaning favorable to AI companies in several rulings, though not uniformly |
| Training-data provenance (pirated or scraped data) |
Whether using unauthorized or pirated copies to build a training set is separately unlawful |
Leaning unfavorable to AI companies; this is where the largest settlement exposure has appeared |
| Output infringement |
Whether a specific generated image, character, or passage improperly reproduces a protected work |
Highly fact-specific, decided case by case rather than as a general rule |
| Cross-border training claims |
Whether infringement occurred if the training itself happened outside the plaintiff's jurisdiction |
Some claims narrowed or dismissed on jurisdictional grounds even when the underlying conduct was contested |
Why fair use arguments split the way they did
The transformative-use argument that helped some AI companies rests on the idea that a model trained on a work learns statistical patterns rather than simply redistributing the work, similar to reasoning that once protected search-engine indexing. That argument tends to fail when the acquisition method was unlawful, or when the resulting product competes directly with the original's market. Thomson Reuters v. ROSS Intelligence is an early example of the second failure mode: a court found that training a competing legal research tool on Westlaw's editorial headnotes was not fair use, largely because the product directly substituted for the original in the same market. A court can find training transformative while still finding that obtaining the underlying copies through piracy is separately punishable, and that split is roughly where several rulings have landed. In practice, this means legal exposure depends heavily on data-sourcing practices and market impact, not just on what the model ultimately does with the data.
Common mistakes
- Treating any single ruling as the final word on AI and copyright. Rulings so far have been narrow, fact-specific, and sometimes explicitly limited to the parties involved, not general precedent covering every AI company or use case.
- Ignoring data provenance as a legal risk. Businesses building on top of models or fine-tuning on scraped data sometimes overlook that how training data was obtained can matter more than how the model is later used.
- Assuming a licensing deal by a large competitor de-risks smaller players. Licensing agreements are negotiated case by case and do not extend legal protection to companies that were not party to them.
- Conflating training-data claims with output-infringement claims. These require different evidence and different legal tests; a strong defense on one does not automatically cover the other.
FAQ
Has any court definitively ruled that AI training is fair use?
Some courts have ruled that training on lawfully acquired material was transformative fair use in the specific cases before them, but these rulings are fact-specific and have not been treated as settling the question for the industry broadly.
Why did Anthropic settle instead of going to trial?
The settlement followed a ruling that separated a favorable finding on training-method fair use from an unfavorable finding on using pirated copies to build the dataset, which exposed the company to potentially enormous statutory damages if the pirated-book claims went to a full trial.
Do artists and authors have a real path to compensation outside of suing?
Increasingly yes. Licensing agreements between AI companies and publishers, image libraries, and other rights holders have become more common as a negotiated alternative or complement to litigation.
How does this affect a business just using AI tools, rather than building them?
Direct legal exposure is lower for end users than for model builders, but businesses generating commercial content with AI tools should understand a vendor's data-sourcing practices and licensing terms, since that context can matter if output-infringement questions arise later.
Where to go next
For related legal and provenance questions, see our guides to AI regulation by country, what an AI audit actually involves, and how AI content watermarking works.