Is Training AI on Copyrighted Books Legal? U.S. Courts Are Still Drawing the Lines

U.S. courts are beginning to define when copyrighted books can be used to train AI, with fair use, piracy and market harm emerging as key legal issues.

Aug 23, 2026 - 13:49
Aug 23, 2026 - 14:00
 2
Is Training AI on Copyrighted Books Legal? U.S. Courts Are Still Drawing the Lines
Image Credit: TechAmerica.ai / AI-generated image

Whether artificial intelligence companies can legally train models on copyrighted books remains one of the most consequential unresolved questions surrounding generative AI. Early U.S. court decisions suggest that the answer can depend not only on what a company does with a book, but also on how it obtained the material and whether its use harms an existing market.

The distinction became particularly important in litigation involving Anthropic. In Bartz v. Anthropic, U.S. District Judge William Alsup ruled that using copyrighted books to train large language models could qualify as fair use because the training was transformative. But he reached a different conclusion about Anthropic’s creation of a permanent library containing millions of books obtained from pirate websites.

AI training and piracy are separate copyright questions

Anthropic ultimately agreed to a $1.5 billion settlement covering claims involving pirated books. The outcome did not establish that training AI on copyrighted books is inherently illegal. Instead, it highlighted a distinction between training a model and obtaining or retaining unauthorised copies of copyrighted works.

The financial impact is substantial, although it must be viewed in the context of the rapidly growing AI industry. Reuters reported that Anthropic is projecting revenue of roughly $190 billion to $200 billion for 2028, according to people familiar with the company’s financials.

The broader legal debate centres heavily on fair use, a doctrine that permits some uses of copyrighted material without the copyright owner’s permission. Courts generally consider four factors, including the purpose and character of the use, the nature of the copyrighted work, how much material was used and the effect on the potential market for the original.

That framework comes from the Copyright Act of 1976, the foundation of current federal copyright law. The statute has been amended repeatedly since its passage, but courts are now applying its principles to AI systems and training methods that did not exist when the law was written.

Competition can change the fair-use analysis

A separate case involving Thomson Reuters illustrates why the commercial purpose of an AI system can matter. Thomson Reuters sued Ross Intelligence after Ross used material derived from Westlaw headnotes to develop an AI-powered legal research service.

In a 2025 decision in Thomson Reuters v. Ross Intelligence, Judge Stephanos Bibas rejected Ross’s fair-use defence. He found that Ross’s use was commercial and not transformative because it used Thomson Reuters material to help develop a legal research tool competing with Westlaw.

The ruling is important but does not settle whether training generative AI models such as ChatGPT or Claude is fair use. Bibas specifically noted that the Ross system before the court was not generative AI, limiting how broadly the decision can be applied to other AI copyright disputes.

Those different outcomes show why there is no simple rule that all AI training on copyrighted material is either legal or illegal. A court may look differently at a model trained to learn patterns from legitimately obtained works than at a company using unauthorised copies to create a directly competing product.

AI-generated content raises a different copyright issue

The law governing AI-generated material is also distinct from the rules governing training data. In Thaler v. Perlmutter, the U.S. Court of Appeals for the District of Columbia Circuit upheld the Copyright Office’s refusal to register an image whose sole listed creator was an AI system.

The court held that the Copyright Act requires a human author for a work to qualify for copyright protection. Importantly, it did not hold that everything made with AI assistance is automatically ineligible. The court noted that works made with the assistance of AI may still qualify depending on the human creative contribution and how the technology was used.

That makes copyright questions involving AI highly fact-specific. Training models on copyrighted material, acquiring pirated training data, generating material similar to existing works and seeking copyright protection for AI-assisted creations are related issues, but they are not legally identical.

Dozens of lawsuits involving AI developers, authors, publishers and other copyright owners remain pending. Until more appellate courts weigh in or Congress adopts legislation directly addressing generative AI, individual court decisions will continue to shape how companies obtain training data and how copyright holders challenge its use.

For now, the emerging picture is more nuanced than a simple yes-or-no answer: some uses of AI training may qualify as fair use. At the same time, piracy, direct market substitution, and other circumstances can create separate copyright liability.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Angry Angry 0
Sad Sad 0
Wow Wow 0
Shivangi Yadav Shivangi Yadav reports on startups, technology policy, and other significant technology-focused developments in India for TechAmerica.Ai. She previously worked as a research intern at ORF.