Back to articles
Policy & Regulation

AI Training on Copyrighted Books: Why the Legal Answer Is Still Unclear

3 min read

The copyright dispute around generative AI is often described as a question of whether a company may read books without permission. That framing is understandable, but it leaves out several legal distinctions. Developers train models on books, articles, academic papers, and other online material, while authors argue that their work has been commercially exploited without consent. AI companies respond that training extracts patterns and capabilities rather than delivering the original text to users.

US copyright law does not contain a bespoke rule for ingesting hundreds of millions of works into a modern language model. Courts therefore have to apply older concepts, especially fair use, to a technology that operates at a very different scale from earlier forms of copying.

Key points

  • The legality of training may differ from the legality of obtaining the data. In the Anthropic litigation, Judge William Alsup concluded that training on books could be lawful. He viewed the model’s process as more comparable to reading and learning from literature than to reproducing books for readers. However, Anthropic had obtained some works from illegal online shadow libraries. That conduct led to a reported $1.5 billion settlement. The result should not be read as a blanket approval of every AI training practice: the court’s criticism focused on the unlawful source of the books rather than declaring training itself illegal.
  • Fair use is the central test. Courts generally examine the purpose and character of a use, the amount taken, and the effect on the market for the original work. A use that changes the purpose of the material may receive stronger protection, while a use that substitutes for the original market may be viewed more skeptically.
  • Competition matters. In Thomson Reuters’ case against Ross Intelligence, Ross used legal content to build an AI research platform that competed with Thomson Reuters. Judge Stephanos Bibas found that the use did not have a sufficiently different purpose or character to qualify as transformative. Authors may argue that chatbots also compete with them by generating synthetic books, but that theory has not yet produced the same clear courtroom result.
  • Training rights and output rights are separate questions. In Thaler v. Perlmutter, a court held that a work created entirely by AI is not eligible for copyright protection. That decision raises a different set of questions about human contribution, AI assistance, and how anyone could reliably determine the proportion of a work produced by a machine.

Why it matters

The early cases are not creating a single industry-wide rule. Instead, they are identifying the facts that may decide future disputes: whether the training corpus was lawfully acquired, whether the model’s purpose is transformative, whether the resulting product competes with the copyright owner, and whether outputs reproduce protected expression.

For publishers and authors, the stakes include more than retrospective compensation. They are also concerned about whether synthetic content will weaken the market that supports professional writing. For AI companies, the practical lesson is that provenance, licensing strategies, filtering, and safeguards against memorized output may matter as much as the abstract argument that models learn in a manner similar to human readers.

The legal landscape remains unsettled. Different courts may weigh market harm and transformation differently, and decisions at early stages of litigation can still be revised. Until more cases reach final appellate resolution—or lawmakers update the statute—copyright compliance for AI training will remain a matter of managing legal risk rather than following a definitive checklist.

Source: TechCrunch AI

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Flock Safety Faces Backlash as CEO Calls for a Privacy-Safety ‘Compromise’
Policy & Regulation
cctest.ai

Flock Safety Faces Backlash as CEO Calls for a Privacy-Safety ‘Compromise’

Flock Safety is under mounting scrutiny over the possible misuse of its cameras, drones, and automated license plate recognition tools. CEO Garrett Langley says the country needs to balance public safety with privacy, while critics demand stronger controls and accountability.

Read more