Anthropic's $1.5B Settlement Over Pirated Books: A Landmark AI Copyright Case
A federal judge approved a $1.5 billion settlement where Anthropic will pay authors ~$3,000 per book for using pirated copies to train Claude. The ruling upholds fair use for training but punishes the source of the data.

A federal judge has approved a $1.5 billion copyright settlement between Anthropic and a class of authors, resolving claims that the company used pirated copies of books to train its Claude chatbot. The deal, approved by District Judge Araceli Martínez-Olguín, provides roughly $3,000 per book to affected authors and publishers.
About 91% of the more than 482,000 books covered by the settlement have been claimed, making this what plaintiff attorney Justin Nelson called “the largest known copyright recovery in history.” The case was originally filed by bestselling thriller novelist Andrea Bartz and two other authors in 2024.
The Mixed Ruling
The settlement follows a mixed ruling last summer from now-retired Judge William Alsup. He found that training AI chatbots on copyrighted books is not illegal under fair use, but that Anthropic wrongfully acquired millions of books through pirate websites. That distinction matters: the settlement addresses the acquisition method, not the training itself.
Anthropic’s deputy general counsel, Aparna Sridhar, emphasized that the ruling “shows that training AI on books is fair use under copyright law.” The company is paying to resolve the piracy sourcing issue while preserving the core legal precedent that training on copyrighted data is permissible.
What This Means for AI Training
This is the first major settlement in dozens of pending AI copyright lawsuits. It creates a template: pay for the data you scrape, but don't concede that training on copyrighted material is illegal. For engineers building LLMs, the takeaway is clear—you still need to vet your data sources, but the legal path for training on public copyrighted works remains open.
The settlement does not set a per-book royalty rate for future training data licenses. That question will be fought in other cases, including those against OpenAI and Meta.
The court found that training AI on books is fair use under copyright law, but acquiring those books through pirate websites was not. Pay for your data sources, not for your training methodology.
Source: AP News
Discussion
0 Comments
Be the first to start the discussion.