Anthropic2 mins read

Anthropic’s $1.5B Book Piracy Settlement: Why AI Labs Still Won on Fair Use

Anthropic must pay authors $1.5 billion over pirated book downloads, but the settlement does not overturn a key fair-use win for AI training on legally obtained books.

Illustration for Anthropic book authors copyright settlement coverage
Image credits:The Decoder

The Settlement Is Huge — But Its Target Is Narrow

Anthropic book lawsuit illustration
Image credits:The Decoder

Anthropic has to pay book authors $1.5 billion in what The Decoder describes as the largest copyright settlement in class action history. The approved settlement centers on Anthropic downloading books from the piracy databases LibGen and PiLiMi between 2021 and 2022.

The article says roughly 482,460 works were listed, 91.3 percent were claimed, and authors would receive about $3,000 each. Anthropic must also destroy the pirated files.

Why AI Labs Still Got the Key Legal Signal They Wanted

The payout covers piracy, not AI training itself. That distinction matters because Judge Alsup had previously ruled that training AI on legally obtained books is “transformative” and falls under fair use.

For AI labs, the practical takeaway is that the court separated illegal acquisition from the act of training on legally obtained material. The Decoder frames that as a major legal win despite the record settlement amount.

What Authors Can Still Challenge

The settlement does not close every path for authors. According to the article, authors retain claims over AI outputs that reproduce original works and over Anthropic’s future conduct.

That keeps attention on how AI systems generate text after training, not just what data was used to build them. It also means the settlement is not a blanket resolution of copyright concerns around generative AI.

The Open Question: Mass Web Scraping

The Decoder notes that whether mass scraping of internet content without authors’ consent counts as legal acquisition remains unresolved. That question is central because web content is described as a main source of training data for AI labs.

The immediate lesson is clear: pirated datasets carry major legal risk, while fair use arguments for training may be stronger when materials are legally obtained. The broader AI copyright debate is likely to continue around acquisition, consent, outputs, and future conduct.

Discover More