AI Copyright3 mins read

Sony and Warner Sue Anthropic Over Alleged Copyrighted Music Training Data

Sony Music, Warner Music, and other publishers allege Anthropic used tens of thousands of copyrighted musical compositions to train Claude without permission, escalating AI’s copyright fight over training data.

The Lawsuit Targets How Claude Was Trained

Sony Music, Warner Music, and other publishers are suing Anthropic in federal court in Northern California over allegations that the company used tens of thousands of copyrighted musical compositions to train Claude without permission. The complaint focuses largely on song lyrics, along with sheet music and derivative works. Anthropic CEO Dario Amodei and co-founder Benjamin Mann are named as individual defendants over their alleged role in directing and overseeing torrent-based acquisition of copyrighted files.

Why the Data Source Matters

The case centers not only on whether copyrighted works were used in AI training, but how Anthropic allegedly acquired them. The complaint alleges torrenting, scraping lyrics from licensed platforms such as MusixMatch and LyricFind without publisher consent, use of datasets that allegedly contain unauthorized content, and scanning and destroying used songbooks and sheet music collections. That framing matters because the plaintiffs treat unlawful acquisition as a standalone infringement issue.

Potential Damages and Prior Copyright Pressure

The publishers are seeking up to $150,000 per infringed work and up to $25,000 per violation for alleged unlawful removal of copyright management information. The Decoder notes the case follows Anthropic’s September 2025 agreement to pay $1.5 billion to authors and publishers over pirated books used during AI training. The new complaint argues that Anthropic faces a similar vulnerability if the alleged acquisition methods are proven.

Synthetic Data Claims Could Broaden the Fight

The plaintiffs also challenge Anthropic’s handling of allegedly torrented material through synthetic data and reinforcement feedback. They allege that at least one commercial Claude model was trained on synthetic data generated by a non-commercial model that had learned from LibGen and PiLiMi texts. The article notes that the full scope of these claims is expected to emerge during discovery.

Discover More