The Medicare Advantage startup’s funding talks underscore rising investor focus on AI and healthcare.
TechCrunch examines why training AI models on copyrighted books remains legally complicated, with courts weighing fair use, market impact, piracy, and whether AI-generated works can be copyrighted.


AI models powering tools such as ChatGPT, Gemini, Claude, and other chatbots are trained on vast collections that can include books, articles, academic papers, and other internet-accessible material. TechCrunch frames the central legal tension clearly: many published authors may have contributed to AI systems without knowledge or consent, while those systems could affect their livelihoods.
The key issue is whether training on copyrighted books is treated more like a person reading and learning from a work, or like unlawful copying. The answer remains contested, and the article emphasizes that the reality is not simple.
TechCrunch highlights a major Anthropic ruling in which Judge William Alsup ordered a $1.5 billion copyright settlement connected to writers whose works were used to train AI models. But the article notes an important distinction: the judge ruled that Anthropic’s AI training was lawful, while penalizing the company for pirating books from illegal online shadow libraries.
That distinction matters for AI companies and authors alike. It suggests courts may separate the legality of model training from the legality of how training materials were obtained.
The article explains that many disputes turn on fair use, especially whether use of copyrighted work is sufficiently “transformative.” Courts may consider the purpose and nature of the use, how much of the work was used, and the effect on the market.
TechCrunch also points to a Thomson Reuters case against Ross Intelligence, where a judge found the use was not transformative because the AI-based legal platform would directly compete with Thomson Reuters. That makes market competition a practical factor to watch in future AI copyright cases.
Training data is only one part of the copyright puzzle. TechCrunch also discusses the separate question of whether AI-generated material can be copyrighted, citing Thaler v. Perlmutter, where the court ruled that a work that is 100% AI-generated is not copyrightable.
That raises follow-up questions for writers, publishers, and platforms: how much AI assistance changes copyright status, how AI involvement can be proven, and where courts draw the line between tool use and authorship.
The near-term takeaway is uncertainty. TechCrunch reports that many AI companies remain tied up in pending litigation, meaning there is no definitive answer yet on the broader legality of training AI models on copyrighted books.
Early court decisions are already shaping company behavior, author strategies, and policy debates. For now, the most important questions are how training data was obtained, whether the use competes with the original market, and how courts interpret older copyright law in a rapidly changing AI industry.
The Medicare Advantage startup’s funding talks underscore rising investor focus on AI and healthcare.

The orbital data center startup is raising capital as launch capacity becomes a strategic bottleneck.

Pew Research found signs of AI authorship in 35% of post-ChatGPT web pages studied.
Orion Hindawi returns as Tanium CEO as the cybersecurity company responds to AI-driven software shifts.