
The legal complexities of training artificial intelligence on copyrighted works
The intersection of artificial intelligence training and copyright law remains legally ambiguous. Recent court rulings suggest that ingesting published works to learn patterns may be permissible, while acquiring those works through illicit means or building direct market competitors carries heavy penalties.
Published by Jin · 2 min read · 24 AUG 2026
As artificial intelligence systems ingest vast quantities of published books, articles, and academic papers to improve their capabilities, the question of copyright legality remains a subject of intense debate. While many authors argue that using their creative output without consent threatens their livelihoods, legal realities involve nuanced distinctions between reading, copying, and direct market competition.
The distinction between consumption and infringement
In a notable legal decision, a federal judge ordered Anthropic to pay a $1.5 billion settlement to a group of writers. However, the court ruled that the actual training of large language models on copyrighted text was lawful, comparing the process to a writer studying literature. The financial penalty was instead assessed because the company acquired books through illegal online shadow libraries.
Legal experts note that copyright law focuses primarily on copying and market impact rather than the act of consuming or reading a work. Because current copyright guidelines date back to 1976, judges are tasked with interpreting half-century-old statutes to address modern technological capabilities.
Fair use and competitive intent
Central to these legal battles is the doctrine of fair use, which permits the unlicensed use of copyrighted material for purposes such as criticism, education, or transformative creation. Courts generally evaluate whether an AI model's training data serves a transformative purpose or directly competes with the original author's market.
In a separate case involving Thomson Reuters and research firm Ross Intelligence, a judge ruled that training an AI model on proprietary legal content to build a directly competing platform does not constitute fair use. When a system is designed to rival the original publisher, courts tend to rule against the AI developer.
Looking ahead
As pending litigation continues to wind through the courts, the legal framework governing artificial intelligence remains unsettled. Questions surrounding both the legality of training datasets and the copyright status of purely AI-generated outputs continue to force a reevaluation of intellectual property law for the digital age.
Source — Original announcement ↗
Worth a read?
Comments · 0