A guide to the legal questions at the center of lawsuits over AI models trained on copyrighted material

Generative AI models are trained by processing enormous volumes of text, images, audio, and other data scraped from the internet and other sources, a significant portion of which is protected by copyright. The central legal question is whether using copyrighted material to train an AI model, without a specific license from each rights holder, constitutes copyright infringement, or whether it falls under existing exceptions designed to permit certain uses of copyrighted works without permission.
This is not a single settled question but a live legal dispute being argued in multiple courts and jurisdictions simultaneously, with different legal systems applying different standards.
In the United States, much of the legal debate centers on the doctrine of fair use, which allows certain uses of copyrighted material without permission when factors such as the purpose of the use, the nature of the copyrighted work, the amount used, and the effect on the market for the original work weigh in favor of allowing it. AI companies have generally argued that training a model on copyrighted text is a transformative use, since the resulting model does not reproduce the original works but learns statistical patterns from them. Rights holders and creators have argued the opposite: that large-scale copying for commercial AI training harms the market for their work and does not qualify as fair use, particularly when a model can be prompted to generate outputs closely resembling protected works.
Copyright frameworks vary significantly outside the United States. Some jurisdictions have introduced specific exceptions for text and data mining that can apply to AI training under certain conditions, often requiring that the underlying content was lawfully accessed in the first place. Other regions have moved toward requiring more explicit licensing arrangements or transparency obligations, requiring AI developers to disclose more about the copyrighted material used in training. This has resulted in a patchwork of rules, meaning what is legally permissible in one country may not be in another.
| Theme | What's Being Argued |
|---|---|
| Training as infringement | Whether the act of copying data for training itself requires a license |
| Output similarity | Whether specific AI-generated outputs improperly reproduce protected works |
| Fair use / transformative use | Whether training qualifies for existing copyright exceptions |
| Licensing markets | Whether a market for licensing AI training data should be recognized |
Expect continued litigation and, in some jurisdictions, new legislation specifically addressing AI training data, as lawmakers weigh competing interests between fostering AI innovation and protecting creators' rights. Licensing agreements between AI companies and content owners are likely to keep expanding as a practical, market-based alternative or supplement to waiting for courts to fully resolve the underlying legal questions.
Is training AI on copyrighted material definitely illegal?
There is no single, settled answer. It depends on the jurisdiction, the specific facts of how the data was obtained and used, and ongoing court rulings that continue to shape the legal landscape.
Have any court cases been fully resolved?
Some cases have reached preliminary rulings or settlements, but the overall legal landscape remains unsettled, with many significant cases still working through the courts as of late 2026.
Are AI companies now licensing content instead of just using it freely?
Some major AI companies have signed licensing deals with publishers and content platforms, though this coexists with ongoing litigation over past and current training practices.
Does this affect only large AI companies?
The core legal questions apply broadly to anyone training AI models on copyrighted data, though large, well-resourced AI companies are the primary current defendants in most major lawsuits.
The legal status of AI training data remains one of the most consequential unresolved questions in technology law, with outcomes likely to shape how AI models are built and licensed for years to come. Following jurisdiction-specific rulings and legislation remains essential for both AI developers and content creators trying to understand their rights and obligations.