Artificial Intelligence and Pirated Content

Artificial intelligence, notably in its form based on large language models (LLMs), relies on ingesting a very large volume of text data. However, as revelations have emerged, it appears that some of these models were trained on copyrighted content without prior authorization. The use of pirated content, notably books from databases such as LibGen, raises questions about respect for the law, the sustainability of technological development and the place of authors in this new ecosystem.
This first part examines the recently revealed facts, notably around the Meta / LibGen case, retracing the motivations, the technical processes and the first lawsuits.
The Meta Scandal and the Use of LibGen
In March 2025, court documents shed light on Meta's use of pirated content to train its LLaMA model (Large Language Model Meta AI). The company allegedly used data from Library Genesis (LibGen), a platform well known in academic circles for offering free access to digital books, the vast majority of them illegally downloaded.
AgencePDN gets pirated content removed: see our solutions by sector.
The volume of data involved is considerable: about 183,000 books were allegedly used, representing nearly 32 TB of information. Meta's internal exchanges, revealed by The Atlantic and other specialized media, show that the company was fully aware of the illicit nature of these sources. Members of the legal team reportedly warned of the legal risks, but management allegedly authorized the use of this content so as “not to fall behind” competitors such as OpenAI or Anthropic.

The Strategic Interest of Use by AI
Resorting to pirated content, despite its legal danger, offers considerable advantages. Three main factors can explain this drift.
- First, the availability and density of data: published works (novels, textbooks, essays, etc.) are rich, structured, carefully written texts. They are therefore very effective for training a model to produce coherent, sustained, relevant language. Moreover, their thematic diversity is a decisive asset.
- Second, the cost of lawful access is particularly high. Publishing houses, content aggregators and authors demand significant payment, proportional to the value of their output. Yet developing a language model requires corpora made up of tens or even hundreds of millions of documents, which makes individual licensing economically prohibitive.
- Finally, speed of access plays a central role. Using LibGen or other similar databases allows immediate download of massive content, without contractual constraint or negotiation. This makes it possible to train models quickly and cheaply.
An Exploited Legal Vacuum
The tech giants often rely on American fair use to justify the use of protected content. This concept allows, under certain conditions, use without authorization for research, commentary or parody. But this basis is contested in several ongoing cases, notably because:
- fair use is an exception clause, assessed case by case by the courts;
- its application in the field of AI is still largely undetermined;
- it applies only on American territory and is incompatible with European legislation.
In France, three representative organizations – the Syndicat national de l’édition (SNE), the Société des gens de lettres (SGDL) and the Syndicat national des auteurs et des compositeurs (SNAC) – have taken the matter to court. They denounce the unauthorized appropriation of protected works, often available in bookstores or in official digital libraries.
The Creators' Point of View
Many authors, sometimes with no international reputation, have discovered that their books were in the training corpora used by some AIs. Citizen initiatives have made it possible to cross-reference the metadata of AI models with that of pirate databases to identify the works concerned.
The feeling of dispossession is real. Not only were authors not consulted, but they see that their creations are being used to generate texts automatically, sometimes in their own style, with no payment at all. Some speak of a new type of intellectual property theft, in which works are no longer simply copied or distributed illegally, but absorbed to feed tools that could compete with their very profession.
Toward a Deep Transformation?
The Meta, OpenAI and Stability AI cases are probably only the first in a series. In the United States, class actions have been filed by writers, visual artists and publishers. In Europe, several national courts are taking up the subject, often in the absence of settled case law.
The central question remains: can an AI be trained on a work without authorization if it does not explicitly reproduce its content? The debate pits supporters of a functional interpretation (which looks at the final result) against those of a strict property-based approach (under which any use must be paid for).
The use of pirated content in training artificial intelligence models reveals the major tensions between technological innovation, respect for copyright and the economic viability of creative work. The case of Meta and the use of LibGen crystallizes these issues: it shows that the line between technical exploration and circumvention of the law has now been crossed by some digital players, under the guise of efficiency and competitiveness.
Join us in mid-May for our second part, in which we will address the ethical consequences of these practices, the regulatory avenues being considered internationally, and the prospects for moving toward a fairer model of AI training. In the meantime, if you have a film, a series, software or an ebook to protect, don't hesitate to call on our services by contacting one of our account managers; PDN has been a pioneer in cybersecurity and anti-piracy for more than ten years, and we surely have a solution to help you. Happy reading, and see you soon!
Share this article


