PDN

AI Training and Pirated Content: Issues and Outlook

5 min readPart 2 of 2

Illustration for the article: AI Training and Pirated Content: Issues and Outlook

In the first part of this article, we noted that several major artificial intelligence models were trained, at least partly, on copyrighted content obtained illegally. This practice, though largely concealed until recently, is now becoming a subject of litigation on an international scale.

This second part looks at the ethical implications of the phenomenon, at the – still timid – responses from legislators, and at the possible ways to govern AI training in a fairer, more transparent way that is more respectful of creators.

A Major Ethical Break

The unauthorized use of protected works in developing advanced technologies represents a deep ethical flaw. Unlike marginal uses (such as quotation or parody), AI systems use works in their entirety, often for commercial purposes, and on an industrial scale. This raises several problems:

AgencePDN gets pirated content removed: see our solutions by sector.

  • The absence of consent: authors, publishers, researchers and artists were neither consulted nor informed of the use of their content.

  • The separation of exploitation from remuneration: while AI generates profits for the companies that commercialize it, rights holders receive no compensation.

  • The blurring of traceability: once ingested, works become invisible in the training corpus; it is then almost impossible to know whether a generated text is influenced by a specific author.

Beyond the legal question, then, it is a philosophy of intellectual property that is at stake. Copyright rests on the idea that creation involves an intellectual, emotional and often economic investment that deserves recognition and protection. By absorbing these human productions without permission, AI shakes that conception in favour of an extractive paradigm inherited from the platforms' business model.

The Position of the Tech Giants

Faced with criticism, some technology companies are adopting a defence strategy built on several lines:

  • Technological progress as justification: they claim that training on as much data as possible is a sine qua non for developing models that are useful to society (translation, medicine, accessibility, etc.).

  • Fair use as a legal basis (under American law), even though it is contested in many jurisdictions.

  • The dilution of responsibility: some players claim not to have known exactly which sources were used, particularly when the data comes from subcontractors or intermediary public databases.

  • The emergence of open source models as a lever for democratization: companies such as Meta promote free access to their models to justify a certain tolerance of training practices.

However, this line of defence appears increasingly fragile. It deliberately ignores the fundamental principles of copyright and rests on a utilitarian logic, in which the potential good generated for the greatest number would justify the harm inflicted on individual creators.

Legal Responses Under Way

On the litigation front, proceedings are multiplying in several countries. In the United States, class actions against OpenAI, Meta or Stability AI are trying to establish case law that protects authors. In Europe, national courts are beginning to take a position. We also note:

  • A complaint under way in France against Meta, which we discussed in the first part of our article
  • Challenges at the European Parliament, where the directive on artificial intelligence could include provisions on the origin of data.

Some legislators argue for introducing a compulsory licensing right for AI training, along the lines of what exists for reprography or radio. This would legalize existing practices while ensuring a form of redistribution to rights holders.

Toward Transparency Obligations?

Another possible lever concerns the transparency of training corpora. Today, most models are “black boxes” with respect to their source data. Yet without clear information on what was used, it is difficult for creators to defend their rights.

Proposals are emerging to require AI developers to:

  • Publish a complete or representative list of the works used during training.
  • Provide an interface allowing rights holders to check whether their content is present.
  • Allow a clear, simple and accessible withdrawal or objection (“opt-out”) mechanism.

This transparency would not solve every problem (notably those tied to past use of data), but it would be a first step toward a fairer and more responsible model.

Rethinking the Value Chain

The debate over AI and pirated content is not simply a matter of law or morality. It deeply questions the value chain in the digital economy. If works can be absorbed by machines without compensation, what is the residual value of human creation in the 21st-century economy?

Two pitfalls must be avoided here:

  • Technophobia, which would reject AI on principle as a systematic threat.
  • Technological solutionism, which brushes aside ethical problems under the pretext of innovation.

Between the two, a path is possible: it goes through recognizing the role of creators, putting fair remuneration frameworks in place, and integrating cultural rights into technology governance.

Several avenues are currently being studied internationally:

  • Require developers to publish the exact sources of their training corpora;

  • Set up a mandatory collective licensing mechanism, along the lines of what exists for music or television;

  • Create a clear right to object (opt-out) for authors who refuse to have their works used;
  • Introduce an automatic royalty, redistributed to rights holders through collective management organizations.

Training artificial intelligence on pirated content is a practice that raises complex issues, at once legal, economic, political and ethical. While some companies may have believed themselves above the law in a climate of technological euphoria, it is now clear that a rebalancing is needed.

The tools exist: transparency obligations, licensing mechanisms, withdrawal rights, legislative frameworks. What is still needed is a strong political will, at both the national and international level, to put them into effect. The future of creation, justice and responsible innovation depends on it.

Join us in June for our series on cryptocurrencies. In the meantime, if you have a film, a series, software or an ebook to protect, don't hesitate to call on our services by contacting one of our account managers; PDN has been a pioneer in cybersecurity and anti-piracy for more than ten years, and we surely have a solution to help you. Happy reading, and see you soon!

Share this article

Is your content pirated? We can get it removed.