Where AI training data meets copyright that already had owners

Level High Timing Pre deployment

What this risk is

The use of copyrighted creative works — books, articles, code, images, music, video — to train AI systems without the consent of the rights holders, without compensation, and in ways that enable the AI to produce outputs that compete with and economically substitute for the human-created works that trained it.

This is the most legally active AI risk area as of 2025, with major litigation pending in the US, UK, and EU.

How it occurs · Mechanisms

AI developer position: Training on publicly available data constitutes fair use (US) or text and data mining exception (EU); the AI learns “from” works, it doesn’t reproduce them.

Rights holder position: Training is reproduction requiring licensing; the AI produces outputs that substitute for the original works; this is economic harm at scale without compensation.

The unresolved question: Courts in the US, UK, and EU are still determining where the legal line falls. The outcome will fundamentally shape how AI models can be trained.

Mitigations · Governance

  • Licensing programs — License content from rights holders; pay for training data (as some AI companies now do)
  • Opt-out mechanisms — Honor robots.txt and other creator-specified opt-out signals
  • Compensation frameworks — Revenue sharing models for creators whose work trains AI
  • Provenance tracking — Track which works were used in training for licensing and compliance purposes
  • Content authenticity — Label AI-generated content as such to preserve market differentiation for human creators

Risk you cannot name is risk you cannot manage.

Map your AI portfolio against this taxonomy with Zertia.