What is the text and data mining exception?
The text and data mining exception is a copyright law provision, most developed in the EU and UK, that allows automated analysis of copyrighted works without separate permission, for specific purposes and subject to rights holders opting out.
It exists because copyright law was written for humans reading and copying works, not for software statistically analyzing millions of them at once. Rather than requiring every AI company to negotiate a license before extracting patterns from text, the exception carves out a lane for that kind of automated processing, provided the use fits within what the law defines, which varies by country and often distinguishes research use from commercial use.
It is not the single, uniform shield it is sometimes described as. The EU version lets a rights holder reserve their rights and opt out of the exception for commercial mining specifically; the UK's is narrower. The US has no equivalent statute at all and instead relies on fair use case law, a separate legal doctrine covered under copyright and AI training. A company citing this exception to justify training in one country may have no comparable defense in another.
Related
- Copyright and AI trainingThis is the unresolved legal question of whether training a model on copyrighted text or images, without a license, counts as infringement, and it is currently being fought out in courts rather than settled by any clear rule.
- Opt-outOpting out is telling an AI crawler, usually through robots.txt, not to access your site, and it works only if that crawler chooses to honor the request.
- Content licensingContent licensing is a paid agreement that gives an AI company the right to use a publisher's material, typically for training a model, in exchange for money or other consideration.
Want to know where you actually stand on this? Run a free visibility check or try the free tools.