What is copyright and AI training?
This is the unresolved legal question of whether training a model on copyrighted text or images, without a license, counts as infringement, and it is currently being fought out in courts rather than settled by any clear rule.
Publishers often assume that blocking a crawler protects their copyright, and that allowing one means consenting to infringement. Neither follows automatically. Blocking access is a technical decision about who can fetch a page; whether the resulting use of any copyrighted material is lawful is a separate legal question that depends on the court, the specific use, and doctrines like fair use in the US, a principle that can excuse copying without permission when a judge decides the new use is different enough in purpose or effect, decided case by case rather than by fixed rule.
The practical result is that a robots.txt entry is not legal advice and not a waiver either way. A site that blocks every AI crawler has not necessarily protected itself from a model trained on a copy obtained elsewhere, and a site that allows every crawler has not necessarily granted any rights beyond what the crawler's stated purpose covers.
Related
- The text and data mining exceptionThe text and data mining exception is a copyright law provision, most developed in the EU and UK, that allows automated analysis of copyrighted works without separate permission, for specific purposes and subject to rights holders opting out.
- Content licensingContent licensing is a paid agreement that gives an AI company the right to use a publisher's material, typically for training a model, in exchange for money or other consideration.
- robots.txtrobots.txt is a file at a site's root that tells crawlers which parts of the site they may fetch. It is a request, not an enforcement mechanism.
Want to know where you actually stand on this? Run a free visibility check or try the free tools.