Skip to content
Glossary

What is content licensing?

Content licensing is a paid agreement that gives an AI company the right to use a publisher's material, typically for training a model, in exchange for money or other consideration.

These deals sit apart from the everyday mechanics of crawling. A crawler reading a public page and a licensing agreement are two different relationships: one is unpaid access under terms set by robots.txt, the other is a negotiated contract, often covering back catalogs the crawler could never reach on its own, like paywalled archives.

The detail people miss is that a licensing deal usually governs training data, not what happens when a user asks a live question. A publisher can be paid to have its archive folded into a model's training set and still be effectively invisible when that same model does a live, retrieval-based answer, if its current site blocks the separate crawler used for that purpose. The two are governed by different mechanisms and often by different bots entirely.

Related

Want to know where you actually stand on this? Run a free visibility check or try the free tools.