What is the difference between a training crawler and a search crawler?
A training crawler collects data models learn from. A search crawler builds the index an assistant cites when answering. Blocking one has nothing to do with the other.
This is the distinction most AI-visibility tools collapse, and collapsing it produces a specific false claim: that a blocked crawler makes you invisible to AI. Across our index of 33,670 domains with a readable robots.txt, 5,497 block at least one AI crawler, but only 2,576 block one that affects citation. The other 2,921, or 53%, remain entirely quotable.
The decisions differ in kind. Refusing training is a position about your content being used to build a competitor to you, and costs you nothing in visibility. Refusing search crawlers is a decision to be absent from the answers your buyers read.
Practically: before acting on any tool reporting that AI crawlers are blocked on your site, look at which agents. More than half the time the finding does not mean what the tool says it means.
Related
- OAI-SearchBotOAI-SearchBot is the crawler OpenAI uses to build the index ChatGPT cites from. It is not the training crawler, and blocking it is a different decision from blocking GPTBot.
- GPTBotGPTBot is OpenAI's web crawler. It reads pages to train and ground OpenAI's models, including ChatGPT.
- AI crawlerAn AI crawler is a bot operated by an AI company to fetch web pages for training, indexing, or answering a live question.
- CitationA citation is a source an AI assistant links or refers to in support of an answer it has given.
Want to know where you actually stand on this? Run a free visibility check or try the free tools.