Which sites block AI crawlers
We read robots.txt on 200 of the most visited sites on the web and recorded which AI crawlers each one disallows. Every page publishes the rule responsible and the date it was taken, so any of it can be checked in a few seconds.
The distinction that matters throughout: blocking a training crawler keeps a site out of the models. Blocking a search crawler affects whether an assistant can cite it. 40 of these sites do the second; 4 do only the first.
Blocks a crawler that affects citation (40)
These disallow at least one agent that builds the index assistants quote from, or that fetches a page live to answer a question.
- facebook.com5
- instagram.com6
- twitter.com6
- amazon.com6
- whatsapp.net1
- netflix.com1
- pinterest.com6
- x.com6
- tiktok.com6
- whatsapp.com1
- yahoo.com2
- msn.com6
- chatgpt.com1
- baidu.com6
- reddit.com6
- amazonvideo.com4
- nytimes.com6
- netflix.net1
- flickr.com1
- cnn.com6
- theguardian.com3
- ebay.com1
- forbes.com1
- t-online.de1
- bbc.com4
- t.co6
- bbc.co.uk4
- weather.com1
- imdb.com6
- meraki.com6
- launchpad.net4
- reuters.com4
- nature.com2
- weibo.com6
- amazon.co.uk6
- wsj.com4
- trustpilot.com1
- washingtonpost.com1
- bloomberg.com5
- amazon.de6
Blocks training crawlers only (4)
Out of the models, still quotable. Reporting these as invisible to AI, as most checkers do, is wrong.
Every domain checked
Measured 2026-08-08. Sites whose robots.txt could not be established are excluded rather than recorded as open, because a rule we were never shown looks exactly like no rule at all.
Check a domain that is not here
The same check runs on any site, free and without an account, and reports training and search crawlers separately.