How many sites block AI crawlers
Published estimates for this range from 3.5% to 88.9%. They are not really in conflict. Each one measured a different set of websites and then got quoted without it. Here is ours, and here is the denominator it depends on.
Out of the 2916 whose robots.txt we could establish, not out of the 5000 we asked.
What we threw out
2084 of 5000 domains refused us their robots.txt or failed to answer. They are excluded rather than counted as unblocked.
A rule we were never shown looks exactly like no rule at all. Counting those in the clear would push the headline down by however many strict firewalls the sample happened to contain, and sites that run one are not a random subset.
By crawler
Every count is out of the same 2916 domains. Blocking is a per-agent decision and reporting it per agent is the only way to see the finding below.
| Crawler | Feeds | Domains | Share |
|---|---|---|---|
| GPTBot | ChatGPT | 502 | 17.2% |
| ClaudeBot | Claude | 457 | 15.7% |
| Google-Extended | Gemini | 424 | 14.5% |
| PerplexityBot | Perplexity | 218 | 7.5% |
| ChatGPT-User | ChatGPT | 214 | 7.3% |
| OAI-SearchBot | ChatGPT | 174 | 6.0% |
| Perplexity-User | Perplexity | 156 | 5.3% |
| Claude-SearchBot | Claude | 155 | 5.3% |
| Claude-User | Claude | 150 | 5.1% |
| Googlebot | Gemini / AI Overviews | 66 | 2.3% |
The interesting part is not the headline
GPTBot is disallowed on 502 domains. OAI-SearchBot on 174. Both are OpenAI. GPTBot trains and grounds the models; OAI-SearchBot builds the index ChatGPT cites from when it answers a question.
So roughly 65% of the sites refusing OpenAI are refusing to be trained on while remaining entirely citable. That is a deliberate distinction and, as far as we can tell, a correct one.
It also explains a result that otherwise looks like a paradox. BuzzStream found that 88.2% of news sites blocking GPTBot still get cited by AI. If most blockers never blocked the search crawler, there is nothing to explain.
Which means the sentence you will read everywhere, including in cold emails, is wrong for most of the sites it gets said to. “GPTBot is blocked, so ChatGPT cannot see you” is only true if OAI-SearchBot is blocked too. Check which one before you act.
Who this is, by name
We ran the same check over 118 well-known brands. 101 had an establishable robots.txt. These 27 disallow at least one search crawler, which is the group that has actually opted out of being cited.
| Domain | Search crawlers disallowed |
|---|---|
| cnn.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User |
| economist.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User |
| huffpost.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User |
| nbcnews.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User |
| nytimes.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User |
| usatoday.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User |
| yelp.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User |
| bbc.co.uk | OAI-SearchBot, ChatGPT-User, Perplexity-User |
| bloomberg.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot |
| quora.com | OAI-SearchBot, ChatGPT-User, Claude-SearchBot |
| theverge.com | ChatGPT-User, Claude-SearchBot, Perplexity-User |
| vox.com | ChatGPT-User, Claude-SearchBot, Perplexity-User |
| cnbc.com | OAI-SearchBot, ChatGPT-User |
| instacart.com | Claude-SearchBot, Perplexity-User |
| newyorker.com | Claude-SearchBot, Perplexity-User |
| reuters.com | Claude-SearchBot, Perplexity-User |
| theatlantic.com | Claude-SearchBot, Perplexity-User |
| wired.com | Claude-SearchBot, Perplexity-User |
| wsj.com | Claude-SearchBot, Perplexity-User |
| buzzfeed.com | Perplexity-User |
| disneyplus.com | Perplexity-User |
| espn.com | ChatGPT-User |
| ft.com | Perplexity-User |
| hulu.com | Perplexity-User |
| netflix.com | Perplexity-User |
| newsweek.com | ChatGPT-User |
| theguardian.com | Claude-SearchBot |
Opted out of training, still fully citable
These 10 disallow a training crawler and no search crawler. An engine can still fetch and quote them today. Almost every article written about AI crawler blocking describes this group and the one above as the same decision, and they are opposites.
| businessinsider.com | ClaudeBot |
| canva.com | GPTBot, ClaudeBot |
| ebay.com | GPTBot, ClaudeBot, PerplexityBot |
| figma.com | GPTBot, Google-Extended |
| forbes.com | GPTBot, ClaudeBot, PerplexityBot |
| garmin.com | PerplexityBot |
| medium.com | GPTBot, ClaudeBot |
| strava.com | GPTBot, ClaudeBot, Google-Extended |
| tripadvisor.com | GPTBot, ClaudeBot |
| washingtonpost.com | ClaudeBot, PerplexityBot |
Every row is a line in that site's public robots.txt on August 7, 2026. Check any of them yourself at /robots.txt. We are reporting what each file says and nothing beyond it: a disallowed crawler is a rule the site published, not a claim that any engine has stopped naming them.
Method
The panel. Tranco daily list Q2X24 (generated 2026-08-05), ranks 2000-50000, excluding hostnames matching a fixed infrastructure/CDN/ad-tech pattern, then a systematic every-9th sample of the 47,341 eligible rows. Frozen: membership never changes. New domains belong to a new panel with its own manifest.
What counts as readable. A 2xx we could parse, or a 404 or 410. A 404 is an answer: the server replied and there is no such file, so nothing is disallowed there. A 403, a timeout or a 5xx is not an answer, and those domains are excluded.
What we did not do. We did not treat a server refusing our live request as a block. We send a crawler's User-Agent from our own IP, which is what an impersonator looks like, so a bot-protection service refusing us is that service working correctly rather than evidence about the crawler.
Run. Completed August 7, 2026, from a single vantage point, one moment in time. The panel is frozen and will be re-run on the same domains, so the next edition is a change rather than a new sample.
What this does not show
The 41.7% exclusion rate is the biggest weakness here and it is not random. Sites behind strict firewalls are dropped systematically, and if those sites block AI crawlers at a different rate than the ones that let us read, the headline is off by an amount we cannot measure from here.
robots.txt is also a request rather than a wall. It tells you what a site asks crawlers to do, not what any crawler did. And access is not citation: a crawler permitted to read a page may never fetch it, and one that fetches it may never quote it.
We publish this weakness because the alternative is a cleaner number that means less. If you want to check any of it, the crawler checker below runs the same measurement on one domain and shows the responsible robots.txt line.
Run it on one domain, and clear what it finds
The checker is free and needs no account: run the same check on your own domain and it names each agent and the robots.txt line responsible. The harder case is the one this study cannot see from the outside: a file that permits every agent while the CDN in front of it refuses them anyway. That is the case below.
AI assistants have to read your website before they can mention it, and we find whatever is blocking them from reading yours and clear it.
- 10 AI agents tested against robots.txt, CDN and WAF
- The exact rule behind every block, quoted
- A corrected robots.txt, written and applied for you
- All 10 re-fetched after deploy, response codes in writing
Get the next edition
The panel is frozen, so the same domains get re-checked and the next edition reports what moved rather than a new sample. One email when it runs, and nothing else.
Why we have 5,000 domains of this
Standing measures whether AI assistants actually name a brand when someone asks for a recommendation, across five engines, with a confidence interval on every score. Whether the crawlers can reach a site at all is the floor underneath that, which is why we ended up counting it. The same crawler check runs on one domain, free and without an account, and reports the exact robots.txt line responsible.