Skip to content
Research · August 2026

How many sites block AI crawlers

Published estimates for this range from 3.5% to 88.9%. They are not really in conflict. Each one measured a different set of websites and then got quoted without it. Here is ours, and here is the denominator it depends on.

Disallow at least one AI crawler in robots.txt
18.3%
535 of 2916 domains

Out of the 2916 whose robots.txt we could establish, not out of the 5000 we asked.

What we threw out

2084 of 5000 domains refused us their robots.txt or failed to answer. They are excluded rather than counted as unblocked.

A rule we were never shown looks exactly like no rule at all. Counting those in the clear would push the headline down by however many strict firewalls the sample happened to contain, and sites that run one are not a random subset.

By crawler

Every count is out of the same 2916 domains. Blocking is a per-agent decision and reporting it per agent is the only way to see the finding below.

CrawlerFeedsDomainsShare
GPTBotChatGPT50217.2%
ClaudeBotClaude45715.7%
Google-ExtendedGemini42414.5%
PerplexityBotPerplexity2187.5%
ChatGPT-UserChatGPT2147.3%
OAI-SearchBotChatGPT1746.0%
Perplexity-UserPerplexity1565.3%
Claude-SearchBotClaude1555.3%
Claude-UserClaude1505.1%
GooglebotGemini / AI Overviews662.3%

The interesting part is not the headline

GPTBot is disallowed on 502 domains. OAI-SearchBot on 174. Both are OpenAI. GPTBot trains and grounds the models; OAI-SearchBot builds the index ChatGPT cites from when it answers a question.

So roughly 65% of the sites refusing OpenAI are refusing to be trained on while remaining entirely citable. That is a deliberate distinction and, as far as we can tell, a correct one.

It also explains a result that otherwise looks like a paradox. BuzzStream found that 88.2% of news sites blocking GPTBot still get cited by AI. If most blockers never blocked the search crawler, there is nothing to explain.

Which means the sentence you will read everywhere, including in cold emails, is wrong for most of the sites it gets said to. “GPTBot is blocked, so ChatGPT cannot see you” is only true if OAI-SearchBot is blocked too. Check which one before you act.

Who this is, by name

We ran the same check over 118 well-known brands. 101 had an establishable robots.txt. These 27 disallow at least one search crawler, which is the group that has actually opted out of being cited.

DomainSearch crawlers disallowed
cnn.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User
economist.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User
huffpost.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User
nbcnews.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User
nytimes.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User
usatoday.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User
yelp.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot, Perplexity-User
bbc.co.ukOAI-SearchBot, ChatGPT-User, Perplexity-User
bloomberg.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot
quora.comOAI-SearchBot, ChatGPT-User, Claude-SearchBot
theverge.comChatGPT-User, Claude-SearchBot, Perplexity-User
vox.comChatGPT-User, Claude-SearchBot, Perplexity-User
cnbc.comOAI-SearchBot, ChatGPT-User
instacart.comClaude-SearchBot, Perplexity-User
newyorker.comClaude-SearchBot, Perplexity-User
reuters.comClaude-SearchBot, Perplexity-User
theatlantic.comClaude-SearchBot, Perplexity-User
wired.comClaude-SearchBot, Perplexity-User
wsj.comClaude-SearchBot, Perplexity-User
buzzfeed.comPerplexity-User
disneyplus.comPerplexity-User
espn.comChatGPT-User
ft.comPerplexity-User
hulu.comPerplexity-User
netflix.comPerplexity-User
newsweek.comChatGPT-User
theguardian.comClaude-SearchBot

Opted out of training, still fully citable

These 10 disallow a training crawler and no search crawler. An engine can still fetch and quote them today. Almost every article written about AI crawler blocking describes this group and the one above as the same decision, and they are opposites.

businessinsider.comClaudeBot
canva.comGPTBot, ClaudeBot
ebay.comGPTBot, ClaudeBot, PerplexityBot
figma.comGPTBot, Google-Extended
forbes.comGPTBot, ClaudeBot, PerplexityBot
garmin.comPerplexityBot
medium.comGPTBot, ClaudeBot
strava.comGPTBot, ClaudeBot, Google-Extended
tripadvisor.comGPTBot, ClaudeBot
washingtonpost.comClaudeBot, PerplexityBot

Every row is a line in that site's public robots.txt on August 7, 2026. Check any of them yourself at /robots.txt. We are reporting what each file says and nothing beyond it: a disallowed crawler is a rule the site published, not a claim that any engine has stopped naming them.

Method

The panel. Tranco daily list Q2X24 (generated 2026-08-05), ranks 2000-50000, excluding hostnames matching a fixed infrastructure/CDN/ad-tech pattern, then a systematic every-9th sample of the 47,341 eligible rows. Frozen: membership never changes. New domains belong to a new panel with its own manifest.

What counts as readable. A 2xx we could parse, or a 404 or 410. A 404 is an answer: the server replied and there is no such file, so nothing is disallowed there. A 403, a timeout or a 5xx is not an answer, and those domains are excluded.

What we did not do. We did not treat a server refusing our live request as a block. We send a crawler's User-Agent from our own IP, which is what an impersonator looks like, so a bot-protection service refusing us is that service working correctly rather than evidence about the crawler.

Run. Completed August 7, 2026, from a single vantage point, one moment in time. The panel is frozen and will be re-run on the same domains, so the next edition is a change rather than a new sample.

What this does not show

The 41.7% exclusion rate is the biggest weakness here and it is not random. Sites behind strict firewalls are dropped systematically, and if those sites block AI crawlers at a different rate than the ones that let us read, the headline is off by an amount we cannot measure from here.

robots.txt is also a request rather than a wall. It tells you what a site asks crawlers to do, not what any crawler did. And access is not citation: a crawler permitted to read a page may never fetch it, and one that fetches it may never quote it.

We publish this weakness because the alternative is a cleaner number that means less. If you want to check any of it, the crawler checker below runs the same measurement on one domain and shows the responsible robots.txt line.

Run it on one domain, and clear what it finds

The checker is free and needs no account: run the same check on your own domain and it names each agent and the robots.txt line responsible. The harder case is the one this study cannot see from the outside: a file that permits every agent while the CDN in front of it refuses them anyway. That is the case below.

Let AI assistants read your site

AI assistants have to read your website before they can mention it, and we find whatever is blocking them from reading yours and clear it.

  • 10 AI agents tested against robots.txt, CDN and WAF
  • The exact rule behind every block, quoted
  • A corrected robots.txt, written and applied for you
  • All 10 re-fetched after deploy, response codes in writing

Get the next edition

The panel is frozen, so the same domains get re-checked and the next edition reports what moved rather than a new sample. One email when it runs, and nothing else.

Why we have 5,000 domains of this

Standing measures whether AI assistants actually name a brand when someone asks for a recommendation, across five engines, with a confidence interval on every score. Whether the crawlers can reach a site at all is the floor underneath that, which is why we ended up counting it. The same crawler check runs on one domain, free and without an account, and reports the exact robots.txt line responsible.