Research
Three datasets, each published with the population it was measured on.
- Domains in the panel
- 5,000
- Crawlers checked on each
- 10
- Membership changes
- Never
Published
Every Y Combinator company, checked for AI crawler blocks
Not a sample. All 4,226 active companies. 24 block a crawler that decides citation.
Read itHow many sites block AI crawlers
Published estimates elsewhere run 3.5% to 88.9%, because almost none carry their population.
Read itThe AI Crawler Access Index
The widest of the three, re-run on a fixed list, raw dataset published.
Read itWhat a figure here is a share of
Out of the domains that answered, never out of the domains we asked.
- 2,916answered
- 2,084excluded, never counted as clear
- 535disallow an AI crawler
- 2,381disallow none
None of the three names the domains it counted. The per-domain index does, with the rule responsible on each.
What to do with a number like that
The measurements above are the same check at three populations, and the free crawler check is that measurement on one domain, with the responsible robots.txt line shown. Where it reports a block the file itself does not explain, the refusal is happening in front of the file.
AI assistants have to read your website before they can mention it, and we find whatever is blocking them from reading yours and clear it.
- 10 AI agents tested against robots.txt, CDN and WAF
- The exact rule behind every block, quoted
- A corrected robots.txt, written and applied for you
- All 10 re-fetched after deploy, response codes in writing