Every Y Combinator company, checked for AI crawler blocks
Not a sample. We read robots.txt on all 4,226 companies listed as active in the YC directory. Startups block AI crawlers at less than half the rate of the general web, which is the opposite of what most coverage of this assumes, and the interesting number is not the headline.
Out of the 3,755 whose robots.txt we could establish, not out of the 4,226 we asked.
What we threw out
471 of 4,226 refused us their robots.txt or failed to answer. Excluded, not counted as unblocked.
A rule we were never shown looks exactly like no rule at all. A 404 is an answer, since there is no file and nothing is disallowed. A 403 is not, and treating the two the same is a false all-clear.
Startups block less than the web does
Two different populations, checked the same way on the same day. Neither number is 'how many sites block AI crawlers'. Each is a fact about the set it came from, which is the entire point.
| Population | Checked | Readable | Block one or more |
|---|---|---|---|
| Y Combinator, all active | 4,226 | 3,755 | 6.7% |
| General web, frozen panel | 5,000 | 2,916 | 18.3% |
The gap is not a mystery. The general-web panel is full of newspapers, magazines and publishers, who block AI crawlers deliberately and in several cases are in active litigation about it. Startups mostly want to be found. That is exactly why a single “X% of sites block AI” figure quoted without its population is not information.
The number that actually matters
Blocking is not one decision. GPTBot trains the models. OAI-SearchBot builds the index ChatGPT cites from. A company blocking the first and not the second has refused to be trained on while staying entirely quotable, which is a deliberate and reasonable position.
trains OpenAI models
builds the index ChatGPT cites from
the ones where citation is genuinely affected
Among YC companies, GPTBot is disallowed 18.9× more often than OAI-SearchBot. On the general web the same ratio is 2.9×.
That difference is the most interesting thing we found. Startups are not blocking less because they are careless about it. They are drawing the line in a more sophisticated place: keep us out of the training data, keep us in the answers. Publishers, who make up much of the general-web panel, more often block both.
Almost every crawler checker on the market collapses these into one “AI bots blocked” count and reports a training-only block as “invisible to AI”. It is wrong, and on this population it is wrong 229 times out of 253. If you take one thing from this page, take that.
By sector
Sectors with at least 50 companies whose robots.txt we could establish. Below that a percentage is mostly noise, and printing it anyway is the sort of false precision this page exists to argue against.
- Industrials30/3129.6%
- Consumer29/3508.3%
- Healthcare33/4806.9%
- B2B129/20056.4%
- Fintech24/4006.0%
- Education4/785.1%
- Real Estate and Construction2/882.3%
The 24 companies a search crawler cannot reach
These are the cases where the block genuinely affects whether an AI assistant can cite the company, rather than only whether it can train on them. Every row is a line the company published in its own robots.txt and can be checked in about four seconds. We are not claiming any of these are mistakes; some will be deliberate.
| Company | Domain | Search crawlers disallowed |
|---|---|---|
| Clau | clau.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Eligible | eligible.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Koshex | koshex.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Markhor | markhor.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Metofico | metofico.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| MICSI * | micsi.co | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| NALA | nala.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| UpCodes | up.codes | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Margin | usemargin.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| XTraffic | xtraffic.com | ChatGPT-User, Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Degla Inc | degla.ai | ChatGPT-User, Claude-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Kiwi Biosciences | fodzyme.com | Claude-SearchBot, Claude-User, OAI-SearchBot, Perplexity-User |
| Twine | usetwine.com | ChatGPT-User, OAI-SearchBot, Perplexity-User, PerplexityBot |
| Adni | adni.co | ChatGPT-User, Claude-User, Perplexity-User |
| Fenrock AI | fenrock.ai | ChatGPT-User, PerplexityBot |
| ZBiotics | zbiotics.com | ChatGPT-User, PerplexityBot |
| Canary Technologies | canarytechnologies.com | Claude-SearchBot |
| Mentra | mentra.glass | ChatGPT-User |
| Morphle Labs | morphlelabs.com | Perplexity-User |
| Sunflower | sunflowersober.com | ChatGPT-User |
| TAG | tagme.pk | ChatGPT-User |
| Taloflow | taloflow.ai | ChatGPT-User |
| Async | withasync.com | ChatGPT-User |
| yhangry | yhangry.com | ChatGPT-User |
* Re-checked August 7, 2026. Every row above was measured again before publishing. 1 of 24 could no longer be confirmed:
- micsi.co — Site unreachable on re-check; robots.txt could no longer be established.
The counts elsewhere on this page are left as measured on August 7, 2026. A site going down afterwards does not make the original reading wrong, and editing the numbers to match a later day would make the study unreproducible.
How this was measured
Population. Every company listed as Active in the Y Combinator public company directory, deduplicated by domain. Taken from the public YC directory at yc-oss.github.io. This is a census, so there is no sampling frame to defend and no sampling error to report.
Method. One request to /robots.txt per domain, parsed for a Disallow applying to each of ten named AI user-agents. No JavaScript, no login, nothing a site owner cannot reproduce with curl.
Denominator. Every figure is out of the 3,755 domains whose robots.txt we established, never out of the 4,226 we asked. A 2xx we could parse counts. A 404 counts, because it is an answer: there is no file, so nothing is disallowed there. A 403, a timeout or a TLS failure does not count.
What this does not show. Not whether these companies get cited. robots.txt is a request, and a blocked crawler is not the same as an absent one. It also says nothing about pages behind a login, or about content an engine already learned before a rule was added. Anyone telling you a robots.txt line makes a company invisible to AI is overreading a text file.
Check your own
The same check that produced this page runs on any domain, free and without an account. It reports the training and search crawlers separately, because collapsing them is the mistake this whole page is about. Run it on your site.
The same question, asked of the general web
Our frozen panel of 5,000 domains, re-run on the same list each edition so a change means something. It is the number to compare this one against, and the reason the comparison is honest is that both travel with their denominators.