t-online.de blocks crawlers that AI assistants cite from
t-online.de disallows 4 AI crawlers in robots.txt, including 1 that builds the index AI assistants quote from. Checked 2026-08-07, with the responsible line published.
Disallow: /
This affects citation
One of the disallowed agents builds the index assistants quote from, or fetches a page live to answer a question. Blocking that is a different decision from refusing to be trained on.
Crawlers that affect citation
These build the search index assistants quote from, or fetch a page live when a question needs it.
- PerplexityBotPerplexity
builds Perplexity's index
Training crawlers
Blocking these keeps content out of the models. It does not stop an assistant citing the site.
- GPTBotChatGPT
trains and grounds OpenAI models
- ClaudeBotClaude
trains and grounds Anthropic models
- Google-ExtendedGemini
controls Gemini grounding and training
Permitted
OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, Perplexity-User, Googlebot
What this does and does not show
Method. One request to t-online.de/robots.txt on 2026-08-07, parsed for Disallow rules applying to each of the 10 named agents. No JavaScript, no login, nothing that cannot be reproduced with curl in a few seconds.
This is a snapshot. robots.txt changes whenever the site owner edits it, and this page reports one day. Anyone can re-run the same check live, and if it disagrees with what is here, the live result is the true one.
Access is not citation. A permitted crawler may never fetch the page, and a fetched page may never be quoted. Equally, a blocked crawler does not make a site invisible: content an engine learned before a rule was added does not disappear, and other sites writing about t-online.de remain readable.
Check your own site
The same check runs on any domain, free and without an account. It reports training and search crawlers separately, because collapsing them is the mistake this page exists to correct.