Can AI read your site?
Run the crawler, page and markup checks on one domain, at a link you can send.
If something is failing
The report groups its findings this way for a reason: each one gates the next.
Common questions
Four fetches, and every one of them is something you can repeat with curl. robots.txt, parsed against each agent, quoting the rule responsible when one shuts a crawler out. The homepage fetched as each crawler, because robots.txt is a request rather than a wall and a CDN rule that refuses anything without a browser fingerprint is the common real failure. The homepage fetched as a browser, for whether there is text a machine can read without executing JavaScript and whether any of it is quotable. The structured data in the markup, for what the page declares about itself and whether the JSON-LD parses at all.
No. Every check here is an HTTP request, so it says whether an engine can read and quote your pages. Whether any engine names you is a different measurement that needs the engines asked, and it comes with a confidence interval.
No. GPTBot, ClaudeBot and Google-Extended govern whether your content trains a model. The agents that decide whether you can be cited are different ones, and this report checks them separately.
llms.txt; Blocked training crawlers (GPTBot, ClaudeBot, Google-Extended); Word count and keyword density; A single score out of 100; Anything we could not establish. Each of those is a point somebody else would award you, and none is evidence about whether an assistant names you. A proposed convention, not a standard. No engine has published that it reads one, and none of the five we scan has said it affects whether you are cited. We ship a generator for it because it costs nothing and may one day matter. Scoring a page on it today would be inventing a requirement. These govern whether your content trains a model, not whether an assistant names you. Allowing them donates your writing to a training corpus; refusing them does not remove you from an answer. The agents that do gate answers are checked above, separately, and those are worth acting on. Borrowed from search ranking, where it was already weak. An assistant is choosing a passage to quote, and length is not what makes a passage quotable. We check that there is text at all, which is a different and much more common failure. Nobody can convert page properties into a probability that an engine names you, including us, and a number implies otherwise. This report counts checks passed and lists blocking failures separately, because one robots.txt line that shuts out every answer engine is not four missing meta descriptions. If a fetch was refused or a file was unreadable, the report says so rather than scoring it as a fail. Absence of evidence is not evidence of absence, and a tool that treats a Cloudflare refusal as a finding will tell most of the internet it is broken.
Every finding on this page is a live HTTP request you can repeat yourself, robots.txt, the homepage as a browser, the homepage as each crawler, and the structured data in the markup. Nothing here asks ChatGPT, Claude, Gemini, Perplexity or AI Overviews anything, so nothing here says whether they name you. That question needs the scan, and the scan reports its confidence interval.
Yes, for any domain, with no signup. It costs us four HTTP requests, so there is nothing to meter.
More free tools
From readable to recommended
This report covers what an engine can fetch. The free check takes the next measurement: it asks the engines your buyers use, five times per question, and returns the score with the confidence interval it is accurate to.