Can AI read this page?
Can an AI engine reach this page, tell what it is, and find a passage worth quoting?
Free, unlimited, no signup. Any page. Reads the HTML you serve and your robots.txt, and asks no AI engine anything, so it costs nothing to run.
If a check fails
Blocking failures first, because nothing below one of those can help until it passes.
Common questions
12 checks in three groups. Reach, where every check is blocking because nothing below can help until they pass: The page loads; Answer engines are allowed in; No noindex directive; A bot sees the same page a person does. Understand, whether an engine can work out what the page is without inferring it from the layout: Title; Meta description; One clear h1; The page declares what it is about; Canonical URL. Quote, whether there is a passage here worth lifting, because assistants quote passages rather than pages: The text is in the HTML; Subheadings to quote from; Answers a question in the reader's words. Each one comes back with its own fix written out when it fails.
Because those govern training, not citation. GPTBot, ClaudeBot, Google-Extended decide whether your content is used to train and improve models, and allowing them donates your writing to a training corpus. Refusing them does not remove you from an answer. The agents that decide whether an assistant can quote you are different ones, and these are what this check counts: OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User, Googlebot. Blocked training crawlers are reported and never counted against you, because allowing them is a decision about your content rather than a fix for your visibility. Google-Extended is the arguable one and sits on the training side deliberately: Google documents it as governing Gemini Apps and Vertex AI rather than Search or AI Overviews, and AI Overviews is what we measure, so counting it as blocking would report a risk we do not observe.
Because nobody can convert page properties into a probability that an assistant names you, and a number implies somebody can. Every scanner that prints one is adding up things it can see and presenting the total as a thing it cannot. This reports how many checks passed and lists the blocking failures separately, because one robots.txt line that shuts out every answer engine is not the same size of problem as a missing meta description.
We request the page twice, once with a browser user-agent and once with an openly declared bot one, and compare how much text comes back. A large gap means an engine is reading a different, shorter page than your visitors see, usually because of a paywall, an interstitial or a bot wall. It is the finding most likely to explain a page that looks fine and gets quoted from oddly.
It matters if the content only exists after JavaScript runs. Several of the agents that gate citation fetch HTML and do not execute scripts, so a client-rendered route can serve them an empty shell. The served-text check counts the words present in the HTML before any script runs, which is what those agents actually see.
llms.txt; Blocked training crawlers (GPTBot, ClaudeBot, Google-Extended); Word count and keyword density; A single score out of 100; Anything we could not establish. Every one of those earns points somewhere else in this category. None of them is something we can show affects whether an engine names you, so scoring them would be inventing a requirement and then selling the fix. A proposed convention, not a standard. No engine has published that it reads one, and none of the five we scan has said it affects whether you are cited. We ship a generator for it because it costs nothing and may one day matter. Scoring a page on it today would be inventing a requirement. These govern whether your content trains a model, not whether an assistant names you. Allowing them donates your writing to a training corpus; refusing them does not remove you from an answer. The agents that do gate answers are checked above, separately, and those are worth acting on. Borrowed from search ranking, where it was already weak. An assistant is choosing a passage to quote, and length is not what makes a passage quotable. We check that there is text at all, which is a different and much more common failure. Nobody can convert page properties into a probability that an engine names you, including us, and a number implies otherwise. This report counts checks passed and lists blocking failures separately, because one robots.txt line that shuts out every answer engine is not four missing meta descriptions. If a fetch was refused or a file was unreadable, the report says so rather than scoring it as a fail. Absence of evidence is not evidence of absence, and a tool that treats a Cloudflare refusal as a finding will tell most of the internet it is broken.
No, and this is the honest limit of every page-level scan including this one. These checks establish that an engine can read you, identify you and find something to quote. Whether it then names you depends on what it has read about you elsewhere: third-party pages, comparisons, forums, listicles, and what your competitors have that you do not. Measuring that means asking the engines, which is a different tool.
Two page requests and a robots.txt request, so effectively nothing, which is why it has no cap and needs no account. Our visibility check does cost real money because it queries five AI engines many times over, and that is the one thing on this site that is metered.
More free tools
The page is readable. Now ask the engines.
Everything this checks is a property of your page, which is where every other page-level scanner stops and calls the total a score. The measurement that follows it is what an assistant says when a buyer asks, and the only way to take that one is to ask the engines directly, repeatedly, and report the band.