How we measure
Most AI visibility scores are a number with no stated method. Here is ours, in enough detail to argue with.
- Prompts per scan
- 10
- Runs per prompt
- 5
- Engines asked
- 5
The six checks
Three of them read your website and cost nothing. Three of them ask the engines and cost real money every time. The line between those two is where most of this category goes wrong.
- 1Crawler accessWhether each of the 10 named AI agents is allowed to fetch your site, read from robots.txt and from what your server actually returns rather than from what the file says.Free to run
- 2Readable answerWhether the answer to a buying question exists in the HTML your server sends, rather than appearing after JavaScript runs. An engine that does not execute your page cannot quote what is not there.Free to run
- 3Structured dataWhether your markup parses and says what you are. Treated as hygiene rather than a citation lever, because a controlled test found no causal lift on its own.Free to run
- 4Whether you are namedAcross 10 buying questions, 5 engines and 5 runs of every question, how often you appear at all, reported with the interval it is accurate to.Costs a scan
- 5What they pulled fromThe sources each engine cited, split into the ones you already appear in and the ones you do not. This is the list the fix roadmap is ordered by.Costs a scan
- 6Who they named insteadWhich competitors were named, how often, and on which engines. A score with nobody to compare it to is a number without a scale.Costs a scan
Checks one to three tell you whether an engine can read and quote your page. They cannot tell you whether it names you, and no property of your website can. That is what four to six are for, and it is the reason the first three are free here and the last three are what you pay for.
There is no score out of a hundred and no letter grade anywhere in this. Both require inventing the thresholds that separate one band from the next, and a count of checks passed carries the same information without the invention.
How a scan works
The same 10 buying questions, never edited between scans.
5 engines, 5 runs each, search grounding asserted on every call.
0 to 100, published with the 95% interval the runs support.
A move is reported once it clears this domain's own drift, measured from 5 scans.
250 answers behind one score, every scan, on the same questions as the scan before it.
Choices that cost us something
Each one makes the product slower, dearer or less flattering. The alternative is a number that looks better and means less.
Questions a skeptic should ask
Because a basket that changes between scans measures two things at once: the engines, and our own edits. A fixed set of 10 questions means a change in the score is a change in the answers, not a change in what we asked. It also means two customers in the same category are asked the same question, so the numbers are comparable.
Asking an engine the same question 5 times does not give the same answer 5 times. The band is the 95% confidence interval on that sampling: half-width 1.96 times the standard deviation over the square root of the number of runs. A score published without one cannot be checked, so every score we publish carries it.
Because most changes are noise. An alert fires when a scan-to-scan move is larger than the drift that domain's own history shows is normal, which needs at least 5 scans to establish. Before that we use a deliberately wide provisional band rather than guessing.
It is reported as degraded and excluded from the score rather than counted as a zero. An engine that did not answer is missing data, not evidence you are unmentioned, and scoring it as absence would show a drop that never happened.
No. Each engine is called through its own API with search grounding explicitly asserted, and a call that cannot confirm grounding fails rather than returning an ungrounded answer. Routing through an aggregator would silently break that, because the grounding setting is not something an aggregator can guarantee on our behalf.
Because the model each engine reported is recorded with the answer. Engines change models without announcing it, and a swap moves scores on its own, so the first thing to check when a number jumps is whether the machinery changed underneath it. Without that record the two are indistinguishable and the score stops being evidence.
Yes, by question, and never across a week. Two customers in the same category ask the same question, so the second scan of that question is nearly free and that is why the price is what it is. The week is the hard limit: cache for longer and real drift ends up inside the number rather than being measured by it.
Yes. Every report lists the full basket, the engines that answered, the engines that did not, and the model each one reported. If you cannot check the questions, the score is an assertion rather than a measurement.
Check it yourself
The free check runs 4 of the 10 questions at 3 runs each, and shows the same band. No account, no card.