Why two AI visibility tools report different numbers for you
You open Dashboard A and see that your brand appears in 42, plus or minus 5, out of 100 queries. You open Dashboard B and see a score of 12, plus or minus 3.
Your immediate assumption is that one tool is broken. Either Dashboard A is inflating your visibility to make you feel good, or Dashboard B is missing obvious recommendations.
In almost every case, neither tool has a software bug. They are reporting different numbers because they are measuring fundamentally different things. AI visibility tools rarely publish their sampling methodologies, their exact prompt baskets, or their definition of a mention. When two tools measure non-deterministic systems using different inputs and different definitions, divergence is guaranteed.
Understanding why these dashboards disagree requires looking at four structural differences: prompt baskets, sampling frequency, engine selection, and mention logic.
Prompt baskets and intent scope
An AI visibility tool does not measure the entire internet. It runs a set of text queries through model APIs and evaluates the output. The specific prompts in that basket dictate the score.
If Tool A tests a narrow basket of high-intent buyers looking for your exact product category, your score will be higher. If Tool B tests a broad basket of general industry questions, your brand will appear less frequently.
Most visibility vendors hide their prompt baskets or allow users to generate custom lists on the fly without standardization. Comparing scores across tools that use different prompt baskets is like comparing the traffic of two websites using completely different keyword sets.
Without a fixed, reproducible prompt basket, a score tells you how well you performed against an undisclosed list of questions on a single afternoon.
Single-run sampling versus repeated runs
Large language models are probabilistic. If you ask ChatGPT the same question ten times in a row, you will not get ten identical answers. The model chooses words based on calculated probabilities, meaning citations appear and disappear between runs.
Most visibility vendors run each prompt exactly once per update cycle. They do this because running thousands of API calls across multiple engines is expensive.
A single run provides zero statistical confidence. If a tool runs a prompt once and your brand appears, it records a positive hit. If it runs it five minutes later and the model omits your brand, it records a miss. A single-run scan yields a binary noise signal rather than a stable measurement.
To produce a statistically valid score, a tool must query the model repeatedly. At Standing, our method uses a fixed prompt basket, run five times per question per engine across five covered engines: ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews.
Because we execute five runs per question, we do not present scores as isolated integers. Reporting a bare single-number score without an uncertainty interval is mathematically incomplete. We report results with a Wilson interval within a 95% confidence interval, showing a range like 34, plus or minus 6. Competitors run each prompt once, which is why none of them publishes a confidence band. When you compare a single-run tool to a statistical sampling tool, the numbers will never match.
Engine coverage and crawler mechanics
AI visibility depends heavily on which engines a vendor monitors and whether your technical setup allows those engines to read your site.
The primary engines that drive user recommendations are ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews. Some visibility platforms only track ChatGPT. Others attempt to aggregate multiple engines into a single blended index. If Dashboard A measures only ChatGPT while Dashboard B averages three different platforms, their scores will differ.
Divergence also happens at the infrastructure level. AI engines retrieve information in two distinct ways: pre-training on historical web data and live web searching during prompt execution.
Many site owners confuse training crawlers with answer crawlers. Blocking a training crawler prevents an AI company from using your content to train future model weights. It does not stop an engine from retrieving your site during a live search answer today. Blocking an answer crawler, however, immediately cuts off live search citations.
In our crawler index version 2026-08-08, we evaluated a panel of 54,082 domains, of which 33,670 returned a readable robots.txt file. Across those 33,670 readable domains, 5,497 block at least one AI agent, but only 2,576 block at least one answer crawler.
The distinction between training and answer crawlers is evident when looking at individual agent block rates across those 33,670 readable domains:
- GPTBot (training): blocked by 5,080 domains
- ClaudeBot (training): blocked by 4,603 domains
- Google-Extended (training): blocked by 4,275 domains
- PerplexityBot (search): blocked by 2,147 domains
- ChatGPT-User (search): blocked by 2,074 domains
- OAI-SearchBot (search): blocked by 1,579 domains
- Perplexity-User (search): blocked by 1,402 domains
- Claude-SearchBot (search): blocked by 1,400 domains
- Claude-User (search): blocked by 1,384 domains
- Googlebot (other): blocked by 610 domains
Even technology-focused companies exhibit this pattern. In our YC census conducted on 2026-08-07, out of 4,226 active Y Combinator companies evaluated, 3,755 had readable robots.txt files, and 253 blocked at least one AI crawler.
If one visibility tool measures an engine that relies on live web search while your website blocks ChatGPT-User or Perplexity-User, that engine cannot cite your site today. A different visibility tool measuring an offline base model may still report citations based on old training data.
Definitions of a mention
The fourth source of disagreement is how each platform defines a mention.
When an engine produces a response, what counts as a recommendation? Is a plain-text mention of your company name counted? Does the tool require an active hyperlink back to your domain? How does the tool handle comparative context?
If a model writes "Company X is a legacy vendor that lacks modern features," a naive parsing tool will record a positive mention for Company X because the string appeared in the output. A more rigorous parsing tool will categorize that output as a negative citation or discard it entirely.
Some vendors count indirect mentions, such as listing a product feature that belongs to you without naming your brand. Others strictly count cited domain sources in footnote links. If Dashboard A counts every raw text string while Dashboard B counts only hyperlinked sources in top-line recommendations, Dashboard A will report significantly higher numbers.
How to evaluate AI visibility tooling
When evaluating visibility software, recognize what to buy and what to ignore.
Do not buy tools that guarantee citations, specific rankings, or model mentions. AI engines are non-deterministic systems that change their retrieval weights constantly. Any vendor promising guaranteed placement is selling claims that cannot be fulfilled.
Do not buy tools that claim schema markup or an llms.txt file will improve your AI recommendations. We have seen no evidence that adding schema markup or an llms.txt file improves citation frequency in live search engines.
Where competitors are genuinely better: if you need end-to-end marketing automation, continuous content generation, and automated landing page rewrites, several enterprise marketing suites offer broader workflow tooling than we do. Those tools are built for teams that want an all-in-one execution platform.
Standing is built specifically for measurement rigor. We do not sell content creation engines or promise magic ranking fixes. We provide clear, repeatable tracking across five major engines: ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews.
Our pricing is straightforward:
- Track: $100 per month for 3 domains.
- Optimize: $300 per month for 5 domains with custom prompts.
- Agency: $500 per month for 50 domains with a monthly re-scan.
We publish our sampling methodology, run five repetitions per prompt per engine, and express every metric as a score with a Wilson interval within a 95% confidence interval. When your numbers change on Standing, you can verify whether the shift represents actual movement in model behavior or simple statistical noise.
See where you actually stand
Run a free check on your domain. Five AI surfaces, the real buying questions, about a minute. No signup.
Run a free check