Why your AI visibility score moves when you changed nothing
You run a scan on Monday and score 34. You run the same scan on Friday and score 41. You changed nothing in between.
Did anything happen?
Most AI visibility tools cannot answer that question, and the ones that show you a chart with an upward arrow are answering it wrong. This is the single most important thing to understand about the category, and it's the thing least often explained.
Where the variance comes from
There are three separate sources, and they need to be kept apart because only one of them is news.
Sampling variance. Ask ChatGPT "what's the best CRM for a small sales team" five times and you will not get five identical answers. These models sample from a probability distribution. Names appear and drop out between runs with no change in the world at all. This is the largest source of movement and it is pure noise.
Retrieval variance. Grounded answers depend on what a live search returned at that moment. Search results shift hour to hour. Two identical questions minutes apart can be answered from different sources.
Real change. A new review lands on a site the engine trusts. A competitor gets written up. Your own page finally gets indexed. This is the only one you'd want an alert about.
The first two are noise, and they are big. A tool that reports a single number and shows you the line between Monday and Friday is presenting mostly noise as though it were mostly signal.
What a single scan can and cannot tell you
A single run of a single question is one draw from a distribution. Reporting it as a fact is the central error in this category.
The fix is standard statistics, not anything clever. Ask each question several times, treat the runs as a sample, and report the interval rather than the point:
half-width = 1.96 × standard deviation / √(number of runs)
That's the 95% confidence interval. A score of 40 with a half-width of 8 means: given how much the answers varied, the honest reading is somewhere between 32 and 48.
Now go back to the opening question. Monday 34, Friday 41. If the band on each is ±8, those two intervals overlap almost entirely. Nothing happened. The seven-point "improvement" is a tool showing you a coin landing differently.
This is why the band matters more than the score. Not because uncertainty is interesting, but because without it you cannot tell whether any given change is worth acting on, and acting on noise is how you end up attributing a random fluctuation to whatever you happened to do that week.
Why more runs is not a free fix
The obvious response is: run it more times, and the band shrinks.
It does, but slowly. The square root in the denominator means going from 5 runs to 20, which is four times the cost, cuts the band in half. To halve it again you need 80. Every run is a paid API call to a frontier model, so precision has a linear price and a square-root benefit.
That's the real constraint in this category, and it's why so many tools quietly report a single run: it's the cheapest possible answer, and the cost of it is invisible to the buyer.
Distinguishing noise from a change
Within-scan variance tells you how precise one measurement is. It doesn't tell you whether a difference between two scans is real, because between-scan movement includes drift the sampling band never sees. Engines swap models underneath you without announcement, and a model change can shift the whole distribution.
So the threshold for "something happened" has to be built from a domain's own scan-to-scan history, not from the within-scan band. A robust way to do it:
- Take the absolute differences between consecutive scans for that domain.
- Take the median of those differences, scaled by 1.4826. That is a median-based estimate of the standard deviation that isn't wrecked by one outlier the way a mean would be.
- Alert only when a move exceeds that.
This needs history. Around five scans before the estimate means much, which is an awkward but honest thing to say to a customer in their first month. The alternative, alerting from scan two on a threshold you invented, produces alerts that feel responsive and mean nothing.
What this implies about the tools
Three questions to ask any AI visibility product, including ours:
How many times do you ask each question? If the answer is once, the number is one draw from a distribution presented as a measurement.
What's the confidence interval on this score? If there isn't one, the tool cannot distinguish its own noise from your progress. It will show movement every week regardless of whether anything happened, because that's what sampling does.
When an engine fails, what do you record? If a timeout is scored as "not mentioned," you'll see drops that never occurred. Missing data and confirmed absence are different things, and conflating them manufactures bad news.
There's a commercial reason these questions are rarely answered. A product whose core artefact is a number going up is awkward to reconcile with a band saying last month's improvement was noise. It's easier to sell the 40 than to sell "somewhere between 32 and 48."
We think the second one is the only version worth paying for, which is why every score we report ships with its band and why alerts here are gated on significance rather than on movement. You can read the full method, including what it can't tell you, or run a free check and see the band on your own domain.
See where you actually stand
Run a free check on your domain. Five AI surfaces, the real buying questions, about a minute. No signup.
Run a free check