What is a sample size?
Sample size is the number of prompts, queries, or pages included in a single measurement, and it sets how much random noise a result carries.
Sample size is just a count: how many prompts you ran, how many domains you crawled, how many pages you checked. But that count decides how much weight a single measurement can bear. Run the same ten prompts twice and the citation rate can swing wildly, just from which model responses happened to mention you. Run five hundred and the swing shrinks, not because you did anything differently, but because averaging over more draws cancels out more of the noise.
The mistake we see most often is treating a bigger sample as proof of accuracy rather than proof of precision. A large sample tightens the range around your estimate, it does not fix a biased one. If your prompt basket only contains questions your own team would think to ask, running it against a thousand variations of those same questions still won't tell you what a real buyer asks. Precision and correctness are different problems, and sample size only buys you the first.
For AI visibility specifically, sample size interacts directly with how low the numbers usually run. Citation rates in the low double digits need a meaningfully larger basket than a coin-flip metric would, to say anything with confidence, which is why a single afternoon of manual prompting rarely settles an argument about whether a brand's visibility moved.
Related
- Sampling varianceSampling variance is the run-to-run variation you get from asking a language model the same question more than once.
- Confidence bandA confidence band is the range around a measured score that reflects how much of the number is sampling noise rather than signal.
- Wilson intervalThe Wilson interval is a method for calculating the plausible range around a measured rate that stays sensible even when the rate is small or the sample is limited, unlike the textbook formula most people default to.
- Margin of errorA margin of error is the range around a measured number, such as a citation rate, within which the true value probably falls given the sample it came from.
Want to know where you actually stand on this? Run a free visibility check or try the free tools.