Standing vs Evertune
A research-led pure-play with the deepest disclosed sampling in the category.
We make Standing, so read this accordingly. Every claim about Evertune below comes from their own published pricing, documentation or help centre, verified August 6, 2026. Where they have not published something we say so rather than guessing, and the first section is where they beat us.
Where Evertune is better
- Sampling depth, decisively. 100 runs per prompt per model against our 5. This is the one place in the category where a rival is straightforwardly ahead of us on methodology, and it is not close.
- A 150M real-prompt panel for basket generation, and support for up to 250 customer prompts.
- A published methodology page, which most of this category does not have.
The prompt basket
Evertune: Not published publicly.
Not applicable. Their unique prompt count is also not published: the headline figure of 100,000 prompts tracked counts responses rather than questions, and roughly 90 unique is implied by arithmetic but never stated by them.
Standing: a fixed basket built from your category, priced per domain rather than per prompt. It is the smallest basket in the category and the only fixed one, which is the trade: fewer questions, but the same questions every scan, so a change in the score is a change in the answers rather than a change in what we asked. It also means two customers in one category are directly comparable.
How many times each question is asked
Evertune: 100 runs per prompt per model, stated on a dedicated public methodology page and repeated in their help centre. This is the deepest disclosed sampling in the category and it beats ours twenty to one. It is a documented method rather than marketing.
Standing: 5 runs per prompt per engine on paid scans, 3 on the free check, published. Roughly 300 responses per paid scan. This is the number that decides whether a score is a measurement or a single draw from a distribution.
What each publishes about uncertainty
Evertune: None published, which is the odd part. They assert the conclusion of rigour without showing a statistic: statistically significant is used as an adjective, with no null hypothesis, alpha, interval width or standard error anywhere. With 100 runs they could publish roughly a ±10 point band at 50% and choose not to. Nothing tells a customer how large a week-over-week move must be to be real.
Standing: a 95% confidence band on every score, shown in the product, with alerts gated on whether a move exceeds that domain's own historical drift. An engine that fails more than half its calls reports that it did not respond rather than a measured zero. Our own weakness, stated: the band covers within-scan sampling and does not cover platform drift between scans, which is why the model version is on the report.
Where Standing is better
- We publish the band their sampling depth would let them compute. Depth without a stated interval means the customer still cannot tell a real move from a fluctuation.
- Published pricing, and a free check with no account.
- A stated policy for engine failure.
Which to pick
Choose Evertune if
You want the deepest sampling available and can work with a vendor that does not publish pricing. On raw sampling depth they are the strongest option in the category.
Choose Standing if
You want the uncertainty reported to you rather than absorbed into an unqualified claim of significance.