Skip to content

Can you trust your AI visibility score? A measurement spec to hand any vendor

AI answers vary from run to run, and repeat runs are not independent. The spec to ask any vendor for, and how a corrected margin of error changes a report.

By Published Updated 15 min read
A bright point of light at the centre of a dark field, surrounded by soft concentric halos and small clustered dots scattered around it.

The short answer

Only if the report shows how it was measured and how wide its margin of error is. AI engines often give different answers to the same prompt, and repeat runs of one prompt are not independent, so many reports understate their error. Ask for the prompt set, runs, engines, countries and languages, dates, and an interval on every number.

Key takeaways

  • An AI visibility score is an estimate from a sample of answers, so it means little without its margin of error.
  • Repeat runs of the same prompt tend to agree with each other, so treating them as independent makes the interval look narrower than it is.
  • For a fixed number of answers, more distinct prompts usually narrow the interval more than more repeats of the same prompts.
  • Each country and language is its own sample: report each with its own interval, or do not split the score.
  • Ask every vendor, including us, for the same seven items before comparing two numbers.

Report the interval, not the point

An AI visibility score is a sample estimate, so the first thing to ask for is the interval around it. A vendor runs a set of prompts through ChatGPT, Gemini or another engine, counts the answers that name your brand and divides. That share describes a sample of answers, collected on certain days, in one country and one language. Run the same prompts next week and the share will move, even if nothing about your brand has changed.

So the useful question is how far the number could move from noise alone. A margin of error answers it, and it is the figure most reports leave off the cover page. Our view: a score without an interval should not be used to approve a budget, judge an agency or compare two tools.

A disclosure before we go further. We sell an AI visibility audit whose main output is share of answer with confidence intervals, so we have an interest in you asking for intervals. The spec below applies to our reports too.

Why does the same prompt give different answers?

Because an AI engine writes a fresh answer every time, from sources it picks at run time, so two runs of one prompt are two draws rather than two copies. Two forces drive the spread: how the answer is generated, and how many sources the engine reads before it writes.

Engines generate answers, they do not look them up

A search results page is retrieved. An AI answer is written. The model predicts its reply word by word, and most consumer engines add a little randomness to each choice so that replies read naturally. Ask the same question twice and the wording, the order of brands and sometimes the brands themselves can change.

Other things add to the drift. Engines update their models without notice. Answers that use live search depend on what the search step returned at that moment. Logged-in users may get replies shaped by memory, history or location. A report built from one run of each prompt captures a single draw from all of that and calls it your score.

It helps to think of this as polling rather than rank tracking. A rank is a position you can look up. A presence rate is a probability you have to estimate, and every estimate comes with error.

Engines read different numbers of sources

Engines also differ in how widely they read. In Semrush's analysis of 126 million US AI search prompts, collected January–April 2026, ChatGPT cited an average of 15 sources per response and Gemini an average of 3.1 Semrush sells an AI visibility toolkit, so it has a stake in this category, but the gap is wide enough to matter to any buyer.

US prompts, January–April 2026

Average sources cited per response

ChatGPT15
Gemini3
Source: Semrush 2026 AI Visibility Index, 126 million US prompts, January–April 2026.

An engine that cites many sources per answer gives a brand more chances to appear, and more ways to drop out on the next run. An engine that cites few concentrates attention on a short list. The same brand can be steady on one engine and erratic on another, and a single blended score hides the difference.

Our view: ask for every engine to be reported separately, each with its own interval, before anyone shows you a blend.

What others have already worked out

Most of the statistics in this post are not new, and the people who published them first deserve the credit. Tracking vendors and researchers have spent this year working out how many prompts and runs a reliable measurement needs.

Gumshoe's February 2026 piece works through how much data it takes to reach a given margin of error across prompts, personas and models.2 The cloro article, published in August 2026 and updated in September, covers prompts, runs and alert thresholds for brand monitoring.3 Obsero's May 2026 piece argues for tracking topics rather than single prompts.4

MaxAEO has gone furthest on method. Its early July 2026 guide covers prompt-runs and the power needed to detect a change.5 Its later July piece treats repeat runs as clustered, applies the design effect, compares prompts with repeats at a fixed budget, separates presence from position and proposes a one-line reporting standard.6

Three preprints form the research base. "Don't Measure Once" by Schulte, Bleeker and Kaufmann (April 2026) argues for measuring visibility repeatedly, as a distribution.7 Dmitrij Żatuchin's variance-components paper (July 2026) breaks answer variation into its sources, including language and model.8 His "Dice Roll Method" (September 2026) sets out a standard protocol for repeated-query audits.9

This post adds the buyer's side. It offers a seven-item spec you can hand any vendor, including us, the same sample report shown with a naive and a corrected interval, and reporting per country and language. The spec is close to MaxAEO's reporting standard and to the Dice Roll protocol. The maths is not new. Our aim is to put it in the hands of the person who receives the report.

How wide is the margin of error?

Wider than most reports suggest. At an illustrative presence rate of 30%, 150 answers give a margin of about ±7 points if every answer is independent, and repeat runs of the same prompts push it further out.

The simple arithmetic

The textbook interval for a proportion is the estimate plus or minus 1.96 standard errors, at a 95% confidence level. The standard error depends on the rate and the number of answers.10 The NIST handbook gives this normal approximation and recommends the Wilson interval for most combinations of sample size and rate, because it behaves better for small samples and rates near the extremes.10

The arithmetic is unforgiving. At the same illustrative rate of 30%, 50 answers give a margin of about ±12.7 points. You need 150 answers to get near ±7 and 750 to get near ±3. Halving a margin takes four times as many answers, which is why small trackers swing so much from week to week.

Why repeat runs are not independent

Most trackers run each prompt several times, and that is good practice. The trap is counting every run as a fresh observation. Answers to one prompt tend to resemble each other, because the engine sees the same words and often reads the same sources.

Survey statisticians call this clustering and correct for it with the design effect: deff = 1 + (m − 1) × ICC. Here m is the number of runs per prompt, and ICC measures how alike the runs of one prompt are.11 Divide the number of answers by the design effect and you get the effective sample size, the number of independent answers your data is worth.

MaxAEO reports that, in its own data, answers to the same prompt were strongly correlated across runs.6 We have not checked the vendor's data, and your prompt set will behave in its own way. That is why a report should state the correlation it assumed or measured.

Illustrative

Illustrative: margin of error at a 30% presence rate

50 answers, counted as independent±12.7 pts
150 answers, counted as independent±7.3 pts
450 answers, counted as independent±4.2 pts
750 answers, counted as independent±3.3 pts
150 prompts × 3 runs, correlation 0.3: effective sample 281±5.4 pts
150 prompts × 3 runs, correlation 0.5: effective sample 225±6.0 pts
150 prompts × 3 runs, correlation 0.7: effective sample 188±6.6 pts
Illustrative. Correlation values are assumptions, not measurements. 95% level, normal approximation; methods from NIST and the UN sampling guidelines.

More prompts or more repeats?

For a fixed number of answers, spreading them across more distinct prompts usually narrows the interval more than repeating fewer prompts. The illustrative table below holds the budget at 450 answers and assumes a within-prompt correlation of 0.5.

DesignMargin of errorWhat it tells you
450 prompts × 1 run±4.2 ptsBroad, but no view of run-to-run variation
150 prompts × 3 runs±6.0 ptsBroad, plus a read on how stable answers are
50 prompts × 9 runs±9.5 ptsStable answers to a narrow set of questions
25 prompts × 18 runs±13.1 ptsPrecise per prompt, vague about the brand

Illustrative. Assumed within-prompt correlation of 0.5, 95% level, normal approximation with the design effect.

Our view: a few repeats per prompt are worth paying for, because they show how stable each answer is. Past that point, buy breadth. And ask any vendor for the effective sample size, never only the raw count of answers.

The same report, two intervals

The method can change the story more than the data does. Here is one illustrative quarterly report for a fictional brand, Northwind Digital, read two ways.

The vendor ran 150 buyer prompts three times each on one engine, collected 450 answers and found the brand named in 135 of them. That is a share of answer of 30%. Every number in this example is illustrative.

The same report, two intervals

Counted as independent

  • Share of answer: 30% (Illustrative)
  • 95% interval: 25.8–34.2% (Illustrative)
  • Sample: 450 answers, as if all were separate

Corrected for repeat runs

  • Share of answer: 30% (Illustrative)
  • 95% interval: 24.0–36.0% (Illustrative)
  • Effective sample: 225 (150 prompts × 3 runs, assumed correlation 0.5)
Illustrative. Northwind Digital is a fictional brand, and the correlation is an assumption, not a measurement.

The point estimate is identical in both columns. What differs is how much you should believe it. On the left, the interval spans about eight points. On the right, it spans twelve.

That gap decides real questions. Suppose next quarter's illustrative report shows 34%. Read the first way, it looks like a rise to the top of the old range, and someone books it as a win for the new content programme. Read the corrected way, 34% sits comfortably inside what the first quarter could have produced by chance. Nothing has been shown yet.

Strictly, comparing two quarters calls for an interval on the change itself, which is wider than either quarter's own. A careful vendor computes it for you. Check, too, whether anything besides your work moved between quarters: the engine version, the prompt set, the country or the access method. If any of them changed, the two numbers are not the same measurement.

Country and language are separate samples

Treat them as separate. An answer to a prompt asked in German, from Germany, is a different measurement from the same prompt asked in English from the United States, and pooling the two hides what is happening in each.

Three glass bell jars on a dark reflective surface, each holding its own small cloud of glowing particles in a slightly different shade of blue or violet.
Each market is its own sample, and pooling them blurs what is happening inside each one.

Why the language of the prompt matters

Engines tend to retrieve sources in the language of the question, and the sources differ by market. A brand with strong coverage in English trade press may barely exist in Spanish-language reviews or Hindi-language forums.

A July 2026 preprint on Central and Eastern European brands found that the language of the query explained more of the variation in answers than which brand was asked about.8 That study covers one region and a handful of models, so treat it as a reason to check rather than a rule.

Our view: test it on your own markets before deciding whether one global score is enough. If answers differ a lot between languages, a pooled number averages two different realities into one that describes neither.

What to ask for per market

Splitting a prompt set across markets shrinks each market's sample, and the interval widens to match. The numbers in the example below are illustrative.

So ask for one of two things. Either per-market numbers, each with its own interval and sample size, or a single score clearly labelled as pooled. Do not accept per-country breakdowns of a pooled sample that carry no interval of their own.

Is your position in an answer meaningful?

Less than it looks. Position within a single answer is a weak signal, while presence across many runs is the stronger one.

Rank tracking taught marketers to care about position, and many AI visibility reports carry the habit over: named first, named third, named fifth. In a written answer, order often follows the structure of the reply rather than a judgement of quality. An engine that groups options by price or by company size lists them in that order, and the next run may group them differently.

Position also has a hidden sample-size problem, because it exists only when the brand appears. If Northwind Digital is named in 135 of 450 illustrative answers, its average position rests on those 135 answers alone, so its interval is wider than the headline share's.

Our view: treat presence across runs as the main number and position as a diagnostic. Read position next to the answer text. Being listed fifth as "the best fit for large teams" may be worth more than being listed first with a warning about price. Ask the vendor to report how the brand was described as well as where it appeared.

Seven things to ask any vendor, including us

Ask for these seven items in writing before you compare any two numbers, whether the numbers come from a tool, an agency or Sigzen AI.

Seven things to ask any vendor

  1. The prompt set

    Fixed and shared with you: how many prompts, and how they were chosen.

  2. Runs per prompt

    How many times each prompt was run, and how the runs were spaced.

  3. Engines and access

    Each engine, its model or mode, and how it was reached: app or API, logged in or not.

  4. Country and language

    For every run, not only for the report as a whole.

  5. Collection dates

    The first and last day answers were collected.

  6. Definitions

    What counts as a mention, a citation and a recommendation.

  7. Intervals

    Every number with its interval, the method (naive or corrected) and the effective sample size.

Our checklist for buyers. It sits close to MaxAEO's reporting standard and the Dice Roll protocol.

Items one to five describe the sample. Without them you cannot tell whether two reports measured the same thing. An answer from a logged-in app with memory switched on is not comparable with one from an API call. A prompt set chosen after the vendor has seen the results is not a fair test.

Item six is where definitions hide. A mention is your name in the text. A citation is a link to your site or to a source about you. A recommendation is the engine telling the user to pick you. Reports that blend the three can show progress on the easiest one and call it success. When the description itself is wrong, fix that first: our guide to tracing a wrong AI answer to its source sets out the steps.

Item seven is what this post is about. Every number needs its interval, a statement of whether that interval was corrected for repeat runs, and the effective sample size. You can see how we report each item in our published method and glossary, and hold us to the same list. If a vendor cannot supply an item, that does not prove the score is wrong. It means you cannot check it, so give it less weight.

How Sigzen AI measures, and what to do this week

Sigzen AI's audit runs 150 buyer prompts across six engines, each prompt on three separate days, and puts an interval on every number it reports. It takes two to three weeks and costs $3,500 (₹1.5L), as listed on our pricing page. You get your share of answer against five competitors, a map of the sources engines trust in your category and three prioritised fixes. The method and glossary are on our playbooks page, and our own dashboard on the proof page is labelled sample data until real runs replace it.

The free AI Visibility Score is a screen, not a baseline. It tells you quickly whether engines name you at all. It is not built to detect a change over time, so do not use it to judge a programme. If you are comparing audit prices, our dated list of published GEO prices shows what other providers charge and what they include.

This week, pull your latest visibility report and check it against the seven items. Write down what is missing and ask the vendor for it. Then find the last change in your score that you celebrated or worried about, and check whether it was larger than its interval. If the real question is whether a rise reaches revenue, our plan for proving AI visibility in pipeline reads the score next to four other signals.

Sources

  1. Semrush, 2026 AI Visibility Index: 126M US prompts, Jan–Apr 2026 (Jun 2026)
  2. Gumshoe, How much data do you need to measure AI visibility with confidence? (Feb 2026)
  3. cloro, AI visibility sample size: how many prompts and runs (Aug 2026, updated Sep 2026)
  4. Obsero, How many prompts do you need to track AI visibility? (May 2026)
  5. MaxAEO, How many prompts to test AI visibility? Sample size and prompt-run math (Jul 2026)
  6. MaxAEO, AI visibility sample size and clustered runs (Jul 2026)
  7. Schulte, Bleeker and Kaufmann, Don't Measure Once: measuring visibility in AI search, arXiv preprint (Apr 2026)
  8. Żatuchin, preprint on answer noise across languages (arXiv, Jul 2026)
  9. Żatuchin, The Dice Roll Method: a standardized protocol for repeated-query auditing of LLM brand recommendations, arXiv preprint (Sep 2026)
  10. NIST/SEMATECH, e-Handbook of Statistical Methods, §7.2.4.1 Confidence intervals
  11. United Nations Statistics Division, Designing Household Survey Samples: Practical Guidelines, ST/ESA/STAT/SER.F/98, §3.5.2 (2008)

Questions readers ask

  • Enough to see how much the answers vary, which usually means a few repeats per prompt rather than dozens. Repeats of one prompt tend to agree with each other, so each extra run adds less information than a new prompt would. Beyond a few repeats, spend the budget on more distinct prompts, and ask the vendor for the effective sample size: the number of independent answers the data is worth.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.