Brand mentions or citations: which AI visibility KPI is stable enough to report?
A new stability study finds brand mentions hold steady day to day while cited sources churn. What that means for the number you put in front of a board.

The short answer
Report brand mentions as the headline and use citations as the diagnostic underneath. Otterly's stability study, published 11 September 2026, found brand mentions far steadier from day to day than the sources engines cite. Citations still tell you why an engine mentions you, but they move too much to carry a monthly target on their own.
Key takeaways
- Otterly's September 2026 study of daily US answers across the major AI engines found brand mentions much more stable than cited sources.
- Engines differ in how often they reuse sources from one day to the next, so a citation KPI blends engine behaviour with your own performance.
- In Semrush's analysis of 126 million US prompts from January to April 2026, ChatGPT cited about 15 sources per response against about 3 for Gemini.
- Larger prompt sets reduce reported variation more than repeat runs of a few prompts do.
- Our view: put brand mention share with an interval at the top of the report and show the citation map beneath it.
In this article
What the study measured
Otterly, which sells AI search monitoring software, published a study on 11 September 2026 asking how stable AI visibility metrics are from one day to the next.1 It ran a fixed set of prompts once a day in the US across ChatGPT, Perplexity, Gemini, Google AI Mode, Google AI Overviews, Microsoft Copilot and Claude, over several weeks, in three topic areas.
The question matters because most AI visibility reports move every month, and few say how much of that movement is noise. If a metric jumps around when nothing has changed, a team can spend a quarter reacting to it. This note covers what the study found, what it means for the headline number and what it does not settle.
Mentions held; citations churned
The central finding is that brand mentions were far more stable than cited sources.1 Most brands an engine named on one day were named again on later days. Many cited sources, by contrast, appeared once and never returned, and a small set of URLs collected a large share of all citations while most URLs were cited only once.
Engines also behaved differently. Otterly found that Perplexity reused its sources from day to day much more often than ChatGPT or Google AI Mode did.1 That means a citation-based KPI blends two things: how well you are doing, and how restless each engine is about where it looks.
Otterly's own recommendation is to use brand coverage as the top-line KPI and citations as the diagnostic layer beneath it. Our view: that is right, and it matches how we already report.
Why do citations move more than mentions?
Citations move more because there are more of them per answer and engines pick them from a wide pool. In Semrush's 2026 AI Visibility Index, built on 126 million US prompts from January to April 2026, ChatGPT cited about 15 sources per response against about 3 for Gemini.2 Semrush sells search and AI visibility software.
An engine that lists fifteen sources can swap several of them between runs without changing what it says about you. Your brand is mentioned either way; the supporting links differ. A mention is the conclusion of the answer. A citation is the evidence, and engines treat evidence as interchangeable.
Two KPIs, two jobs
Brand mention share (headline)
- Steadier from day to day
- Answers "are we in the answer?"
- Report with an interval
- Set targets on it
Citation sources (diagnostic)
- Churns between runs
- Answers "why are we, or a rival, in it?"
- Report as a map of trusted sources
- Use it to choose fixes
Prompt set size matters more than repeats
The study also found that larger prompt sets cut the variation in reported brand coverage sharply.1 A report built on a handful of prompts can swing widely with no change in reality; a report built on many prompts swings less.
This agrees with MaxAEO's July 2026 guidance on sample size, which argues that repeat runs of the same prompt are clustered rather than independent, so adding prompts buys more precision than adding repeats.3 MaxAEO also sells AI visibility measurement. We set out the arithmetic behind this in our measurement spec for buyers, and the cadence question in how often to re-run checks.
Our view: if a vendor's monthly report rests on a few dozen prompts, treat month-to-month changes in either KPI as noise until the interval says otherwise.
What the study does not settle
It is a vendor study, published by a company that sells monitoring tools, and it has not been peer reviewed. It covers US answers only, three topic areas and one run per prompt per day. Stability in other countries, languages and categories may differ, and the period is short.
It also measures stability, not value. A steady mention rate is easier to report, but it does not prove that being mentioned sends buyers your way. For that link to revenue, you still need the referral and pipeline signals we described in which AI search number belongs in a board report.
Finally, "brand coverage" in the study and "share of answer" in our reports are close but not identical definitions. When you compare vendors, ask for the formula, not the label.
How we report it
A Sigzen AI audit puts brand mention share at the top, with a confidence interval, and the citation-source map underneath it. Our AI answer monitor runs 150 buyer prompts across six engines (ChatGPT, Gemini, Perplexity, Claude, Copilot and Google AI), each prompt on three separate days. It uses official APIs where they are offered and a licensed AI-answer data provider otherwise.
The citation map then answers the practical question: which sources do the engines trust in your category, and are you in them? Targets go on the mention share; fixes come from the map. The audit's scope and price are on the audit page and our pricing page.


