Skip to content

A model update just changed your AI visibility. How to measure before and after

When an engine update moves your AI visibility: freeze the prompt set, compare equal windows, separate engine changes from yours, report with intervals.

By Published 7 min read
Measurement and method, 7 min read — Two matching glass grids separated by a seam of light, the right one lit in a shifted pattern of brighter and darker cells.

The short answer

To measure the effect of a model update on your AI visibility, freeze the prompt set, compare equal windows before and after the release date, and check whether competitors moved too. Log your own site changes so you can separate them from the engine's. Report the change with an interval, and check the tracker itself before believing a sudden drop.

Key takeaways

  • Freeze and version the prompt set before an update; a changed prompt list makes any before-and-after comparison meaningless.
  • In Profound's data on 687 customers (July 2026), the share of ChatGPT Shopping retrieval from its product-feed catalog jumped from 8.26% to 61.54% on 10 July.
  • Promptwatch data reported by Semrush showed Reddit's share of ChatGPT citations falling from 3.8% to 0.5% in August 2026, with possible collection issues flagged.
  • Competitors measured on the same prompts are the control: if everyone moved, the engine changed; if only you moved, look at your own changes.
  • Report a before-and-after change with an interval, clustered by prompt, and say when the interval crosses zero.

Why a model update needs its own method

When an AI engine ships a new model or changes how it retrieves sources, your visibility can move overnight for reasons that have nothing to do with your work. The question leadership asks is simple: did we lose ground, and was it us? Answering it needs a before-and-after method set up in advance, not a screenshot taken the morning after.

Two changes this summer show why. Both moved results for many brands at once, and in one of them the measurement itself may have been part of the story. We use both as worked cases below. Google's spread of AI Overviews across branded searches, covered in our note on branded overviews, is a change of the same kind.

Our view: most teams discover an engine change through a falling dashboard and then try to explain it backwards. The method below reverses that order. It defines the comparison first, so the explanation is tested rather than assumed.

Worked case one: ChatGPT Shopping's feed shift in July

A retrieval change can reshape results on a single date. Profound, which sells AI visibility software, analysed ChatGPT Shopping results for 687 customers from 1 July to 24 August 2026. On 10 July, the day after it says GPT-5.6 models were released, the share of Shopping recommendations retrieved from ChatGPT's feed-integrated catalog rather than web search rose from 8.26% to 61.54%.1

The effects spread across merchants. Comparing 7–9 July with 10–12 July in the same data, the top ten merchants' share rose from 22.5% to 41.8%, and the number of unique merchants referenced fell from 13,524 to 10,607.1 A merchant without a product feed could have lost visibility on 10 July while changing nothing on its site.

That is the pattern of an engine change: a step on a known date, felt across many brands, in one direction. Our note on the feed shift covers what merchants did about it. For measurement, the lesson is that the date matters as much as the size of the move.

Two engine changes, summer 2026

Large moves on known dates

8.26% → 61.54%ChatGPT Shopping retrieval from its feed catalog, 10 JulProfound, Jul 2026 1
3.8% → 0.5%Reddit's share of ChatGPT citationsSemrush, 18 Jul–17 Aug 2026 2
Sources: Profound; Semrush, reporting Promptwatch data.

Worked case two: the August Reddit drop

A sudden drop can also be partly a measurement artefact, which is why the tracker needs checking too. Semrush reported Promptwatch data showing Reddit's share of ChatGPT citations falling from 3.8% (18 July–7 August 2026) to 0.5% (14–17 August), a decline of 86%.2 Promptwatch called the size of the drop provisional and said it could not rule out a problem in its own data collection.2

Notice the windows. The "after" period covers four days, against three weeks before. A short window can catch a real change early, but it can also catch a temporary one or a collection fault. Our note on the Reddit drop looked at what brands relying on Reddit threads should watch.

Our view: before acting on a sudden drop, confirm it with a second source or a second collection method, and wait for an "after" window as long as the "before" one. A brand that rewrote its community strategy on four days of data would have been acting on a number its own source questioned.

The method: five steps

The method has five steps, and the first two must happen before the update you want to measure. That is the main reason to run AI visibility tracking continuously rather than on demand.

Freeze the prompt set

Fix the prompts, give the set a version number, and do not edit it during the comparison. Adding prompts after an update mixes a change in what you measure with a change in the engine. If the set must change, run old and new versions side by side for a period. Our guide to building a prompt set from real buyer questions covers how to choose them.

Compare like-for-like windows

Use equal windows on either side of the release date, with the same number of runs per prompt, the same days of the week and the same engines. Leave a gap of a few days around the release, because rollouts are gradual. Avoid windows that straddle a holiday or a sale for one side only.

Separate the engine's changes from yours

Keep a dated change log of your own work: pages published, prices changed, feeds fixed, PR coverage. Then track three or four competitors on the same prompts. If everyone moved together on the release date, the engine changed. If only you moved, look at your log first.

Check the tracker

Look for gaps in collection, changes in how many answers came back, and changes in the tracker's own method. A fall in citations alongside a fall in total answers collected is a warning sign, not a finding. Studies can also disagree for method reasons alone, as our note on a new ChatGPT referral dataset shows.

Report with intervals

Report the change with an interval, not a single number. Answers to the same prompt are correlated, so the interval should be clustered by prompt rather than treating every run as independent. MaxAEO's July 2026 piece on sample size and clustered runs explains why that matters for AI visibility.4 If the interval crosses zero, say so plainly.

A before-and-after window around a release

  1. Frozen prompt set, competitors included, daily runs: the baseline.

  2. Logged from the engine's announcement; a few days left out either side.

  3. Same prompts, engines and run count: the comparison window.

  4. Report: change for you and each competitor, with intervals and the change log.

Illustrative. Window lengths are our default; shorten them only with more runs per day.

What to measure: mentions or citations?

Measure both, but expect them to behave differently. Otterly's September 2026 stability study, built on daily runs across several engines, found brand mentions generally more stable from day to day than the sources engines cite, and found that larger prompt sets reduced the variation it reported.3

That suggests a sensible division of labour. Use mention rate to decide whether your visibility changed. Use citations to explain how it changed, such as which sources gained or lost after the update. Our note on mentions and citations as KPIs goes further.

How we set this up

At Sigzen AI, our AI answer monitor runs a fixed, versioned prompt set across six engines: ChatGPT, Gemini, Perplexity, Claude, Copilot and Google AI. It uses official APIs where engines offer them and a licensed AI-answer data provider otherwise. Audits run each prompt on three separate days, and ongoing monitoring keeps the same set running so a baseline already exists when an engine changes.

The score itself follows the specification in our buyer's spec for an AI visibility score: named prompts, named engines, stated run counts and intervals. Our guide to how often to rerun checks covers cadence between updates. A score that cannot show its prompts, runs and interval cannot tell you whether a model update moved it.

Sources

  1. Profound, ChatGPT Shopping feed retrieval: 687 customers, 1 Jul–24 Aug 2026 (Sep 2026)
  2. Semrush, Reddit's citations in ChatGPT: tracker data, 18 Jul–17 Aug 2026 (Aug 2026)
  3. Otterly, AI search visibility stability study (Sep 2026)
  4. MaxAEO, AI visibility sample size and clustered runs (Jul 2026)

Questions readers ask

  • Our default is four weeks on each side, with a few days left out around the release date because rollouts are gradual. Equal length matters more than the exact number. If you must decide sooner, use shorter windows with more runs per prompt each day, and say in the report that the result is provisional.

Keep reading

Free, in two minutes. Enter your domain and we'll score it against three competitors across six engines.

No account needed. The report is emailed within 24 hours.