There is a way to be wrong about AI visibility that looks exactly like being right. You run a prompt, the assistant names your brand, you write down 100%. Next month it names three competitors, you write down 0%, and now you have a trend line with a cliff in it. Nothing about your brand changed. You sampled a distribution twice and drew a line between the points.
What follows is our methodology for measuring AI visibility, written so a careful reader can attack it. Vendors, us included, have an incentive to publish numbers without the design behind them, and the design decides whether the number means anything.
Table of Contents
An answer is a sample, not a lookup #
A search results page is close to a lookup. Two people in the same country querying the same term minutes apart get results similar enough that “position 4” describes reality. Rank tracking was built on that stability and mostly earned it.
A generated answer does not work that way. The model produces text by sampling from a distribution over possible continuations, conditioned on your prompt, on whatever was retrieved first, and on its own configuration. Ask the same question twice and you can get two different brand lists, two orderings and two sets of cited sources. Neither is the true answer. Both are draws.
A single observation therefore carries a variance you cannot see from inside it. You get one answer, it reads like a fact, and it is a coin flip whose bias you never estimated. A visibility percentage with no sampling design behind it is an anecdote with a percent sign on it.
Where the variance comes from #
Model non-determinism. The same prompt to the same model on the same day can produce different text. Temperature and sampling settings drive this and you do not control them through a consumer assistant, so this much variation survives any measurement design.
Model and version changes. Providers ship new versions, retire old ones and rewrite system prompts, none of it announced in a form you can subscribe to. An update can move a mention rate by more than a year of content work, in either direction.
Grounding. Whether the assistant ran a live web search or answered from parametric memory changes what the rest of the answer is built on. A grounded answer usually cites sources and reflects whatever the index held at that moment; an ungrounded one reflects training data, which is older and stickier. The same assistant does both depending on phrasing, and you rarely see which happened except by inferring it from whether sources appeared.
Personalization and account memory. A logged-in user with a history of asking about running shoes gets different answers than a fresh session. Memory, custom instructions and earlier turns all condition the output, so a clean-session measurement captures one case among many.
Locale and language. The same question in Spanish and in English is not the same question. Different sources are available, different brands are prominent, training coverage differs by language, and country affects retrieval.
Index timing. For grounded answers, what the retrieval layer held at the moment it was queried matters. A news cycle can move the cited set within hours.
Prompt phrasing. “Best CRM for small business” and “what CRM should a 10 person company use” are different queries with different answer sets. Small wording changes move results more than people expect.
| Source of variance | Can you control it? |
|---|---|
| Prompt phrasing | Yes, fully. Fix it and log it. |
| Locale and language | Yes. Choose it and keep it as a dimension. |
| Account state and session history | Yes, if you measure from clean sessions. |
| Cadence and sample size | Yes. Your main lever against noise. |
| Model non-determinism | No. Average it out with repetition. |
| Model and version updates | No. Log the date and annotate. |
| Whether the answer was grounded | No. Record whether sources appeared. |
| Index churn | No. Treat as background movement. |
The design that makes a number defensible #
These are the rules a survey researcher would apply to a poll, translated into this domain.
- Fix the prompt set, and change it only with an annotation. Swapping prompts mid-series is like changing the questions in a tracking survey and calling the result a trend. If you add or remove prompts, record the date and expect a discontinuity there.
- One answer is an observation, a period of answers is a measurement. Never report a single response as a state. Pick a period, weekly for most brands, and treat the answers inside it as your unit. More prompts and more frequent runs both increase the sample and both cost money.
- Aggregate at the group level, not the single prompt. One prompt’s mention rate bounces hard. A group of 30 related prompts moves less, and moves for reasons you can investigate. Single prompt detail is for diagnosis, not reporting.
- Keep the denominator explicit. Visibility is mentions divided by answers. “We appear in 38% of answers” means nothing until it says “of 412 answers”. A rate over 12 answers and a rate over 1,200 do not belong in the same chart without it.
- Separate branded from non-branded prompts. A prompt that names your brand will mention your brand. Mixing those into a headline figure inflates it and makes competitive comparison meaningless, since every brand wins on its own name. Report non-branded prompts as the competitive number.
- Keep model and locale as dimensions, never silently averaged. A blended number across five assistants and three countries hides what you needed to know, which is that one model dropped you and the others did not. Average only when you say so.
- Compare like periods, and prefer direction over decimals. Four weeks against the previous four, not this week against the best week you ever had. Report direction and rough size: a gap between 38.4% and 37.9% sits inside the noise.
- Log the model version and the date with every observation. When a series steps, the first hypothesis should be a provider change, not a marketing win, and you can only test that if the version and timestamp sit on the raw observations.
Rules 4 and 5 are where most published visibility numbers fall apart, vendor marketing included. Which of these metrics are worth tracking at all is a separate question, covered in our piece on AI Search metrics, and the position-weighted variant has its own AI Visibility Score post.
Reading a change without fooling yourself #
Your visibility number moved. Work through these before you tell anyone why it did.
- Did the prompt set change? Additions, removals, edits, a new collection someone created on Tuesday. Check this first: it explains more sudden moves than anything else.
- Did the model change? A new version, a rollout in your country, a change in whether the assistant searches the web by default. Provider-side changes tend to move every brand in a category at once, a useful signature.
- Did coverage change? This is the one people miss. If the assistant answered 400 of your prompts last month and 340 this month, because it refused, timed out or came back empty, the rate now covers a different population. Track answered prompts as its own series.
- Is the move bigger than this group’s normal spread? See below.
- Does it hold across models and locales? A change in one assistant, or one country, is usually about that assistant or country. A change everywhere is either a category shift or something you did.
Build a noise band from your own history
Rules of thumb borrowed from someone else’s dataset are worth little here, because the spread depends on your prompt count, cadence and category. Build the band yourself.
Take one prompt group and a stretch of history where you know nothing structural changed, ideally eight weeks or more. Look at the period-over-period changes and note how large they typically got. That range is what the group does when nothing is happening. A move inside it is not news. A move outside it is worth investigating, and still is not proof.
This is descriptive, not a significance test, and we would rather say so than dress it in statistics it does not deserve. It is still enough to stop a team celebrating a three point rise the same group produced twice last quarter for no reason. Related failure modes are in our post on AI rank tracking myths.
What we do, and where our method stops #
We run a fixed prompt set on a schedule, separately per model and per locale, and keep every answer: the text, which brands were mentioned, where in the answer they appeared, the sentiment of the mention and every source cited. Because the history stays in your account, the baseline you compare against is your own. The limits of that design are below.
- Sampling cannot prove what an assistant says to everyone. We can tell you your brand appeared in a known share of a known number of answers under known conditions. “ChatGPT recommends you 38% of the time” is a claim nobody can support.
- Personalized and logged-in answers differ from ours. We measure a clean session. A user with memory enabled and a year of chat history is a different conditioning.
- Coverage varies by model and country. Some assistants answer fewer of a given prompt set, or answer differently in some markets, so comparing raw rates without the answered counts will mislead you.
- Some models expose no sources. Where an assistant returns text without citations we have mentions but no source attribution, and any source analysis has to say so.
- Provider-side changes move series for reasons that have nothing to do with you. That is a property of the surface we measure, and we cannot engineer it away.
The last point is measurable in its own right. We publish Tremor, a public tracker of how much AI answers change on their own week to week, so nobody has to take our word for how unstable the background is. Source measurement has its own caveats, covered in the guide to tracking which sources AI answers cite. Google’s reporting has a comparable limit: the Search Console generative AI performance report gives impressions only, with no clicks, no click-through rate and no position, for AI Overviews and AI Mode.
Reporting uncertainty without drowning the room #
A CMO does not want a confidence interval. They want to know whether the thing is working. You can be honest inside that.
Report ranges and directions. “Roughly a third of answers in this category, up from a quarter over two months” is true and useful. “37.2%, up 4.1 points” implies a precision the instrument does not have.
Show the denominator on the slide: one line giving the count of answers, the models and the period. That line prevents the arguments where two people compare numbers built on different populations.
Annotate known events on the chart: model updates, prompt set changes, a campaign launch, a site migration. When the line steps a month later, the annotation is already there.
And refuse to report a number you cannot reproduce. If you cannot say which prompts, which models, which period and how many answers produced it, it does not go in the deck. You will lose the occasional good-looking slide and spend fewer meetings defending something indefensible.
Six questions for anyone else’s study #
The same discipline applies to research you read, ours included. When a vendor publishes a study about which brands dominate AI answers:
- How many answers? Not prompts, answers. Two hundred prompts run once is 200 observations; the same 200 run weekly for three months is roughly 2,600.
- Over what period? A single week cannot tell a trend from a fluctuation, however many prompts it used.
- Which models, and which versions? Results from one assistant do not generalize to the others, and a version since replaced describes a system that no longer exists.
- Which locales and languages? A US English study is a US English study.
- What was in the prompt set, and how was it built? Prompts harvested from real query data behave differently from prompts a model invented.
- Was the prompt set disclosed? If not, the study is unreproducible. Not disqualifying, since prompt sets can be commercially sensitive, but you are trusting the authors rather than checking them.
A study that answers all six can still be wrong. One that answers none cannot be evaluated at all. More of this reasoning is in our piece on the uncomfortable truths about AI search.
FAQ #
How many answers do I need before a visibility number is stable?
There is no universal threshold: it depends on how contested your category is and how often the assistants disagree. Build the noise band above for your own groups and see how large the sample has to get before period-to-period swings settle. Grouping prompts usually does more for stability than adding them.
Should I run prompts daily or weekly?
Weekly is the right default for most brands: a usable sample per period without paying for observations that re-measure the same state. Daily earns its cost during an active test, a launch or a crisis. Monthly is too coarse unless the prompt set is large.
Is a mention in an AI answer worth anything if nobody clicks?
It is worth something different from a click and should be measured separately rather than converted into a traffic estimate. Assistants increasingly answer without sending a visit, which is why Google’s generative AI reporting exposes impressions and no clicks. Treat mention rate as a presence metric and keep referred visits as their own series.
Why do two tools report different visibility numbers for the same brand?
Almost always because they measure different populations: prompt sets, models, locales, treatment of branded prompts, handling of refused prompts, sample sizes. Before deciding one tool is wrong, compare the four things that define the population: prompts, models, locales and period.
Can I just ask ChatGPT myself and see?
You can, and it is a reasonable way to read how your brand gets described and which sources the assistant leaned on. What it cannot do is produce a measurement. Your account carries memory and history, you will run the prompts you expect to do well on, and one session is one draw. Manual checks are for reading; a fixed repeated sample is for reporting.