AI Visibility: What It Is, How to Measure It, and What Moves It
AI visibility is whether assistants mention your brand when buyers ask. This defines the term precisely, sets out the four metrics worth tracking, and explains which ones you can actually influence.
Entities in this article
Ask ChatGPT to recommend a sauna heater for a small room, or Perplexity which agency handles eCommerce SEO, and you get an answer with a handful of brands named in it. Being one of those brands is what people now mean by AI visibility.
It is a genuinely new measurement problem, and most of the confusion around it comes from borrowing vocabulary that does not fit. You cannot rank in an answer. There is no position three. A brand is either present in the response or it is not, and the same question asked twice can produce different sets of names.
This piece defines the term, sets out the four things actually worth measuring, and — the part most coverage skips — separates the metrics you can influence from the ones you can only observe.
A Working Definition
AI visibility is the likelihood that an AI assistant names your brand, unprompted, in response to a question your buyers actually ask.
Three parts of that sentence carry weight.
Unprompted matters because asking “what do you think of Nordica Marketing?” tests recall, not visibility. The model already has the name. Real visibility is measured on questions that do not contain your brand — “best ecommerce SEO agency for Shopify stores” — where the model must choose whom to name.
Likelihood matters because these systems are not deterministic. The same prompt run ten times may name you six times. A single check tells you almost nothing; visibility is a rate, and anyone reporting it as a yes/no is measuring noise.
Questions your buyers actually ask matters because visibility on irrelevant questions is worthless. Being named in “what is generative engine optimization” is nice. Being named in “who should handle SEO for a £5m DTC brand” is revenue.
The Four Metrics Worth Tracking
Most tools report a single composite “visibility score.” Treat those with suspicion — a number that blends four different things tells you something changed but never what, and you cannot act on it.
1. Presence Rate
Of the prompts in your tracked set, what percentage name your brand at all?
This is the base metric and the one to establish first. Run each prompt several times, because a single run conflates a genuine absence with ordinary variance. A presence rate is only meaningful with a stated sample size — “42% across 50 prompts × 5 runs” is a measurement; “we’re visible in AI” is not.
2. Share of Voice
When your brand is named, who else is named alongside you, and how often does each competitor appear across the same prompt set?
This is the more actionable half. A 40% presence rate reads very differently if the leader is at 45% than if it is at 95%. Share of voice tells you whether you are competing or absent from the conversation, and it names the specific competitors the model reaches for — which is a content brief in itself.
3. Citation Rate
When you are named, does the response link to your site?
Presence without citation still builds brand awareness, but citation is what produces traffic and what compounds: a cited page is a page the system has learned to treat as a source. Track which of your URLs get cited, because it is rarely the ones you would guess. Deep, specific, well-structured pages get cited far more often than homepages.
4. Sentiment and Framing
Being named is not automatically good. A model that names you as “a budget option” or “better suited to smaller catalogues” has positioned you, and that framing came from somewhere — usually a review site, a comparison page, or your own copy.
This is the metric teams most often skip and most often regret skipping.
What You Can Actually Influence
Here is the distinction that matters commercially, and where most AI-visibility advice becomes hand-waving.
Directly influenceable:
- What your own pages say and how clearly they say it. Structure, extractability, specificity. Covered in GEO versus SEO.
- Factual consistency across your footprint. Where your site, your feed, your listings, and your review profiles disagree, models hedge or drop the claim. This is the cheapest fix available and almost nobody does it.
- Coverage of the questions being asked. If a prompt returns competitors, and you have no page addressing that question, the absence is self-inflicted.
- Third-party presence. Models draw heavily on comparison sites, roundups, and editorial coverage. Earning those placements is slow, and it works.
Observable but not directly controllable:
- Which sources a given model favours. These weightings change without notice.
- Run-to-run variance. Manageable by sampling, not eliminable.
- Training-data recency. Some models answer from a fixed snapshot; nothing you publish today affects them until the next update.
- Whether a query triggers retrieval at all. Some answers come from parameters alone, with no live sources consulted.
The practical consequence: if a competitor dominates a prompt because it has ten years of editorial coverage you do not have, no amount of on-page work closes that this quarter. Choose prompts where the gap is closeable, and treat the rest as a long game.
How Measurement Actually Works
There is no API that returns “your AI visibility is 34%.” Every tool in this space, and every credible in-house process, does the same thing:
- Define a prompt set — the questions your buyers ask, in their words, without your brand in them.
- Run them repeatedly across the assistants your market uses.
- Parse the responses for brand mentions, competitor mentions, and links.
- Aggregate into rates over time.
The methodology is straightforward. The judgement is in step one, and it is where most implementations fail. A prompt set assembled from keyword-research exports produces search queries, not questions — and people do not talk to assistants the way they type into Google. “ecommerce seo agency uk” is a search. “Which agency should I hire to fix my Shopify store’s organic traffic?” is a prompt. They return different brands.
If you are evaluating tooling for this, we compared what the main options actually surface in the guide to AI search visibility tools. For the narrower question of running a one-off check, see how to run an AI visibility check, and for the ongoing discipline, AI search monitoring.
What Good Looks Like
A defensible AI visibility programme has:
- A prompt set of 30–100 questions, written as questions, refreshed quarterly.
- A stated sample size — at least three runs per prompt, ideally five.
- Competitor tracking on the same prompts. Your own numbers are close to meaningless without them.
- A baseline recorded before any changes, so improvement is demonstrable rather than asserted.
- A named owner. Prompt sets rot. Buyer questions change. Nobody notices unless it is somebody’s job.
That is a modest amount of work — days, not months — and it is the difference between knowing your position and guessing at it.
The Honest Caveats
The numbers are noisy. Presence rates move several points run to run for reasons that have nothing to do with you. Read trends over weeks, never single measurements, and be sceptical of any tool reporting to a decimal place.
Comparability across tools is poor. Two tools measuring the same brand on the same day will disagree, because they use different prompt sets, different models, and different mention-matching rules. Pick one and stay with it; switching resets your history.
Correlation with revenue is not yet established. AI-referred traffic is small for most businesses today and hard to attribute — assistants often strip referrer data. The case for investing now is that the trend line is steep and the work compounds, not that there is a proven near-term ROI. Anyone telling you otherwise is selling something.
Where to Start
If you have measured nothing so far, the first useful week looks like this: write 30 real buyer questions, run each five times across the two assistants your market uses, and record presence, competitors, and citations in a spreadsheet. That baseline costs a day and tells you more than any vendor dashboard will before you know what you are looking at.
Then fix consistency, because it is cheap. Then close coverage gaps on the prompts where competitors appear and you do not.
If you want that baseline built properly — real prompt set, competitor benchmarking, and a prioritised gap list rather than a score — that is where our AI search optimization engagements begin. Book a free strategy call and we will show you what your current position looks like.