Skip to main content

Last updated

The GEO metrics, defined

The words on this page appear across AI visibility dashboards, reports and sales decks. None of them is standardized. There is no governing body, no shared specification and no agreed formula, so two products can print the same word and compute it from different data in different ways. That is the current state of the field, and it is the first thing to know about any number carrying one of these labels.

A metric here is only meaningful with three things attached: the prompt set it was measured on, the sample size behind it, and the date. Without those, the number describes somebody's method, not your brand.

The vocabulary, term by term

Visibility score, or presence rate

The share of prompts in a set where the brand is named at all. Most tracking products report some version of this as their headline figure. It is typically computed as prompts where the brand appeared divided by prompts run, but the details diverge: some average across repeats, some count a prompt as a hit if the brand appeared in any single run, and some weight prompts by an importance score. Those three methods return different numbers from identical raw data. The score cannot tell you whether being named helped, whether anyone read the answer, or whether the prompts reflect real demand.

Detection rate, or mention rate

Usually the same idea measured at the level of individual answers rather than prompts: the share of all generated answers in which the brand name appears in the text. Products differ on whether a mention counts once per answer or once per occurrence, and on what counts as the brand at all: product names, legal entity names, parent company names and common misspellings may or may not be included. A mention rate says nothing about how you were described. Appearing in a list of options to avoid registers exactly like a recommendation.

Citation rate, or share of citations

This one is about links, not text. A citation is a source the answer attributes part of itself to, usually rendered as a footnote or linked domain. Citation rate is normally the share of answers that cite your domain. Share of citations is your domain's portion of all citation slots across the set, the harder number, because answers cite several sources at once. In Pew Research Center's analysis of 68,879 Google searches by 900 US adults during March 2025, 18 percent of searches produced an AI summary, and 88 percent of those summaries cited three or more sources. Being cited is not exclusive. Neither figure tells you anything about traffic, since a citation is not a click.

Share of voice

Your brand's mentions relative to competitors' mentions across the same prompt set. The concept, and why it holds up only as an average over many runs, is covered in brand mentions and share of voice. The detail to add here is that the competitor list is a configuration choice. Adding or removing a competitor changes your share without anything changing in the world.

Answer position, or prominence

Where in the answer your brand appears: first paragraph or last, first in a list or fifth. It is computed from character offset, sentence index or list ordinal. The metric is weaker than the name suggests, because generated text has no fixed slots. There is no position one to occupy. The same prompt run twice produces answers of different lengths and structures, so an ordinal describes one generation rather than a standing. Averaged over many runs it is a coarse signal. It is not a rank.

Sentiment

Whether the answer describes the brand favorably, neutrally or unfavorably. Scoring is normally done by a second model reading the first model's output, so the classifier is itself a variable. Covered separately in sentiment in AI answers.

Mention-to-citation ratio

Mentions divided by citations, or simply the gap between the two rates. The common pattern is named often and linked rarely, and it reads specifically: engines carry your brand in their account of the category, drawn from training data or from what other sites say about you, but your own pages are not what they lean on. The inverse, cited but seldom named, means your pages work as reference material in a category where you are not treated as a participant. The ratio tells you which situation you are in, not which one is fixable.

Source share, or domain share

Not a measurement of you. It is the distribution of third-party domains an engine cites when answering your category's questions, which shows whose material the answers are actually built from. Treated on its own page, source mapping.

Prompt coverage

Used for two different things, worth checking before trusting it. Sometimes it means the share of prompts in the set that returned a usable answer, a data quality figure. Sometimes it means how much of your addressable question space the set represents, which estimates your own judgment rather than measuring anything. The two share a name and nothing else.

The comparability trap

Two products both report 24 percent visibility for the same brand in the same month. That agreement is a coincidence. They are near certainly running different prompt sets, on different engines, with different repeat counts, over different date ranges, under different rules about what counts as a mention. When two such numbers differ, neither product is necessarily wrong.

The practical rule is that cross-tool comparison is invalid and only within-tool trend carries information: the same set, the same engines, the same method, run again later. Even that holds only when the sample is large enough for the movement to exceed the interval around it. And no product is working from privileged data, since Google states that no third-party tool has access to its internal ranking or AI systems. Every external number is sampled from outside the engine, whatever the interface implies.

What to demand from any number

Before accepting a metric from any source, ask for six things:

  • The prompt set: how many prompts, what they are, and who wrote them.
  • The engines, and the model versions where they are disclosed.
  • The languages.
  • The sample size: repeats per prompt, and total answers behind the figure.
  • The date range the answers were collected in.
  • The confidence interval around the figure.

A number that cannot produce those six is a reading, not a measurement. What defensible sampling requires is set out in measuring AI search visibility honestly, and how to build the prompt set every one of these metrics depends on is in designing a prompt set you can measure against.

Passez en pilote automatique

Vupie applique tout ce guide, huit articles premium par mois.

Rejoindre la bêta