Skip to main content

Last updated

Designing a prompt set you can measure against

A prompt set is a fixed, written-down list of questions you re-run against answer engines on a schedule. It is the AI equivalent of a tracked keyword list. The questions are the measurement instrument, and holding them constant is what makes this month's result comparable to last month's. Measuring AI search visibility honestly establishes why a single check is noise and what a defensible sample requires. This page covers the part that guidance leaves open: how to build the set.

The design axes

A prompt set is not a keyword list with question marks added. Each entry is a full question, and each question carries several choices that have to be made deliberately, because every one of them changes the answer that comes back.

AxisValuesWhat varying it buys you
Intent stageProblem-aware, solution-comparison, brand-specificWhere in the buying process you appear. Problem-aware questions are where discovery happens; brand-specific ones are where a reader checks you out.
ScopeCategory, use case, comparison, brandCategory questions are crowded and volatile. Use-case questions are narrower and often where a smaller brand can actually be named.
BrandedWith or without your name in the questionTwo different measurements, described below.
LanguageOne line per market languageThe largest single source of difference in what the engine says.
EngineEach engine you care aboutEngines retrieve and cite differently, so a result on one does not transfer to another.
PhrasingSeveral wordings of the same underlying questionAverages out wording effects instead of letting them read as change.

The branded and non-branded split deserves its own emphasis, because the two answer different questions. A non-branded prompt tests whether you get discovered: your name has to arrive in the answer because the engine chose to put it there. A branded prompt puts your name in the question, so the engine will almost always mention you. What varies there is what it says, which sources it leans on and what it gets wrong. The first measures reach. The second measures the account of you that is circulating. Combined into one score, both become unreadable.

Where the prompts come from

The temptation is to write prompts your product wins. That produces a chart that rises and a set that measures nothing. Prompts have to be sourced from demand that already exists:

  • Your own Search Console queries. Real questions, in real wording, from people who already found you. The same discipline as keyword research applies here.
  • Sales and support logs. Questions asked out loud before and after a purchase, including the awkward ones about price, migration and competitors.
  • Customer conversations. The phrasing buyers use before they know your category's vocabulary.
  • Competitor and category language. The words the market uses to describe the problem, which are frequently not the words on your homepage.

The working rule: if you cannot point to a place where a human actually asked the question, it does not go in the set.

How many prompts, and how many repeats

Your budget is a fixed number of answers, and it splits four ways: distinct prompts, times languages, times engines, times repeats per prompt. Repeats reduce noise on questions you already have. Prompts add questions you do not have. Because answers vary run to run, a wider set measured less often beats a narrow set measured daily. A narrow set measured daily gives you a precise estimate of a small corner of demand and a trend line made mostly of sampling.

Language deserves budget before repetition does. Analysis of variance in LLM brand responses attributes 26.5 to 32.0 percent of the variance to the language of the query and 1.5 percent to brand identity. If you sell in more than one language, the next unit of budget buys more information spent on a second language than on another repeat in the first. What that implies for multi-market brands is set out in multilingual GEO.

There is no correct number of prompts, and any specific figure quoted as one is invented. The number you need is the one that makes your confidence interval narrow enough to support the decision you intend to make. Compute it from your own runs rather than adopting someone else's, using the method in measuring AI search visibility honestly.

Freeze the set, then version it

Changing prompts mid-quarter destroys comparability. Once the instrument moves, movement in the results mixes real change with instrument change and the two cannot be separated afterwards. Freeze the set for at least one full measurement cycle.

Sets do have to change, because markets and products change. Treat each change as a new version rather than an edit. Keep a dated changelog recording what was added, removed or reworded, and why. Where the change is large, run the old and new versions in parallel for one cycle so you have an overlap to calibrate against. Every number you report should say which version of the set produced it.

What to record on every run

  • Engine and model version.
  • Language.
  • Date and time.
  • The exact prompt text, including which phrasing variant it was.
  • Whether the brand was named in the answer text.
  • Whether your domain was cited as a source.
  • Every domain cited, not only yours.
  • The verbatim answer text, wherever the interface allows you to keep it.

The verbatim text is the field most often skipped and the one that pays back. Aggregates can be recomputed later from stored answers; stored answers cannot be recovered from aggregates. Recording naming and citing as two separate fields matters for the same reason: the gap between them is a metric in its own right, defined in the GEO metrics.

A sample, not a census

A prompt set is a sample of the questions your market asks, selected by you and weighted by your judgment about what matters. It estimates visibility. It does not measure it. Two consequences follow. Report ranges rather than points, and never describe a prompt-set result as a share of anything other than that prompt set. The vocabulary those results get reported in, and the ways it gets stretched, is covered in the GEO metrics, defined.

自動運転にまかせる

Vupie はこのガイドの内容をすべて実行します。月に 8 本のプレミアム記事。

ベータに参加