Skip to main content

Senest opdateret

Source mapping: finding the domains AI cites in your category

An answer engine does not draw on the whole web when it answers a question. It retrieves a few pages, attributes the answer to them, and stops. Ask enough questions in one category and the same domains keep coming back. Source mapping is the work of finding out which domains those are in your market, and then checking whether you appear on them. It replaces a guess about where authority comes from with a list of named sites and a status next to each one.

The premise

Citations concentrate. In a Pew Research Center analysis of 68,879 Google searches recorded from 900 US adults in March 2025, the most frequently cited sources in AI summaries were Wikipedia, YouTube and Reddit, which together accounted for 15 percent of the sources listed. The same analysis found that 88 percent of summaries cited three or more sources.

That measurement covers general search behavior across all topics. Your category has its own version of it, and it is usually short: one trade publication, two or three review destinations, an active forum, a reference entry, a marketplace listing. The set is small enough to write down. Your presence or absence on each item is a fact you can check, not a matter of opinion. That is what makes the exercise worth running.

The workflow

  1. Run your prompt set and record every cited URL. Use a fixed set of buyer questions, across phrasings and engines, built as described in building a prompt set. For each answer, record the date, the engine, the prompt, every URL the answer cited and the domain of each URL.
  2. Aggregate by domain. Count the number of distinct answers in which each domain appeared, not the number of URLs. A domain cited once is noise. A domain cited in a quarter of your answers, across several runs, is a recurring source. The list of recurring sources is the map.
  3. Classify each recurring source by type. Sort them into your own site, third-party editorial and news, community and forum, reference works such as Wikipedia, marketplaces and directories, competitor-owned pages, and video. The type decides what can be done about the source, so the classification is not bookkeeping.
  4. Check three things per source. Whether you appear on it at all. Whether what it says about you is accurate and current. Whether presence there is something you can legitimately influence, meaning through the route the site itself publishes, rather than by buying a placement or seeding content.

The output is one row per recurring domain with four columns: type, present, accurate, influenceable. None of this requires a tool. It requires writing down what you saw.

The gap list, in priority order

  1. Sources where you are absent but eligible. The domain is cited in your category, it has a route you can legitimately use, and you are not on it. These are the only gaps that convert straight into a task, so they go first.
  2. Sources where you are present but described wrongly. Old pricing, a discontinued product, a wrong category, a former owner. A wrong description that gets retrieved works against you, and correcting it is usually cheaper than earning a new placement. Most directories, marketplaces and publications have a correction or claim process.
  3. Sources you cannot influence. Competitor-owned comparison pages, and coverage you have no legitimate claim on. Record them, then stop spending on them. Buying your way into that group is the pattern described in brand mentions and share of voice, and it does not survive contact with how these systems weigh sources.

What each source type actually requires

Third-party editorial and news. Coverage requires something newsworthy. There is no submission form and no fee that produces an honest result. What earns it is covered in digital PR and how third-party citations get earned.

Community and forum. Presence requires participation under a stable identity over time. Large communities treat promotional posting as spam and remove it, so the only durable route is being useful in threads you would have answered anyway. See Reddit, forums and UGC in AI answers.

Reference works. These run on inclusion rules you do not control. Wikipedia's notability guideline asks for significant coverage in reliable sources that are independent of the subject, and its conflict of interest guideline strongly discourages editors with a conflict from editing affected articles directly, with paid editing subject to disclosure. A reference entry is therefore a downstream effect of independent coverage rather than a task you can schedule. How engines resolve your organization as an entity is covered in entities and how engines identify you.

Marketplaces and directories. These require an accurate current listing and little else. Claim the listing, correct the facts, and keep the name, category and description consistent with your own site. This is the cheapest part of any gap list to close, and it is usually the part that has been neglected longest.

Video. A video platform appears in a citation set because it holds content on the question, not because a brand is active there. Presence means having published something that answers the question, which is a content decision rather than an outreach one.

Measurement discipline

A source map is a measurement and inherits every problem measurement has. Research across 815,000 prompt-page pairs found that after the same prompt was run three times in ChatGPT, only 2.2 percent of citations remained. A single run tells you almost nothing about which domains matter in your category. Record the date and engine for every run, repeat on a monthly cadence, and compare aggregated months rather than individual answers. The sampling requirements are set out in measuring AI search visibility honestly, and what is worth reporting from the results is in the metrics worth tracking.

The limits of a source map

Three limits are worth stating before anyone treats the list as fixed. Citation sets move, because the underlying index and the retrieved pages change, so a map is valid for the month it was made. Engines differ in what they retrieve and cite, so a map built on one engine describes that engine, as covered in how the engines compare. And the list is a sample of what your prompt set surfaced, not a complete map of the category: widen the prompt set and new domains appear. Treat the map as a working hypothesis that gets more accurate each month, and it will do its job, which is to tell you where the next piece of work is.

Sæt det på autopilot

Vupie anvender alt i denne guide, otte premium-artikler om måneden.

Kom med i betaen