AI visibility · Measurement

AI search visibility starts with retrieval evidence

Measure AI and LLM visibility with a defensible evidence stack: observed queries, estimates, citations, referrals and independent search demand.

12 min practical guide · Updated 2026-08-10

A defensible AI visibility stack

No single number captures the whole path from an information need to a business outcome. Retrieval queries, final citations, referrals and conversions describe different stages and should remain separate.

The layers form a dependency chain, not a funnel with guaranteed conversion. A page can be crawlable but never retrieved; retrieved but not cited; cited without a referral; or visited without producing a business outcome. Measuring each layer separately tells the team where the uncertainty begins.

Record the observation unit and denominator. “Three citations” is meaningless without the tested prompts, provider, mode, location, date and repeat design. “Twenty crawler hits” is not twenty users. “AI traffic” can mix assistant referrals, bots and ordinary campaigns unless the referrer and user-agent rules are explicit.

This layered model also prevents vanity wins. A visibility increase is useful only when the surfaced page and query match the intended job. Exposure for an ambiguous or irrelevant query should remain visible in the raw evidence but should not inflate the target canonical's success claim.

LayerQuestionExample evidence
RetrievalWhat did the surface search for?Observed provider query
ExplorationWhat adjacent paths are plausible?Provider-native fan-out estimate
SelectionWas the source shown or cited?Citation or search impression
OutcomeDid qualified attention arrive?Referral, task completion or conversion

What an AI visibility tool should make explicit

An AI visibility tool is useful only when it names the surfaces, prompts, locations, time window and evidence class behind its result. A score without provenance can hide model drift, sampling choices and missing coverage.

Open Queries does not claim full prompt monitoring, market share or a universal LLM rank. It exposes observed retrieval strings and clearly labeled estimates that can feed a broader visibility workflow.

Open Queries contributes one narrow but valuable input: the search strings explicitly exposed by supported provider interfaces and separately labeled model estimates requested by the user. That evidence can reveal how a broad task was narrowed into entities, qualifiers or source-seeking language.

It does not continuously run a prompt panel, calculate share of voice, identify every citation across providers or infer unseen retrieval. A team that needs those capabilities should combine purpose-built monitoring with Search Console, referral analytics, product analytics and manual source review. Calling a query inspector a complete AI visibility tool would make the measurement architecture less honest, not more complete.

The distinction matters operationally. Query evidence helps improve a content brief; citation evidence tests whether a source appeared; referral evidence shows a visit; and conversion evidence shows a defined product action. Each instrument should be chosen for the question it can actually answer.

A practical AI search monitoring workflow

Monitoring should start with a stable set of questions and evidence definitions, then compare like-for-like observations over time.

Define a versioned prompt or task set only when repeated answer testing is justified. Keep brand, generic category, comparison and problem-solving prompts in separate panels so a movement in branded recall does not masquerade as non-brand discovery. Record provider mode and location when they can affect results.

For every content intervention, attach one primary success signal and one guardrail. An internal-link fix may target discovery while guarding against accidental canonical changes. A source-quality rewrite may target relevant impressions or bounded citation appearance while guarding against lower install-page engagement. More metrics do not create more certainty if the hypothesis is vague.

  1. 01
    Define

    Choose the providers, tasks, markets and evidence classes being monitored.

  2. 02
    Observe

    Collect explicit query traces, citations and referrals without merging them.

  3. 03
    Diagnose

    Map retrieval gaps to one canonical page and a falsifiable content change.

  4. 04
    Evaluate

    Compare GSC, referral and business signals after a defined observation window.

From LLM visibility signal to action

Suppose a surfaced query repeatedly asks for privacy-preserving AI search tooling, while the relevant page discusses only query inspection. The missing answer is not another keyword page; it is a clear privacy boundary, data flow and deletion workflow on the existing canonical.

Imagine that a GEO guide is fetched by search crawlers, appears in one manually tested answer and receives no identifiable referral traffic. The correct report is three separate facts: discoverability is active, one bounded citation observation occurred and referral selection is not observed. It is not “AI visibility increased 100%.”

The next action depends on the objective. If the page is new, wait for indexing and relevant query exposure. If it already earns impressions but is not selected, inspect the direct answer, source specificity and title. If referrals arrive but readers bounce, the problem is likely page value or intent alignment rather than retrieval visibility.

Limits of AI and LLM visibility measurement

Assistant outputs vary by model, time, user context and available tools. Search Console Web data does not isolate every AI feature, and a citation does not prove a recommendation or conversion.

Keep missing data as unknown rather than zero. Report store installations only from the Store surface, GitHub downloads as downloads and D1 activity as product activity.

Provider interfaces and reporting surfaces change. Referrer headers can be absent, citations can vary between runs and Search Console does not break out every Google AI feature as a separate performance dimension. Missing evidence must remain unknown rather than being converted into zero.

A panel can support comparison only within its documented design. Changing prompts, providers, modes or locations resets comparability. Even a stable panel samples possible answers; it does not measure every answer shown to every user.

Build a dashboard that preserves evidence boundaries

Every row should identify its source, unit, window and interpretation limit. This makes the dashboard slower to fake and faster to diagnose.

SignalSource and unitSupportsDoes not support
EligibilityLive crawl, HTTP status, canonical and robots checksThe page can be fetched and is intended for indexingThat any answer system selected it
Google exposureGSC Web impressions/clicks by query and pageObserved Google Search exposure and selectionA separate AI Overview rank
Provider retrieval queryExplicit supported interface eventOne surfaced retrieval actionThe complete hidden retrieval plan
Citation panelVersioned prompt/provider/date observationAppearance in the bounded panelPopulation share of voice
Assistant referralValidated referrer and human-traffic rulesA visit attributed to the available referrerA citation or recommendation
Business outcomeDefined first-party product eventThe measured downstream actionThe channel's causal contribution without attribution design

Diagnose the first missing signal

Start at the earliest layer with missing or contradictory evidence. This avoids rewriting content when the route is broken and avoids technical busywork when the answer is simply weak.

  1. 01
    Not fetchable

    Fix status codes, robots, CDN rules, rendering and canonical output. Do not interpret downstream metrics.

  2. 02
    Fetchable but undiscovered

    Check sitemap processing and contextual inbound links; confirm that the canonical owns a visible site purpose.

  3. 03
    Indexed but irrelevant exposure

    Narrow the title, direct answer and entity context so the page's job is unambiguous.

  4. 04
    Relevant exposure without selection

    Inspect answer usefulness, trust evidence, comparison criteria and snippet-level clarity.

  5. 05
    Citation without referral

    Treat the citation as bounded appearance; review whether the answer resolves the task without a click and whether the source offers unique follow-on value.

  6. 06
    Referral with poor engagement

    Audit the landing promise, first-screen information scent, speed and match between cited passage and page experience.

Primary sources

  • AI features and your website

    Google's documented eligibility, query fan-out, internal-link, structured-data and Search Console guidance for AI Overviews and AI Mode.

    Google Search Central · accessed 2026-08-10
  • Creating helpful, reliable, people-first content

    The people-first content and source-quality principles used in the editorial quality framework.

    Google Search Central · accessed 2026-08-10
  • Open Queries methodology

    The published distinction between observed and estimated queries, provider-native estimation methods and reporting limitations.

    Open Queries · accessed 2026-08-10