The missing retrieval layer

Traditional keyword tools describe people typing into search engines. AI assistants add another layer: a model can reformulate a need, issue several searches and select evidence before presenting an answer.

That layer can reveal useful concepts and qualifiers, but it is not represented by human keyword volume and should not be merged into it.

Traditional keyword research starts with people expressing demand in a search engine. An answer system can add its own retrieval language after that expression: it may resolve entities, seek a current source, disambiguate a product or test a constraint. The intermediate search is neither the original human demand nor the final answer.

That layer is valuable because it can expose why a page fails to satisfy a complex task. A canonical may rank for the broad topic while omitting the current fact, comparison criterion or source type that an answer process needs. Query evidence gives the researcher a concrete place to investigate.

The layer is also easy to abuse. A provider observation is a single event, and an estimated branch is an experimental output. Neither supports a market-size claim. Good AEO and GEO work becomes more rigorous when it accepts smaller, correctly named evidence instead of manufacturing a universal metric.

Four evidence families, four different questions

A disciplined program names the source, unit, collection window and coverage boundary before interpreting the metric or comparing it with another evidence family.

Google Ads volume is useful for prioritizing the relative scale of human Google demand within the export's market, language, network and period. Search Console then shows whether the site actually earned impressions and clicks for query/page pairs. Together they support conventional demand and performance decisions.

Provider-query observations answer a different question: which search string did a supported AI interface expose during this event? Model estimates explore which adjacent strings are plausible under a named generation method. They can enrich a research brief before enough production evidence exists, but they need stronger caveats.

A mature report can show all four families side by side while preserving the unit. The reader should never have to inspect a footnote to discover that a bar combines monthly Google searches, local interface events and model samples.

EvidenceQuestion answeredCommon misuse
Google Ads volumeHow much human Google demand is estimated?Calling it AI-assistant demand
Observed AI queryWhat search string did the interface expose?Calling it the complete hidden process
Estimated fan-outWhat adjacent search is plausible under the experiment?Counting it as observed demand
GSC performanceHow did the site appear in Google Search?Attributing every row to an AI feature

Useful data preserves provenance

Every query should carry its provider, source surface, evidence class, capture time and method version. That metadata allows a later analyst to compare like with like and identify interface or model drift.

Open Queries is designed around this provenance contract before it is designed around dashboards or aggregate scores.

Provenance needs more than a source label. Record the market and export for demand, the property and date window for Search Console, the provider/interface/adapter for observations and the model/template/sample design for estimates. Add the target canonical and editorial decision so the evidence remains connected to its use.

The contract should survive transformations. Deduplication can group literal strings, but each underlying record keeps its class. A chart can aggregate observed events, but its tooltip or table exposes the provider and window. An exported content brief identifies whether a phrase came from demand, performance, observation or estimation.

Unknown values remain unknown. If a provider does not expose a query or Search Console suppresses a low-volume row, filling zero creates a false comparison. The dashboard should make incomplete coverage visible rather than cosmetically complete.

  • Observed remains observed.
  • Estimated remains estimated.
  • Missing remains unknown rather than zero.
  • Human demand and AI retrieval remain separate until interpretation.

From query evidence to a content decision

The useful output is a falsifiable improvement to one canonical page, not another dashboard score or a mandate to publish every adjacent query.

A decision begins with a canonical and reader job. Review the verified demand cluster and Search Console query/page pairs first, then add provider evidence to identify retrieval functions or qualifiers. This order prevents a handful of interesting model outputs from overriding actual market and site evidence.

Write the gap as a falsifiable statement: “The GEO canonical does not show how a team maps a volatile claim to a dated primary source.” The intervention follows naturally: add a claim-ledger workflow and worked example. “Add more GEO keywords” is neither falsifiable nor editorially useful.

Choose one intervention and one primary signal. Internal-link work should be evaluated first through discovery and relevant exposure. Evidence and answer improvements may be evaluated through query fit, bounded citation observations and engagement. Product CTAs require actual downstream events before they are called installations or activations.

  1. 01
    Review the query

    Identify the task, entities, qualifiers and freshness requirements.

  2. 02
    Find the owner

    Map the intent to one existing canonical or document why it is distinct.

  3. 03
    Name the gap

    Specify the missing direct answer, workflow, input, output, source or limitation.

  4. 04
    Publish one intervention

    Change the smallest page surface that fully resolves the gap.

  5. 05
    Wait for evidence

    Use a defined observation window before making the next material rewrite.

What query data can and cannot show

Recurring observed queries can strengthen confidence that a retrieval concept matters. They cannot prove a ranking factor, market size or the provider's complete reasoning. Estimated candidates can reveal adjacent language but cannot prove a production search.

The defensible outcome is better content coverage and clearer measurement, not synthetic certainty.

Convergence raises confidence without proving causality. If demand, Search Console exposure and provider observations all reveal the same missing constraint, the case to improve the canonical is strong. A later traffic movement can still have multiple causes, including broader demand, crawling, competition and product changes.

The data also cannot decide truth. A frequent query may rest on a false premise, and a cited source may be weak. Editorial research must verify claims against the strongest appropriate source and preserve disagreements or uncertainty.

Match each evidence family to a decision

The right metric is the one whose collection process matches the question. This matrix prevents a convenient dataset from becoming the answer to every problem.

DecisionPrimary evidenceUseful secondary evidenceDo not substitute
Choose a canonicalVerified demand + intent analysisCurrent query/page exposureModel-estimated fan-out alone
Find retrieval gapsObserved provider queriesControlled estimates and source reviewGoogle volume as provider frequency
Diagnose Google discoveryGSC page/query and index evidenceCrawl and internal-link checksManual assistant citation
Evaluate citation testsVersioned prompt/provider observationsSource and passage inspectionOne screenshot as share of voice
Report installationsVerified store or product eventInstall-page visits and downloadsGitHub download as an install

Write a one-page evidence memo before publishing

The memo is a compact audit trail. It makes the reasoning reviewable and stops a later editor from turning a cautious observation into an unsupported marketing claim.

  • Decision: improve, hold, consolidate, create or reject.
  • Canonical and reader job: the one page and stable intent in scope.
  • Evidence: source, unit, market or interface, date window and collection version.
  • Gap: the missing question, comparison, source, workflow or limitation.
  • Intervention: exact sections and claims to change.
  • Primary signal and guardrail: what would support the hypothesis and what must not degrade.
  • Earliest review date: enough time for crawling, exposure or product data to arrive.

Primary sources

  • AI features and your website

    Google's documented eligibility, query fan-out, internal-link, structured-data and Search Console guidance for AI Overviews and AI Mode.

    Google Search Central · accessed 2026-08-10
  • Creating helpful, reliable, people-first content

    The people-first content and source-quality principles used in the editorial quality framework.

    Google Search Central · accessed 2026-08-10
  • GEO: Generative Engine Optimization

    The original GEO framing, benchmark design and the finding that optimization effects vary by domain.

    Aggarwal et al., arXiv:2311.09735 · accessed 2026-08-10
  • Open Queries methodology

    The published distinction between observed and estimated queries, provider-native estimation methods and reporting limitations.

    Open Queries · accessed 2026-08-10