The missing retrieval layer
Traditional keyword tools describe people typing into search engines. AI assistants add another layer: a model can reformulate a need, issue several searches and select evidence before presenting an answer.
That layer can reveal useful concepts and qualifiers, but it is not represented by human keyword volume and should not be merged into it.
Traditional keyword research starts with people expressing demand in a search engine. An answer system can add its own retrieval language after that expression: it may resolve entities, seek a current source, disambiguate a product or test a constraint. The intermediate search is neither the original human demand nor the final answer.
That layer is valuable because it can expose why a page fails to satisfy a complex task. A canonical may rank for the broad topic while omitting the current fact, comparison criterion or source type that an answer process needs. Query evidence gives the researcher a concrete place to investigate.
The layer is also easy to abuse. A provider observation is a single event, and an estimated branch is an experimental output. Neither supports a market-size claim. Good AEO and GEO work becomes more rigorous when it accepts smaller, correctly named evidence instead of manufacturing a universal metric.
Four evidence families, four different questions
A disciplined program names the source, unit, collection window and coverage boundary before interpreting the metric or comparing it with another evidence family.
Google Ads volume is useful for prioritizing the relative scale of human Google demand within the export's market, language, network and period. Search Console then shows whether the site actually earned impressions and clicks for query/page pairs. Together they support conventional demand and performance decisions.
Provider-query observations answer a different question: which search string did a supported AI interface expose during this event? Model estimates explore which adjacent strings are plausible under a named generation method. They can enrich a research brief before enough production evidence exists, but they need stronger caveats.
A mature report can show all four families side by side while preserving the unit. The reader should never have to inspect a footnote to discover that a bar combines monthly Google searches, local interface events and model samples.
| Evidence | Question answered | Common misuse |
|---|---|---|
| Google Ads volume | How much human Google demand is estimated? | Calling it AI-assistant demand |
| Observed AI query | What search string did the interface expose? | Calling it the complete hidden process |
| Estimated fan-out | What adjacent search is plausible under the experiment? | Counting it as observed demand |
| GSC performance | How did the site appear in Google Search? | Attributing every row to an AI feature |
Useful data preserves provenance
Every query should carry its provider, source surface, evidence class, capture time and method version. That metadata allows a later analyst to compare like with like and identify interface or model drift.
Open Queries is designed around this provenance contract before it is designed around dashboards or aggregate scores.
Provenance needs more than a source label. Record the market and export for demand, the property and date window for Search Console, the provider/interface/adapter for observations and the model/template/sample design for estimates. Add the target canonical and editorial decision so the evidence remains connected to its use.
The contract should survive transformations. Deduplication can group literal strings, but each underlying record keeps its class. A chart can aggregate observed events, but its tooltip or table exposes the provider and window. An exported content brief identifies whether a phrase came from demand, performance, observation or estimation.
Unknown values remain unknown. If a provider does not expose a query or Search Console suppresses a low-volume row, filling zero creates a false comparison. The dashboard should make incomplete coverage visible rather than cosmetically complete.
- Observed remains observed.
- Estimated remains estimated.
- Missing remains unknown rather than zero.
- Human demand and AI retrieval remain separate until interpretation.
From query evidence to a content decision
The useful output is a falsifiable improvement to one canonical page, not another dashboard score or a mandate to publish every adjacent query.
A decision begins with a canonical and reader job. Review the verified demand cluster and Search Console query/page pairs first, then add provider evidence to identify retrieval functions or qualifiers. This order prevents a handful of interesting model outputs from overriding actual market and site evidence.
Write the gap as a falsifiable statement: “The GEO canonical does not show how a team maps a volatile claim to a dated primary source.” The intervention follows naturally: add a claim-ledger workflow and worked example. “Add more GEO keywords” is neither falsifiable nor editorially useful.
Choose one intervention and one primary signal. Internal-link work should be evaluated first through discovery and relevant exposure. Evidence and answer improvements may be evaluated through query fit, bounded citation observations and engagement. Product CTAs require actual downstream events before they are called installations or activations.
- 01Review the query
Identify the task, entities, qualifiers and freshness requirements.
- 02Find the owner
Map the intent to one existing canonical or document why it is distinct.
- 03Name the gap
Specify the missing direct answer, workflow, input, output, source or limitation.
- 04Publish one intervention
Change the smallest page surface that fully resolves the gap.
- 05Wait for evidence
Use a defined observation window before making the next material rewrite.
What query data can and cannot show
Recurring observed queries can strengthen confidence that a retrieval concept matters. They cannot prove a ranking factor, market size or the provider's complete reasoning. Estimated candidates can reveal adjacent language but cannot prove a production search.
The defensible outcome is better content coverage and clearer measurement, not synthetic certainty.
Convergence raises confidence without proving causality. If demand, Search Console exposure and provider observations all reveal the same missing constraint, the case to improve the canonical is strong. A later traffic movement can still have multiple causes, including broader demand, crawling, competition and product changes.
The data also cannot decide truth. A frequent query may rest on a false premise, and a cited source may be weak. Editorial research must verify claims against the strongest appropriate source and preserve disagreements or uncertainty.
Match each evidence family to a decision
The right metric is the one whose collection process matches the question. This matrix prevents a convenient dataset from becoming the answer to every problem.
| Decision | Primary evidence | Useful secondary evidence | Do not substitute |
|---|---|---|---|
| Choose a canonical | Verified demand + intent analysis | Current query/page exposure | Model-estimated fan-out alone |
| Find retrieval gaps | Observed provider queries | Controlled estimates and source review | Google volume as provider frequency |
| Diagnose Google discovery | GSC page/query and index evidence | Crawl and internal-link checks | Manual assistant citation |
| Evaluate citation tests | Versioned prompt/provider observations | Source and passage inspection | One screenshot as share of voice |
| Report installations | Verified store or product event | Install-page visits and downloads | GitHub download as an install |
Write a one-page evidence memo before publishing
The memo is a compact audit trail. It makes the reasoning reviewable and stops a later editor from turning a cautious observation into an unsupported marketing claim.
- Decision: improve, hold, consolidate, create or reject.
- Canonical and reader job: the one page and stable intent in scope.
- Evidence: source, unit, market or interface, date window and collection version.
- Gap: the missing question, comparison, source, workflow or limitation.
- Intervention: exact sections and claims to change.
- Primary signal and guardrail: what would support the hypothesis and what must not degrade.
- Earliest review date: enough time for crawling, exposure or product data to arrive.
Primary sources
- AI features and your websiteGoogle Search Central · accessed 2026-08-10
Google's documented eligibility, query fan-out, internal-link, structured-data and Search Console guidance for AI Overviews and AI Mode.
- Creating helpful, reliable, people-first contentGoogle Search Central · accessed 2026-08-10
The people-first content and source-quality principles used in the editorial quality framework.
- GEO: Generative Engine OptimizationAggarwal et al., arXiv:2311.09735 · accessed 2026-08-10
The original GEO framing, benchmark design and the finding that optimization effects vary by domain.
- Open Queries methodologyOpen Queries · accessed 2026-08-10
The published distinction between observed and estimated queries, provider-native estimation methods and reporting limitations.