Obsurfable

From Prompt to Recommendation: A Stage Model of AI Brand Visibility

Obsurfable

What makes a brand appear in an AI answer? The GEO industry often compresses the question into a page score: rewrite the FAQ, add statistics, ship schema, wait for citations.

A September 2026 paper by Benjamin Tannenbaum argues that framing is incomplete. Generative search inserts engine-mediated decisions between the user's request and the final recommendation — search activation, query fan-out, evidence exposure, and selection. Page relevance alone cannot explain who gets mentioned.

The paper analyzes 34,960 unbranded prompt-engine observations from 75 anonymized monitoring projects, covering 2,854 distinct prompts and repeated GPT and Gemini runs from June–September 2026. The headline finding: when neither the target brand nor its own domain appears in the observable retrieval path, mention rates are 2.8% for GPT and 3.8% for Gemini. With an own-domain citation but no branded fan-out, they rise to 49.0% and 58.4%. When both own-domain exposure and a branded fan-out occur, mention rates reach 91.4% and 100%.

The practical translation: AI visibility is a multi-stage pipeline, not a single optimization checklist.

Why a single GEO score is not enough

Traditional SEO compresses ranking into two intuitions: relevance and authority. Generative search adds at least two more verbs — retrieve and select.

Tannenbaum's compact mnemonic:

AI visibility ≈ Match × Exposure × Selection + Prior

Each term maps to a distinct failure mode:

StageQuestion it answersWhat failure looks like
MatchIs there a page that answers the real request?Good content exists but targets the wrong prompt
ExposureDoes the engine surface brand-supporting evidence?Page ranks in search but never enters the citation set
SelectionGiven that evidence, does the brand survive into the answer?Domain is cited but brand is not named
PriorCan the brand appear without visible live evidence?Competitor mentioned 74 times with only 7 own-domain citations

A deterministic page score can estimate match quality. It cannot, by itself, predict whether a live engine will expose or select your brand. That distinction matters for budget allocation.

Finding 1: Commercial prompts trigger search fan-out far more often

Among 80 unique real-user prompts in the study's fan-out cohort, 25% triggered at least one observed search-query fan-out. The rate is highly intent-dependent:

Prompt intentFan-out trigger rate
Commercial78.3% (18 of 23)
Informational3.6% (2 of 55)

Risk ratio: 21.5× (95% CI approximately 5.4–85.3).

When fan-out occurs, the engine often constructs an evaluation space — best options, features, comparisons, pricing — that differs from the user's literal wording. A brand competes against the fan-out representation, not just the surface prompt.

Operational implication: Buyer-intent prompt panels matter more than keyword lists. If your monitoring set is mostly informational queries, you may be measuring a retrieval path that rarely activates for commercial decisions.

Finding 2: Page match is real but engine-specific

The study crawled 275 pages from one organization's site and computed a normalized BM25 best-page match score for 199 prompts.

EngineAUC: match → own-domain citationAUC: match → brand mention
GPT0.545 (near chance)0.539
Gemini0.6410.608

For Gemini, the highest match quartile produced a 50.0% own-domain citation rate versus 22.0% in the lowest quartile — a risk ratio of 2.27. For GPT, the match gradient was flat.

The same relevance statistic has different downstream value by engine. A single engine-agnostic "AI readiness" score cannot serve as a universal visibility predictor. See also Why Each AI Engine Reads a Different Search Index.

Finding 3: Evidence exposure dominates selection

The exposure-selection relationship is an order of magnitude larger than page match.

For one organization in the case study:

ConditionGPT mention rateGemini mention rate
Own-domain URL in source list100% (30/30)73.8% (48/65)
No own-domain citation9.5% (16/169)8.2% (11/134)

Risk ratios: 10.6× for GPT, 9.0× for Gemini.

Even in the highest match quartile, high page fit without exposure still loses. For Gemini, among 25 high-match prompts with own-domain citation, the brand was mentioned on 84%; among 25 high-match prompts without citation, only 8%.

Page fit is a favorable upstream condition. Its downstream value depends on whether the engine surfaces evidence. This aligns with Indexably's finding that domain factors dominate page-level signals for most sites — exposure is often gated upstream of on-page work.

Finding 4: The large-panel four-cell ladder

On 34,960 unbranded observations across 75 projects, the study reports a clean four-cell pattern:

Evidence stateGPT mention rateGemini mention rate
Neither signal2.8%3.8%
Own domain only49.0%58.4%
Branded fan-out only64.4%84.8%
Both signals91.4%100.0%

Relative to the neither-signal baseline, own-domain exposure alone corresponds to a 17.4-fold GPT and 15.4-fold Gemini increase in mention probability. The combination of both signals corresponds to 32.4-fold and 26.3-fold increases.

Within-prompt repeated runs strengthen the pattern. A Cochran–Mantel–Haenszel common odds ratio stratified by organization and prompt is 15.3 for GPT and 29.7 for Gemini when own-domain exposure varies within the same prompt cell. That controls for time-invariant brand identity and prompt wording.

Finding 5: The prior-compatible path is large

The study's repeated-run benchmark (16 prompts × 10 ChatGPT runs = 160 answers) reveals a citation-free visibility path:

EntityBrand mentionsOwn-domain citations
Competitor 1747
Competitor 2655
Competitor 33611
Target brand (Org B)129

Four prompts produced zero citations across all 10 runs — 40 citation-free answers. Competitor 1 was mentioned in 30 of 40; Competitor 2 in 28 of 40; the target brand in zero.

This is compatible with training-time brand knowledge, third-party evidence not represented by own-domain citations, or other unobserved information. The data cannot prove the mechanism. They do prove a measurement fact: an own-domain citation model alone cannot account for final brand visibility.

That connects directly to ghost citations — AI can use your content without naming you, and name competitors without citing them.

The fitted stage equation

The paper fits a predictive model separately by engine:

logit P(Mention) = α + β·logit(Prior) + γ·Exposure + δ·Fan-out + θ·Intent controls

On a 30% holdout, the full model achieves AUC 0.963 on GPT and 0.942 on Gemini, compared with 0.937/0.917 for prior history alone and 0.880/0.840 for live signals alone.

Prior visibility is independently persistent. A previous non-mention plus no current own-domain exposure yields next-run mention rates of 1.6% and 1.9%; previous mention plus current exposure yields 80.5% and 83.7%.

The equation is predictive and observational — not a causal description of proprietary engine internals. But it gives practitioners a diagnostic framework: which stage is actually failing?

Practical playbook by failure mode

If match is the bottleneck

Symptoms: No owned page clearly answers buyer prompts; competitors with similar authority get cited on the same queries.

Actions:

  1. Map real buyer prompts (not keyword lists) to existing pages.
  2. Build comparison, product, and docs pages that answer fan-out queries — not just blog posts. See Product Pages vs Blog Posts.
  3. Measure match per engine; do not assume one relevance score generalizes.

If exposure is the bottleneck

Symptoms: Pages rank in traditional search; competitors get cited; your domain rarely appears in AI source lists.

Actions:

  1. Invest in referring-subnet diversity and third-party mentions — not another FAQ rewrite sprint.
  2. Target surfaces AI already cites in your category: earned media, Reddit at BOFU.
  3. Separate retrieval eligibility from citation absorption in your KPIs.

If selection is the bottleneck

Symptoms: Own-domain citations appear but brand is not named; ghost citations on third-party pages.

Actions:

  1. Strengthen entity signals: consistent brand naming, organization schema, authoritative definitions.
  2. Ensure cited pages lead with direct, extractable brand-attributed answers.
  3. Track mention rate conditional on citation — not citation rate alone.

If prior is the bottleneck

Symptoms: Established competitors mentioned frequently with few own-domain citations; your brand invisible despite good content.

Actions:

  1. Accept that model familiarity compounds over time — short-term page rewrites may not move the needle.
  2. Build category presence across review sites, communities, and press — the surfaces that feed model priors.
  3. Monitor competitor mention-to-citation ratios to understand which path dominates in your category.

What this does not mean

It does not mean page optimization is pointless. Match still matters — especially on Gemini, and especially once you are in the retrieval pool.

It does not prove citations cause mentions. The study is careful about confounds. Well-maintained sites with citations also have authority, distribution, and entity footprint.

It does not replace relevance. The authors' own framing says content that answers the prompt still wins — but only after the engine exposes it.

It does not contradict volatility research. AI Citation Volatility shows citation sets rotate quickly. This paper explains why a single snapshot misleads: the pipeline has multiple stochastic stages.

How Obsurfable fits

Obsurfable records the full observation chain: prompts, answers, brands mentioned, and citations — across repeated runs so you can see which stage is failing.

If your page score is high but mention rate is flat, the bottleneck is likely exposure or selection — not another on-page cycle. If competitors appear without citations, the prior path may dominate. The Visibility Director flags prompt-level gaps at the stage that is actually broken.

For the companion paper on cross-engine divergence and engine-agnostic scores, see Why Engine-Agnostic GEO Scores Cannot Predict Citations.

FAQ

Should I stop using GEO page scores entirely?

No. Use them to diagnose match quality for specific prompts. Do not interpret them as end-to-end visibility probabilities without an exposure model.

Why does GPT show weaker match effects than Gemini?

The paper does not identify the mechanism. Plausible explanations include different retrieval architectures, reranking behavior, and how each engine weights domain authority versus page fit. Optimize per engine.

How many prompts do I need to measure stage effects reliably?

The large panel uses thousands of observations. For a single brand, start with 20–50 buyer-intent prompts run repeatedly (weekly minimum) per engine. Single checks are one sample from a distribution — not a measurement.

Does branded fan-out mean I should put my brand name in prompts?

No. The study uses unbranded prompts. Branded fan-out means the engine's search expansion includes the target brand — an engine behavior, not a prompt-writing tactic.

Bottom line

Across 34,960 unbranded observations, brand mentions rise from under 4% to over 91% when own-domain citations and branded fan-outs both appear. Page match is real but engine-specific. Evidence exposure is associated with roughly tenfold selection effects. Competitors can win through a prior path that bypasses own-domain citations entirely.

AI visibility is Match × Exposure × Selection + Prior. Measure which term is failing — then sequence the work accordingly.