Obsurfable

Why Engine-Agnostic GEO Scores Cannot Predict Citations

Obsurfable

The GEO tooling market is filling with deterministic page scores: inspect a URL, output an "AI readiness" number, interpret it as citation probability. The appeal is obvious. A page score is cheap, stable, and easy to explain. A live generative engine is expensive, rate-limited, personalized, and non-stationary.

A September 2026 paper by Benjamin Tannenbaum asks whether that shortcut works. The answer is nuanced: engine-free scores can estimate page quality or query-page fit, but end-to-end visibility additionally depends on engine-specific exposure and selection — and the engines disagree far more than most score vendors imply.

What Bajemon and Rochet found — and what they did not

Bajemon and Rochet (2026) stress-tested deterministic content scoring on modern engines. Their key result: query-agnostic page scores have weak within-query association with citation order, while query-conditioned relevance is materially more informative.

That design is valuable. It isolates citation preference conditional on exposure — what happens after a page is already in the candidate set. It intentionally removes the production process that determines whether a page is searched, retrieved, reranked, placed in context, and made available for citation.

Tannenbaum's paper measures that missing layer.

The cross-engine divergence audit

On 6 June 2026, a fixed benchmark of 15 commercial prompts about AI search visibility software produced 589 citation observations across ChatGPT, Microsoft Copilot, Google, and Perplexity — 528 unique URLs and 356 domains.

Same-prompt cross-engine URL overlap was extremely small:

MetricValue
Mean pairwise Jaccard similarity0.0079 (95% CI 0.0037–0.0132)
Median pairwise JaccardZero
Engine pairs sharing no cited URL84.9%

On the 10 prompts observed on all four engines:

MetricValue
Mean exact-URL Jaccard0.0072
Reciprocal-rank-weighted overlap0.0027
Top-five exact-URL overlapZero in all 60 pairwise comparisons

A single engine captured only 11.4%–42.6% of the four-engine URL union. The union contained 2.43× as many URLs as the broadest single engine on average. Among all URLs observed, 96.4% appeared in only one engine.

Per-engine coverage on 6 June

EngineCitation rowsUnique URLsUnique domains
ChatGPT231201155
Copilot636252
Google14513695
Perplexity150148105
All engines589528356

Even monitoring the broadest single engine misses most of the multi-engine source universe.

Temporal turnover vs cross-engine divergence

A separate 5-to-6 June same-engine comparison found mean URL-set turnover of 67.0% (95% CI 61.1%–72.2%). Citations rotate quickly within an engine.

But cross-engine divergence is much larger than temporal noise. A same-engine adjacent-day citation set was still about 42× more similar than a same-day cross-engine set in this benchmark.

Implication: engine identity matters more than day-to-day volatility for source-set composition. That aligns with Foglift's Q3 2026 finding that AI engines do not share citation lists — and with the broader result that only 2.7% of domains are cited by all five engines.

The visibility vector: six coordinates, not one score

Tannenbaum formalizes a visibility vector:

V(page, request, engine, time) = (Q, E, S, A, B, Y)

CoordinateWhat it measures
QQuery-page fit
EExposure to the generation or citation process
SCitation selection conditional on exposure
ACitation absorption — substantive use in the answer
BBrand/recommendation prominence
YDownstream user outcome (traffic, conversion)

Different experiments identify different coordinates. Treating them as interchangeable creates apparent contradictions that are often only estimand mismatches.

A deterministic engine-free score g(page, request) can estimate Q or S — but end-to-end visibility is:

P(Citation) = P(Exposure) × P(Citation | Exposure)

Two engines can share identical conditional preference over exposed pages and produce arbitrarily different end-to-end citation distributions if their exposure functions differ.

What engine-agnostic scores actually estimate

The paper is explicit: these results do not invalidate engine-free page scoring. They identify its estimand.

What a page score can estimateWhat it cannot estimate without a live engine
Content qualityWhether the engine triggers search for this prompt
Query-page fit (if query-conditioned)Which fan-out queries the engine generates
Conditional citation preference (in fixed-candidate designs)Which pages enter the retrieval pool
Structural extractabilityEngine-specific reranking and context allocation

An "AI visibility score" that never observes an engine should not be interpreted as a stable probability of being surfaced or cited without an explicit exposure model.

This connects to the stage model of brand visibility: match quality (Q) is one term in a multiplicative pipeline. A perfect Q cannot compensate for zero exposure (E).

Why query-conditioned scores still fall short

Even query-conditioned relevance — the stronger variant Bajemon and Rochet identify — only addresses fit given a request. It does not address:

  1. Search activation — whether the engine decides to search at all (commercial prompts trigger fan-out 78.3% of the time; informational prompts 3.6% in Tannenbaum's fan-out cohort).
  2. Query rewriting — one human request becoming multiple search queries with different intent mixes.
  3. Engine-specific retrieval — different indexes, rerankers, and context budgets per platform.
  4. Prior-compatible paths — brands mentioned without own-domain citations.

Martinez's critical GEO survey (2026) reaches a compatible conclusion: GEO is a stochastic, partially observable pipeline spanning search activation through user behavior — not a single ranking task. No reviewed technique shows stable, longitudinal, cross-platform causal effects on organic discoverability.

Practical reporting framework

Tannenbaum proposes reporting four separate quantities instead of one blended score:

1. Page fit (Q)

Query-conditioned relevance between prompt and owned pages. Use transparent lexical or embedding methods. Report per prompt, not site-wide averages.

2. Observed exposure (E)

Whether your domain appears in the engine's stored source list for tracked prompts. Binary per run; aggregate as exposure rate over repeated observations.

3. Conditional selection (S)

Mention or citation rate conditional on exposure. Separates "we got into the pool" from "we survived synthesis."

4. Final visibility (B or citation rate)

End-to-end mention or citation rate across all runs — the number executives want, but only meaningful with explicit denominators and repeated sampling.

MetricFormula (conceptual)Common mistake
Exposure rateRuns with own-domain citation / total runsConfusing with mention rate
Selection rateRuns with mention given citation / runs with citationIgnoring when cited but not mentioned
End-to-end mention rateRuns with mention / total runsSingle-run snapshots
Cross-engine coverageUnion of cited domains across engines / single-engine countMonitoring one engine only

What vendors should disclose

When evaluating an "AI visibility score" product, ask:

  1. Which coordinate does it measure? Fit, exposure, selection, or end-to-end visibility?
  2. Does it query live engines? If not, it estimates page quality — not citation probability.
  3. How many engines? A score from one engine explains at most ~40% of the multi-engine URL universe in Tannenbaum's audit.
  4. What is the sampling design? Single runs capture 62%–77% of brands observed across five runs in a September 2026 study — repeated sampling is not optional.
  5. Is the denominator explicit? "38% visibility" means nothing without "of how many answers."

For a methodology that addresses several of these points, see How to Measure AI Visibility Without Fooling Yourself.

Playbook for marketing teams

Stop

  • Treating a page GEO score as a citation probability.
  • Monitoring one engine and extrapolating to "AI search."
  • Comparing scores from different vendors with different methodologies.
  • Reacting to single-run citation changes.

Start

  • Tracking exposure, selection, and end-to-end visibility as separate KPIs.
  • Running a fixed buyer-intent prompt panel repeatedly per engine.
  • Measuring cross-engine citation union — not just your best engine.
  • Sequencing work to the bottleneck: if exposure is zero, page rewrites will not help. See Page Optimization Only Helps High-Authority Domains.

For high-authority domains

Page fit and conditional selection are high-ROI once you are in the retrieval pool. Engine-agnostic scores may correlate with S. Validate with live repeated runs — do not trust the score alone.

For challenger brands

Exposure is usually the gate. Invest in distribution, earned mentions, and multi-engine footprint before optimizing page structure. A perfect page score with zero exposure measures nothing useful.

How Obsurfable fits

Obsurfable does not sell a deterministic page score. It records live observations: prompts, answers, brands mentioned, citations, and sources — across repeated runs so you can decompose fit, exposure, and selection from your own history.

When a page score says "ready" but Obsurfable shows zero exposure across 50 repeated runs, the score is measuring the wrong stage. When exposure is high but mentions are low, the bottleneck is selection — not another FAQ block.

The public corpus at obsurfable.com lets anyone browse the same observation structure: what AI actually said, what it cited, and which brands it named.

FAQ

Are engine-agnostic scores useless?

No. They can audit page fit, structural extractability, and conditional citation preference in controlled settings. They are useful inputs — not outputs that replace live measurement.

Why is cross-engine overlap so low?

Engines use different search backends, rerankers, query-rewriting strategies, and context budgets. Each engine reads a different search index. Low overlap is expected, not a measurement error.

Should I optimize for the engine with the most citations in my category?

Start there for quick wins, but do not treat one engine as "AI search." Multi-engine monitoring reveals gaps that single-engine scores hide.

How does this relate to Google's "GEO is still SEO" guidance?

Google's framing applies to Google surfaces, which share an index and quality system. Tannenbaum's audit includes Copilot and Perplexity — platforms with different retrieval stacks. Engine-agnostic scores conflate platforms that do not share source pools.

Bottom line

On 15 matched commercial prompts, mean cross-engine citation overlap is 0.0079 Jaccard. 96.4% of cited URLs appear in only one engine. A page score without a live exposure model estimates fit or conditional preference — not whether you will be cited.

Report page fit, observed exposure, conditional selection, and final visibility as separate quantities. Monitor multiple engines with repeated sampling. Sequence optimization to the stage that is actually failing.

An "AI visibility score" that never talks to an engine is measuring a coordinate — not the game.