The GEO tooling market is filling with deterministic page scores: inspect a URL, output an "AI readiness" number, interpret it as citation probability. The appeal is obvious. A page score is cheap, stable, and easy to explain. A live generative engine is expensive, rate-limited, personalized, and non-stationary.
A September 2026 paper by Benjamin Tannenbaum asks whether that shortcut works. The answer is nuanced: engine-free scores can estimate page quality or query-page fit, but end-to-end visibility additionally depends on engine-specific exposure and selection — and the engines disagree far more than most score vendors imply.
What Bajemon and Rochet found — and what they did not
Bajemon and Rochet (2026) stress-tested deterministic content scoring on modern engines. Their key result: query-agnostic page scores have weak within-query association with citation order, while query-conditioned relevance is materially more informative.
That design is valuable. It isolates citation preference conditional on exposure — what happens after a page is already in the candidate set. It intentionally removes the production process that determines whether a page is searched, retrieved, reranked, placed in context, and made available for citation.
Tannenbaum's paper measures that missing layer.
The cross-engine divergence audit
On 6 June 2026, a fixed benchmark of 15 commercial prompts about AI search visibility software produced 589 citation observations across ChatGPT, Microsoft Copilot, Google, and Perplexity — 528 unique URLs and 356 domains.
Same-prompt cross-engine URL overlap was extremely small:
| Metric | Value |
|---|---|
| Mean pairwise Jaccard similarity | 0.0079 (95% CI 0.0037–0.0132) |
| Median pairwise Jaccard | Zero |
| Engine pairs sharing no cited URL | 84.9% |
On the 10 prompts observed on all four engines:
| Metric | Value |
|---|---|
| Mean exact-URL Jaccard | 0.0072 |
| Reciprocal-rank-weighted overlap | 0.0027 |
| Top-five exact-URL overlap | Zero in all 60 pairwise comparisons |
A single engine captured only 11.4%–42.6% of the four-engine URL union. The union contained 2.43× as many URLs as the broadest single engine on average. Among all URLs observed, 96.4% appeared in only one engine.
Per-engine coverage on 6 June
| Engine | Citation rows | Unique URLs | Unique domains |
|---|---|---|---|
| ChatGPT | 231 | 201 | 155 |
| Copilot | 63 | 62 | 52 |
| 145 | 136 | 95 | |
| Perplexity | 150 | 148 | 105 |
| All engines | 589 | 528 | 356 |
Even monitoring the broadest single engine misses most of the multi-engine source universe.
Temporal turnover vs cross-engine divergence
A separate 5-to-6 June same-engine comparison found mean URL-set turnover of 67.0% (95% CI 61.1%–72.2%). Citations rotate quickly within an engine.
But cross-engine divergence is much larger than temporal noise. A same-engine adjacent-day citation set was still about 42× more similar than a same-day cross-engine set in this benchmark.
Implication: engine identity matters more than day-to-day volatility for source-set composition. That aligns with Foglift's Q3 2026 finding that AI engines do not share citation lists — and with the broader result that only 2.7% of domains are cited by all five engines.
The visibility vector: six coordinates, not one score
Tannenbaum formalizes a visibility vector:
V(page, request, engine, time) = (Q, E, S, A, B, Y)
| Coordinate | What it measures |
|---|---|
| Q | Query-page fit |
| E | Exposure to the generation or citation process |
| S | Citation selection conditional on exposure |
| A | Citation absorption — substantive use in the answer |
| B | Brand/recommendation prominence |
| Y | Downstream user outcome (traffic, conversion) |
Different experiments identify different coordinates. Treating them as interchangeable creates apparent contradictions that are often only estimand mismatches.
A deterministic engine-free score g(page, request) can estimate Q or S — but end-to-end visibility is:
P(Citation) = P(Exposure) × P(Citation | Exposure)
Two engines can share identical conditional preference over exposed pages and produce arbitrarily different end-to-end citation distributions if their exposure functions differ.
What engine-agnostic scores actually estimate
The paper is explicit: these results do not invalidate engine-free page scoring. They identify its estimand.
| What a page score can estimate | What it cannot estimate without a live engine |
|---|---|
| Content quality | Whether the engine triggers search for this prompt |
| Query-page fit (if query-conditioned) | Which fan-out queries the engine generates |
| Conditional citation preference (in fixed-candidate designs) | Which pages enter the retrieval pool |
| Structural extractability | Engine-specific reranking and context allocation |
An "AI visibility score" that never observes an engine should not be interpreted as a stable probability of being surfaced or cited without an explicit exposure model.
This connects to the stage model of brand visibility: match quality (Q) is one term in a multiplicative pipeline. A perfect Q cannot compensate for zero exposure (E).
Why query-conditioned scores still fall short
Even query-conditioned relevance — the stronger variant Bajemon and Rochet identify — only addresses fit given a request. It does not address:
- Search activation — whether the engine decides to search at all (commercial prompts trigger fan-out 78.3% of the time; informational prompts 3.6% in Tannenbaum's fan-out cohort).
- Query rewriting — one human request becoming multiple search queries with different intent mixes.
- Engine-specific retrieval — different indexes, rerankers, and context budgets per platform.
- Prior-compatible paths — brands mentioned without own-domain citations.
Martinez's critical GEO survey (2026) reaches a compatible conclusion: GEO is a stochastic, partially observable pipeline spanning search activation through user behavior — not a single ranking task. No reviewed technique shows stable, longitudinal, cross-platform causal effects on organic discoverability.
Practical reporting framework
Tannenbaum proposes reporting four separate quantities instead of one blended score:
1. Page fit (Q)
Query-conditioned relevance between prompt and owned pages. Use transparent lexical or embedding methods. Report per prompt, not site-wide averages.
2. Observed exposure (E)
Whether your domain appears in the engine's stored source list for tracked prompts. Binary per run; aggregate as exposure rate over repeated observations.
3. Conditional selection (S)
Mention or citation rate conditional on exposure. Separates "we got into the pool" from "we survived synthesis."
4. Final visibility (B or citation rate)
End-to-end mention or citation rate across all runs — the number executives want, but only meaningful with explicit denominators and repeated sampling.
| Metric | Formula (conceptual) | Common mistake |
|---|---|---|
| Exposure rate | Runs with own-domain citation / total runs | Confusing with mention rate |
| Selection rate | Runs with mention given citation / runs with citation | Ignoring when cited but not mentioned |
| End-to-end mention rate | Runs with mention / total runs | Single-run snapshots |
| Cross-engine coverage | Union of cited domains across engines / single-engine count | Monitoring one engine only |
What vendors should disclose
When evaluating an "AI visibility score" product, ask:
- Which coordinate does it measure? Fit, exposure, selection, or end-to-end visibility?
- Does it query live engines? If not, it estimates page quality — not citation probability.
- How many engines? A score from one engine explains at most ~40% of the multi-engine URL universe in Tannenbaum's audit.
- What is the sampling design? Single runs capture 62%–77% of brands observed across five runs in a September 2026 study — repeated sampling is not optional.
- Is the denominator explicit? "38% visibility" means nothing without "of how many answers."
For a methodology that addresses several of these points, see How to Measure AI Visibility Without Fooling Yourself.
Playbook for marketing teams
Stop
- Treating a page GEO score as a citation probability.
- Monitoring one engine and extrapolating to "AI search."
- Comparing scores from different vendors with different methodologies.
- Reacting to single-run citation changes.
Start
- Tracking exposure, selection, and end-to-end visibility as separate KPIs.
- Running a fixed buyer-intent prompt panel repeatedly per engine.
- Measuring cross-engine citation union — not just your best engine.
- Sequencing work to the bottleneck: if exposure is zero, page rewrites will not help. See Page Optimization Only Helps High-Authority Domains.
For high-authority domains
Page fit and conditional selection are high-ROI once you are in the retrieval pool. Engine-agnostic scores may correlate with S. Validate with live repeated runs — do not trust the score alone.
For challenger brands
Exposure is usually the gate. Invest in distribution, earned mentions, and multi-engine footprint before optimizing page structure. A perfect page score with zero exposure measures nothing useful.
How Obsurfable fits
Obsurfable does not sell a deterministic page score. It records live observations: prompts, answers, brands mentioned, citations, and sources — across repeated runs so you can decompose fit, exposure, and selection from your own history.
When a page score says "ready" but Obsurfable shows zero exposure across 50 repeated runs, the score is measuring the wrong stage. When exposure is high but mentions are low, the bottleneck is selection — not another FAQ block.
The public corpus at obsurfable.com lets anyone browse the same observation structure: what AI actually said, what it cited, and which brands it named.
FAQ
Are engine-agnostic scores useless?
No. They can audit page fit, structural extractability, and conditional citation preference in controlled settings. They are useful inputs — not outputs that replace live measurement.
Why is cross-engine overlap so low?
Engines use different search backends, rerankers, query-rewriting strategies, and context budgets. Each engine reads a different search index. Low overlap is expected, not a measurement error.
Should I optimize for the engine with the most citations in my category?
Start there for quick wins, but do not treat one engine as "AI search." Multi-engine monitoring reveals gaps that single-engine scores hide.
How does this relate to Google's "GEO is still SEO" guidance?
Google's framing applies to Google surfaces, which share an index and quality system. Tannenbaum's audit includes Copilot and Perplexity — platforms with different retrieval stacks. Engine-agnostic scores conflate platforms that do not share source pools.
Bottom line
On 15 matched commercial prompts, mean cross-engine citation overlap is 0.0079 Jaccard. 96.4% of cited URLs appear in only one engine. A page score without a live exposure model estimates fit or conditional preference — not whether you will be cited.
Report page fit, observed exposure, conditional selection, and final visibility as separate quantities. Monitor multiple engines with repeated sampling. Sequence optimization to the stage that is actually failing.
An "AI visibility score" that never talks to an engine is measuring a coordinate — not the game.