Obsurfable

GEO in 2026: What the New Evidence Review Says Actually Works

Obsurfable

In July 2026, Olivier Martinez published a critical survey on arXiv reviewing 45 studies on Generative Engine Optimization (GEO) published between November 2023 and July 2026. Its conclusion is uncomfortable but useful: no reviewed technique demonstrates a stable, cross-platform, longitudinal effect on organic discoverability or downstream traffic. GEO is real — but the commercial claims built around it have outrun the evidence.

That does not mean optimization is pointless. It means the field is noisier than most playbooks imply, and teams need a sharper measurement model before they rewrite pages or redirect budgets.

The headline finding: visibility is a pipeline, not a rank

The survey's central argument is that GEO is not one ranking problem. It is a stochastic, partially observable pipeline spanning:

  1. Search activation (does the engine run retrieval at all?)
  2. Crawling and indexing
  3. Retrieval eligibility
  4. Reranking and context allocation
  5. Citation selection
  6. Prominence within the answer
  7. Factual absorption and fidelity
  8. User behavior (clicks, conversions)

Tactics that improve one stage can hurt another. A page rewritten to perform better once it is already in context may become less likely to be retrieved at all. Collapsing these stages into a single "AI visibility score" hides the trade-off — and lets a laboratory result stand in for a business outcome.

Martinez proposes a visibility vector that separates discoverability, exposure, citation probability, prominence, absorption, fidelity, and economic outcomes. The practical implication: your dashboard should track retrieval and citation as distinct metrics, not one blended number.

The 40% figure the survey rejects

The most widely circulated number in GEO comes from Aggarwal et al.'s foundational 2024 paper at ACM SIGKDD, which reported gains of up to 40%. Martinez traces that figure to a single metric in a single configuration.

The 40% derives from Position-Adjusted Word Count (PAWC) rising from 19.3 to 27.2 under a strategy called Quotation Addition — roughly 41% in relative terms. That gain occurs in a testbed where five documents have already been supplied to the generator. It does not mean 40% more readers will click. It does not mean a page gains 40% in retrieval probability. It means that, in that specific setup, a source already handed to the model receives a larger position-weighted share of attributed text.

The survey places "GEO increases visibility by 40%" in its lowest confidence tier — rejected as a general claim. Yet that number has traveled far. A Wall Street Journal investigation documented businesses paying substantial sums to shape how ChatGPT describes them, with chatbot referrals growing from under a million visits in early 2024 to more than 230 million monthly by September 2025.

The gap between laboratory gains and business outcomes is the survey's core correction.

When the rewrite backfires: SAGEO Arena

The most operationally pointed result concerns what happens when optimization is tested across the full pipeline rather than inside a fixed context.

SAGEO Arena — an end-to-end environment across 171,003 documents and 2,700 queries — found that optimizing only the body of a page:

MetricChange
Average top-20 presence−9%
Top-10 presence after reranking−16%
Final citation rate−6%

Applying an automated optimization system to the body alone produced larger losses still. The mechanism is straightforward: if a transformation raises the conditional probability of citation given retrieval but lowers the probability of retrieval, the total effect can be negative. A page can be rewritten to perform better once selected while becoming less likely to be selected at all — invisible to any test that starts with the document already in context.

This is why "citation rewrites" have become a flashpoint in 2026 GEO debates. Body-only optimization is not a free win.

What the evidence actually supports

Across the 45 reviewed studies, two levers hold up most consistently:

1. Query-document relevance. The most reproducible factor in a large factorial experiment (252,000 trials across six language models and eighteen factors) was relevance between the query and the document. This points back toward fundamentals: topical authority, intent match, and content that actually answers the question.

2. Position within the context window. Where a passage appears in the retrieved context matters — passages higher in the allocation window receive more attribution. This aligns with practical AEO advice about leading with direct answers, but the survey frames it as a conditional effect, not a universal formatting checklist.

Extractable evidence — verifiable figures, definitions, dated facts, prices — shows a more moderate advantage, and the effect depends on intent. A recent date helps a time-sensitive query but not a stable definition. A fabricated statistic may increase reuse while degrading answer accuracy. The operative criterion is verifiable, dated, properly attributed evidence, not simply "add numbers."

Cyrus Shepard's scored analysis of 54 studies, cited in the survey's coverage, placed URL accessibility (9.5) and search rank (9.4) as the two highest predictors of AI citation on an evidence-weighted scale — pointing back toward traditional retrieval fundamentals rather than novel formatting tricks.

What does not transfer

Several widely promoted tactics fail under cross-domain testing:

  • C-SEO Bench tested optimization methods across two tasks, six domains, roughly 1,900 queries, and 16,360 documents. Only three of 54 method-domain combinations were significantly positive in the main experiment, with none positive in question answering. Several transformations reduced rank outright.
  • In e-commerce, ten of fifteen initial heuristics were neutral or negative; only systematically optimized prompts performed better.
  • Keyword stuffing, imported from conventional SEO, does not transfer and can lower position-adjusted metrics.

Generic heuristics — "add statistics," "use question headings," "insert quotations" — may work in one simulator and fail in end-to-end retrieval. The survey's blunt synthesis: already-retrieved content can causally influence an answer, but no reviewed technique shows stable cross-platform effects on organic discoverability.

Engines do not share a source list

One of the survey's firmer conclusions is that there is no single global ranking to optimize for.

  • An audit of 4,706 queries found 53% of domains cited by Google's AI Overviews did not appear in the organic top 10, and 27% were absent from the top 100.
  • A separate audit across 11,500 queries reported URL-level overlap between organic Google, AI Overviews, and Gemini in the range of 0.11 to 0.18.
  • Research on 89,000 LinkedIn URLs cited across ChatGPT Search, Google AI Mode, and Perplexity found Perplexity cites company pages 59% of the time while ChatGPT Search and Google AI Mode cite individual creators 59% of the time.

This aligns with CiteMetrix's finding that only 11% of cited domains overlap between ChatGPT and Perplexity. Platform-specific strategy is not optional — see Only 11% Overlap: Why ChatGPT and Perplexity Need Separate AEO Strategies.

Activation rates compound the fragmentation. One study of 55,393 trending queries found AI Overviews activated on 13.7% of queries overall but 64.7% for queries phrased as questions. A different representative sample showed AI Overviews on 51.5%. Rates quoted without sample context carry little weight.

Citation does not mean support

The survey draws a sharp line between being cited and being accurately represented.

Across four early generative engines, only 51.5% of sentences were fully supported by their cited sources, and 74.5% of citations actually supported the proposition they were attached to. Newer work reports credible-source shares from 71.4% to 86.3% depending on the assistant and topic. Roughly 11% of 98,020 atomic claims were classified as insufficiently supported.

One audit found approximately 16% of more than 19,000 retrieved pages were labeled AI-generated, while 27.1% of URLs could not be scraped at all. A visible citation may point to a page that does not support the claim — or to synthetic content feeding back into the system.

For all the money flowing into GEO, the evidence on actual traffic and revenue is the thinnest in the field.

One log-based natural experiment found total ChatGPT referrals to a site increased by a factor of 5.7, but untreated pages on the same site had already risen by a factor of 3.5 as the platform grew. After controlling for platform growth, the estimated additional effect was a multiplier of 1.82 (95% CI: 1.31–2.54), with a conservative placebo test yielding a statistically inconclusive result.

Meanwhile, the traffic at stake is real and measured:

  • A randomized field experiment with 1,065 desktop Chrome users found Google's AI Overviews cut outbound organic clicks by 39.8% and raised zero-click searches by 34.5%.
  • When AI Overviews appear, click-through rates at position one fall from 27% to 11%.
  • An Ahrefs study of 300,000 keywords found AI Overviews correlated with a 58% reduction in CTR for top-ranking pages.

The traffic loss is measured. The case for GEO as the remedy is not — at least not yet.

What to stop doing

Based on the survey's evidence hierarchy:

  • Treating one case-study uplift as universal. Most positive results are conditional on a fixed context, a single engine, or a specific query type.
  • Assuming a formatting checklist transfers across ChatGPT, Gemini, Perplexity, and AI Overviews. Cross-engine overlap is structurally low.
  • Running single-shot "did we get cited?" tests and calling it strategy. Citation sets rotate weekly — Google exchanges 56% of AI Mode sources weekly; ChatGPT exchanges 74%.
  • Body-only rewrites without measuring retrieval impact. SAGEO Arena shows this can reduce top-10 presence by 16%.
  • Collapsing citation and traffic into one ROI narrative. The evidence for each is at different maturity levels.

Practical playbook for teams

1. Separate retrieval metrics from citation metrics.

Track whether your pages appear in the retrieval pool at all, not just whether they get cited once selected. A citation gain with a retrieval loss is a net negative.

2. Track weekly volatility for fixed prompt sets.

Run the same prompts across multiple days. A single appearance is not durable visibility. See AI Citation Volatility: Why Rank-Style Tracking Fails.

3. Test changes with controls.

Before/after on matched prompts, with untreated baselines. The survey recommends multiple query phrasings, multiple engines with version and date recorded, and outputs with no citations treated as data — not discarded.

4. Prioritize source quality and distribution over cosmetic rewrites.

Relevance, crawlability, search rank, and third-party corroboration outperform generic formatting heuristics in the reviewed evidence. Build topic depth, earn placements on surfaces each engine trusts, and refresh high-velocity content on platform-appropriate cadences.

5. Measure citation and mention separately.

Roughly 62% of AI citations are "ghost citations" — your URL appears but your brand is not named. Citation rate and mention rate are different KPIs. See Ghost Citations: When AI Uses Your Content but Doesn't Name Your Brand.

How Obsurfable fits

Obsurfable runs your defined prompts against AI search engines and tracks brand mentions, competitor positioning, and cited sources over time — with repetition built in, not bolted on. When a shift like the July 2026 evidence review reframes what "works," you see whether your actual answer-layer presence moved, not whether you followed a checklist.

The Visibility Director can flag gaps where competitors appear in AI answers and you don't, then draft content designed to close them.

For fundamentals, see What Is Answer Engine Optimization?. For the Gemini model shift that weakened rank-citation correlation, see The Gemini 3 Citation Reset.

FAQ

Does this survey mean GEO is useless?

No. The survey confirms that already-retrieved content can causally influence how it is cited and used. The correction is narrower: no reviewed technique proves stable, cross-platform effects on organic discoverability or traffic. Optimization should be measured, staged, and engine-specific — not applied as a universal rewrite formula.

Should I stop doing body-text optimization?

Not necessarily — but test it end-to-end. If a rewrite improves citation given retrieval but reduces retrieval probability, the net effect may be negative. Measure both stages.

Is the 40% figure completely wrong?

It is valid within its experimental setting — five documents already in context, one metric (PAWC), one strategy (Quotation Addition). It is not a general promise of 40% more traffic, clicks, or retrieval probability.

What is the single highest-confidence lever?

Query-document relevance, followed by position within the retrieved context. Both require the page to be retrieved first — which is why crawlability, indexing, and organic authority remain foundational.

Bottom line

The strongest 2026 GEO signal is methodological. Teams that measure repeatedly, across engines, with controls — separating retrieval from citation, citation from mention, and visibility from traffic — will outperform teams relying on static best-practice lists built on a misread 40% figure. The industry needs a standard of proof, not more tactics.