The GEO industry sells page rewrites as the path to AI citations: add FAQs, inject statistics, chunk for extractability, ship schema. Some of that helps — but mostly if you already have authority.
Indexably's study of 18,129 AI-cited pages across five platforms found that domain-level factors account for 77% of predictive importance in their combined model. Page factors account for 23%. When they split pages by domain authority, page-level signals only produced measurable lift in the top quartile. For everyone else, effects were flat or slightly negative.
Page optimization compounds with authority. It does not substitute for it.
What Indexably measured
The team sent 500 prompts across 12 topic categories to ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews. They collected 25,115 citations pointing to 18,129 unique pages.
For comparison, they pulled 4,622 control pages from Brave Search — pages that rank well for the same queries but were not cited by any AI model. That control design matters: it separates "ranks in search" from "gets cited by AI."
They extracted 213 structural signals from every page's HTML and analyzed 2,000 domains (1,000 cited, 1,000 control) for referring domains, backlink diversity, and organic visibility. Then they combined page and domain signals in one analysis.
Full methods are published in the Indexably Method. The authors are explicit: this is correlation, not causation — and their combined model still leaves most of citation unexplained (McFadden pseudo R² ≈ 0.10). Content relevance — whether the page answers the specific question — is the likely missing majority.
Finding 1: Domain authority dominates
In Indexably's logistic regression combining page and domain signals:
| Factor group | Share of predictive importance |
|---|---|
| Domain-level | 77% |
| Page-level | 23% |
"Domain-level" here means measured footprint signals — referring domains, backlink diversity, organic keyword coverage, Domain Rank (via DataForSEO). The authors treat these as proxies for broader entity authority: brand recognition, topical association, and trust built across the web.
DataForSEO Rank showed an effect size of d = 1.075 — roughly 7× stronger than any page signal in their reporting. That matches the direction of Discovered Labs' analysis of 2M citations across 10K pages, where AI-perceived domain authority was about 6× as influential as the strongest non-alignment page-level feature.
The uncomfortable translation: for most sites, rewriting the intro is rearranging deck chairs while the retrieval gate is domain trust.
Finding 2: Page optimization only lifts the top quartile
Indexably split pages into four groups by domain authority. Page-level signals showed positive effects only in the top quartile (d = 0.24–0.37). Lower quartiles: flat or slightly negative.
| Authority quartile | Effect of page optimization on citation odds |
|---|---|
| Top quartile | Positive, measurable |
| Lower three quartiles | Flat or slightly negative |
This is the finding that should reset GEO roadmaps.
- Incumbents / high-authority domains: Extractable structure, clean HTML, and answer-first formatting are high-ROI — they compound existing retrieval eligibility.
- Challenger / mid-authority domains: The same checklist may produce no measurable citation lift until off-site authority moves. Prioritize distribution, earned mentions, and link diversity before another FAQ rewrite sprint.
That does not mean challengers should ship broken pages. It means the bottleneck is usually upstream of on-page GEO tactics.
Finding 3: Backlink diversity beats backlink volume
| Signal | Effect size (Cohen's d) |
|---|---|
| Referring subnets (unique network blocks) | 0.513 |
| Referring domain count | 0.235 |
| Total backlink count | 0.098 |
Referring subnets were more than twice as predictive as raw referring domain count. Total backlink count barely registered.
It is not how many links you have. It is how many independent corners of the internet endorse you. A hundred links from one blog network is weaker than fewer links from diverse, unrelated neighborhoods of the web.
This aligns with Victorious's Q2 2026 finding that referring domains (r = 0.49) and third-party mentions (r = 0.45) correlate with AI mention rates more than owned-site metrics — see AI Recognizes 96% of Brands — Then Mentions Almost None.
Finding 4: Boring HTML beats trendy GEO checklists
The strongest page-level differentiators were not schema, FAQ blocks, or word count. They were fundamentals:
- Proper doctype
- Language declaration
- Canonical tag
- Viewport meta
- Meta description
Word count showed a slight negative correlation with citation. Cited pages had higher vocabulary diversity, shorter paragraphs, and shorter maximum sections. The pattern is "write clearly and structure well," not "write more."
Schema and FAQ formats are not useless — other studies find lifts in specific contexts — but in Indexably's head-to-head against ranking-but-uncited controls, basic technical hygiene differentiated more consistently than the fashionable GEO checklist items.
Finding 5: Each model weights factors differently
Indexably's per-platform breakdown (directionally):
| Platform | Relative preference |
|---|---|
| ChatGPT | Freshness and structured metadata |
| Claude | Weight spread more evenly |
| Gemini | Crawlability |
| Google AI Overviews | Most balanced |
| Perplexity | Distinct mix (freshness-heavy in other studies) |
Optimizing for one model may not help another. That matches the broader finding that only 2.7% of domains are cited by all five engines and that each engine reads a different search index.
What this does not mean
It does not mean on-page work is pointless. For high-authority domains, page structure is a compounding lever. For everyone, crawlable HTML is table stakes for retrieval eligibility.
It does not prove canonical tags cause citations. Well-maintained sites have canonicals and authority. Indexably is careful about that confound.
It does not replace relevance. The authors' own R² says most of what drives citation was not in their structural or domain features. Answering the prompt still wins.
It does not contradict extractability research. Findings like citations concentrating in the first 30% of a page still matter — especially once you are in the retrieval pool. Indexably's point is about who gets into that pool, not how absorption works after selection.
Practical playbook by authority tier
If you are a high-authority domain (top quartile)
- Fix technical fundamentals (doctype, language, canonical, meta, crawlability).
- Lead with direct, extractable answers; shorten dense sections.
- Match structure to platform preferences (freshness for ChatGPT/Perplexity; crawlability for Gemini).
- Measure citation lift after changes with repeated prompt runs — not one-off checks.
If you are a mid- or low-authority domain
- Stop expecting FAQ rewrites alone to move AI citation rate.
- Invest in referring-subnet diversity: relevant publishers, communities, review platforms, niche sites — not PBN-style volume.
- Grow third-party brand mentions on surfaces AI already cites in your category (earned media, Reddit at BOFU).
- Keep pages clean and answer-first so that when authority arrives, structure is ready to compound.
- Track category mention rate and competitor citation share — not vanity "we rewrote 40 pages" metrics.
For every team
Separate retrieval eligibility (can the engine find and trust the domain?) from citation absorption (does the page contribute extractable evidence once retrieved?). Indexably speaks primarily to the first. Most GEO checklists speak to the second. You need both — sequenced correctly.
How Obsurfable fits
Obsurfable measures what page checklists cannot: whether your brand actually appears in AI answers for buyer prompts, how often competitors displace you, and which sources the engines cite instead.
If you are mid-authority and rewriting pages without citation movement, Obsurfable makes that failure visible early — so budget shifts to distribution and entity footprint instead of another on-page cycle. If you are high-authority, it shows which prompts still miss despite structural work.
The Visibility Director flags prompt-level gaps and drafts content aimed at the retrieval and citation layers that are actually failing.
For volatility in citation sets, see AI Citation Volatility. For what the July 2026 evidence review says about rewrite tradeoffs, see GEO Evidence Review 2026.
FAQ
Should challenger brands ignore on-page GEO entirely?
No. Ship crawlable, clear, answer-first pages. Just do not treat on-page GEO as the primary growth lever until domain footprint and third-party mentions improve.
Is schema still worth implementing?
Yes for product, organization, and FAQ consistency — and because other datasets show lifts. Indexably's result says schema was not the strongest differentiator versus ranking-but-uncited controls, not that schema is harmful.
How do I know which authority quartile I am in?
Compare referring domains, organic visibility, and third-party mention volume against cited competitors in your category. If competitors with similar content get cited and you do not, authority and footprint are the first hypotheses — not heading tags.
Does this conflict with Google's "GEO is still SEO" guidance?
No. Google's May 2026 guide says Google AI features rest on core Search quality systems — which is a domain-and-content authority story. Indexably's domain dominance is consistent with that framing for Google surfaces, while still leaving room for platform-specific retrieval elsewhere.
Bottom line
Across 18,129 AI-cited pages, domain factors explain 77% of what Indexably could predict. Page optimization only measurably helps the top authority quartile. Backlink diversity beats volume. Boring HTML beats trendy checklists. Challengers need off-site footprint first; incumbents get compounding returns from structure. Measure citation outcomes — then sequence the work to the bottleneck you actually have.