The GEO industry has spent two years selling structured content as the path to AI citations: add headings, break answers into lists, ship tables, chunk for extractability. A new causal study suggests that advice is half right — and dangerously misread if you apply it to the wrong bottleneck.
CITECHOICE, submitted to arXiv on 14 September 2026, is the first controlled audit of citation allocation in authentic multi-turn agentic search. The researchers froze 129 everyday-query transcripts from a tool-using GPT-5.4 agent, then replayed each conversation with exactly one change: how a competing document was rendered (structured vs prose) or where it appeared in the source list.
The headline: structured rendering concentrates citation credit — it raises target citation count by +0.50 per answer (95% CI [+0.20, +0.84]; Holm-adjusted p = .033). But it does not reliably increase whether the target is cited at all. The pre-specified incidence effect is +4.5 percentage points and statistically inconclusive (95% CI [-1.4, +10.4]; p = .168).
Formatting helps you win credit among sources already in contention. It does not clearly get you into contention.
What CITECHOICE measured
Most AI citation research asks whether citations are correct — does the cited page support the claim? CITECHOICE asks a different question: when several retrieved pages could support the same fact, which one gets the citation?
The protocol:
| Step | Detail |
|---|---|
| Acquisition | A tool-using GPT-5.4 agent answered 130 human-style queries with its own web searches; every message and result object was archived |
| Pair selection | From 129 transcripts, researchers identified 113 same-call document pairs where both documents independently supported the same pre-specified fact — selected outcome-blind, without seeing ranks or citation outcomes |
| Human review | Blinded adjudication confirmed 103 of 113 as strict same-proposition competitions |
| Intervention | A hash-verified 2×2 replay crossing pair order with structured vs prose rendering of one target document; everything else in the transcript held byte-identical |
| Trials | 452 replay trials across 89 independent answer-target families |
The intervention acts on the serialized source snapshot the answer model actually reads — downstream of retrieval and extraction. Among 1,750 archived result-text records, 93.2% contained markdown-style headings, 76.6% contained list markers, and 21.3% contained table-like rows. Extraction and serialization decide which document structure survives into model context.
This is not a synthetic two-document RAG lab. These are real agentic search transcripts with real page snapshots.
Finding 1: Structured rendering concentrates credit, not admission
When two documents support the same fact and both are already retrieved, switching one from prose to structured rendering:
| Outcome | Effect | Significance |
|---|---|---|
| Target citation count (how many times cited in the answer) | +0.50 per answer | 95% CI [+0.20, +0.84]; Holm-adjusted p = .033 |
| Target citation incidence (cited at all vs not) | +4.5 pp | 95% CI [-1.4, +10.4]; p = .168 — inconclusive |
| Total citations in the answer | No increase | Structured formatting did not add more citations overall |
| Competitor citation credit | Unchanged | Credit moved within the pair, not away from unrelated sources |
About half the count gain attached to shared evidence — meaning structured formatting helped the target capture more of the citation markers when both documents were relevant, without expanding the total citation budget.
The practical translation: if your page is already being retrieved alongside a competitor for the same fact, clean structure may help you capture more of the visible credit. If your page is never retrieved, reformatting alone is unlikely to change that.
This aligns with Indexably's finding that page-level signals only lift citation odds in the top authority quartile — page work compounds inside the retrieval pool, not outside it.
Finding 2: Observational rank effects dwarf controlled reordering
The observational citation-rate gap between rank 1 and rank 5 documents was 42.3 percentage points (85.1% cited at rank 1 vs 42.8% at rank 5). When researchers controlled for everything else and swapped the same documents' positions in a frozen transcript, the effect shrank to +7.9 pp in the main replay and 0.0 pp in a held-out confirmation.
| Population | Rank 1 vs rank 5 citation gap |
|---|---|
| Observational (correlational) | 42.3 pp |
| Controlled pair-swap replay | +7.9 pp |
| Held-out confirmation | 0.0 pp |
Observational rank gradients can be several times larger than the allocation change supported by controlled reordering. Search engines place more relevant pages earlier — and the agent chose the queries that produced those rankings. Correlation is not mechanism.
For practitioners: "we rank #3 so we should get cited" is weaker logic than it appears. Position in the retrieved set matters less than position in organic search implies, once you control for the documents themselves.
Finding 3: Citation decisions are noisy — even with identical inputs
Even with frozen transcripts and identical prompts, 15% of binary citation decisions (cited vs not cited) flipped across fresh decoding runs. Decoding noise accounted for an estimated 45% of single-generation family-effect variance.
The aggregate count effect (+0.50) repeated under fresh decoding of 30 frozen families. But individual citation outcomes are not deterministic. One run is not a measurement.
This reinforces what AirOps found about citation volatility: only 30% of brands remained visible in back-to-back responses, and 57% of brands that disappeared from one run resurfaced in a later one. CITECHOICE adds a mechanistic explanation — part of that churn is irreducible evaluation noise, not just retrieval rotation.
What this changes about GEO advice
The advice that survives
Structured, extractable content still matters — conditionally. Headings, lists, and tables that survive extraction into the model's context can redistribute citation credit among competing sources. The first-30% concentration pattern and answer-first formatting remain relevant once you are in the retrieval pool.
Measure repeatedly. A single prompt run cannot tell you whether formatting moved citations. CITECHOICE's 15% decision-flip rate means you need longitudinal observation, not spot checks.
Separate retrieval from allocation. Retrieval is "did the engine find your page?" Allocation is "given that it found your page alongside competitors, did you get the citation?" CITECHOICE speaks to allocation. Most challenger-brand problems are retrieval problems.
The advice that needs reframing
"Add schema and FAQ blocks to earn citations" — CITECHOICE tested structured rendering of page content, not schema markup specifically. But the mechanism is the same class of bet: presentation changes inside the pool. It does not substitute for earned media and third-party mentions that put you in the pool at all.
"Rank higher to get cited" — the observational 42.3 pp rank gradient deflates to near zero under control. Ranking correlates with citation because relevance and authority correlate with both — not because position alone allocates credit.
"One formatting pass will move the needle" — at the observed variance, 89 families provide roughly 80% power only for incidence effects around 8.5 pp or larger. The measured +4.5 pp incidence effect is below that threshold. Formatting's measurable win is in citation count among admitted sources, not in breaking into the set.
Practical playbook
If you are already being retrieved but losing citations to competitors
- Audit how your pages render after extraction — not how they look in a browser. View-source and structured-text exports matter more than visual design.
- Lead with the fact that competes: put the claim AI needs in the first screen of extractable text, with a heading that matches the query's subtopic.
- Use lists and tables for comparative claims where you and a competitor both have legitimate supporting evidence.
- Track citation count and share per prompt over weeks, not single runs.
If you are not being retrieved at all
- Do not prioritize a formatting sprint. Prioritize distribution, earned mentions, and third-party coverage on surfaces the engine already cites in your category.
- Fix crawlability and entity clarity so retrieval is possible when authority arrives.
- Measure whether your domain appears in the retrieved set before investing in allocation tactics.
For every team
Run the two-question diagnostic on your top buyer prompts:
- Retrieval: Does our domain appear in the sources the engine considered?
- Allocation: When we appear alongside competitors, do we get citation credit?
CITECHOICE says presentation moves the second needle. Most GEO programs only measure the first — or conflate the two.
How Obsurfable fits
Obsurfable runs your defined buyer prompts against AI search engines on a repeated schedule — tracking whether your brand is mentioned, whether your domain is cited, and which competitor sources capture the citation instead.
When a formatting change should move allocation but does not, Obsurfable shows whether the bottleneck is retrieval (you never enter the source set) or allocation (you are retrieved but uncited). That distinction saves months of misdirected page rewrites.
The Visibility Director flags prompt-level gaps and drafts content aimed at the layer that is actually failing — distribution when retrieval is the problem, structure when allocation is.
FAQ
Does this mean structured content is useless?
No. It means structured content is an allocation lever, not a retrieval lever. Use it when you are already in the source pool. Do not expect it to substitute for authority, distribution, or earned coverage.
How is this different from the Princeton GEO paper?
The Princeton GEO framework tested synthetic prompt-response pairs with controlled content variations. CITECHOICE uses authentic agentic search transcripts with real page snapshots and causal replay. The Princeton work found formatting changes could increase visibility in controlled settings; CITECHOICE finds that in realistic agentic search, the effect concentrates credit rather than clearly expanding admission.
Should I still implement FAQ schema?
Yes for consistency, crawlability, and product clarity. CITECHOICE's finding is about what moves citations among competing retrieved sources — not that schema is harmful. Sequence the work: retrieval eligibility first, allocation optimization second.
How many prompt runs do I need to detect a formatting change?
Given CITECHOICE's 15% binary decision-flip rate and the inconclusive +4.5 pp incidence effect, plan for weeks of repeated observation across a fixed prompt panel — not a before/after snapshot on 10 queries.
Bottom line
CITECHOICE is the first causal evidence that presentation redistributes citation credit in agentic search — structured rendering adds +0.50 citations per answer when competing sources support the same fact. But it does not reliably increase whether you are cited at all (+4.5 pp, inconclusive). Observational rank effects (42.3 pp) collapse under control (+7.9 pp, then 0.0 pp held out). And 15% of citation decisions flip on identical inputs.
Format for allocation. Distribute for retrieval. Measure both — repeatedly.