Obsurfable

Semantic HTML Doubles AI Citation Rates: What the Beyond the Blue Link Study Found

Obsurfable

The GEO checklist focuses on statistics, citations, and FAQ blocks. A paper presented at the 37th ACM Conference on Hypertext in September 2026 argues the deeper lever is authored hypertext structure — semantic markup, anchor boundaries, and answer-unit architecture.

The core result from controlled retrieval injection on open-weights models: structured HTML yields 2.6× higher citation rates and 4.0× higher extraction fidelity than equivalent unstructured content, with measurably lower attention entropy. Community corpus placement increases citation probability for subjective queries by ΔgSoV = 18.2 percentage points (p < 0.001), independent of content quality.

For brands trying to earn AI citations, the message is structural: how you mark up content determines whether a retrieved page survives synthesis — or gets rejected.

The AEO gap: retrieved but not cited

Traditional SEO optimizes Stage 1 of the generative information retrieval pipeline: getting into the retrieval context. Answer Engine Optimization (AEO) targets Stage 2: maximizing the probability that a retrieved document survives synthesis and receives attribution.

The paper formalizes the failure mode as Synthesis Rejection:

SRR(entity) = 1 − P(cited | retrieved)

A document can be in the retrieval context but never appear in the output. Reasons include low semantic coherence, redundancy with other chunks, poor formatting that raises extraction cost, or safety filters.

The gap between being retrieved and being cited is where most GEO tactics either help or waste effort. Indexably's research found page-level signals matter most for high-authority domains — but within the retrieval pool, structure determines absorption.

Generative Share of Voice (gSoV)

Because GenIR is stochastic (temperature > 0), the paper defines visibility as an expected citation probability across semantically related queries:

gSoV(entity, intent cluster) = E[citation indicator]

Approximated via K = 10 independent inference runs per query, with N = 200 queries per cluster. Standard error < 0.01 under Bernoulli sampling. All estimates reported with 95% Wilson score confidence intervals.

This is the measurement discipline the field needs: repeated sampling, explicit confidence intervals, and separation of retrieval from citation. It aligns with How to Measure AI Visibility Without Fooling Yourself.

Citation-Centric Alignment (CCA): three interventions

The paper's treatment framework, Citation-Centric Alignment, consists of three design interventions deployed in staggered two-month windows across a 12-month longitudinal study (January 2025–January 2026):

1. Intent Cluster Modeling (ICM)

Consolidate content into "Listicle Nodes" designed as embedding cluster centroids — one well-optimized page satisfying hundreds of lexically distinct queries with the same information need.

2. Structured Answer Architecture (SAA)

Decompose content into atomic Answer Units on HTML pages:

  • Header-answer pairing (question + answer in < 60 words)
  • Definition lists (<dl> / <dt> / <dd>)
  • Inverse pyramid structure (direct answer at DOM top)

3. Community Signal Amplification

For subjective queries, authentic community participation on domain-relevant platforms. The paper treats this primarily as a phenomenon to study rather than a strategy to deploy at scale — with controlled evidence in source substitution experiments.

The SAA result: structure beats unstructured PDFs

For factual queries, the primary failure mode was extraction failure from dense, unstructured PDFs. After SAA deployment:

MetricUnstructured baselineAfter SAA
Synthesis Rejection Rate89%12%
Citation rate (controlled injection)Baseline2.6× higher
Extraction fidelityBaseline4.0× higher
Attention entropyH = 3.87H = 2.14

Lower attention entropy means the model spends less effort resolving ambiguous structure — explicit anchors reduce runtime resolution failure.

Dexter Reference Model formalization

Each Answer Unit maps to a Dexter component with explicit internal structure:

  • Heading defines a question anchor
  • Content defines an answer anchor
  • Wrapper marks the component boundary

These anchors are self-describing: semantic HTML tags signal the question-answer relationship without requiring the consumer to parse surrounding context.

In controlled injection experiments:

Component structureRuntime resolution success
Explicit anchor structure91.2%
No anchors34.8% (65.2% failure)

The GenIR synthesis stage is, functionally, a machine-mediated Dexter runtime layer. Well-anchored components exhibit lower traversal failure.

This connects to the first 30% of a page being where AI looks — SAA puts the direct answer at DOM top, not buried in paragraph three.

Longitudinal gSoV lifts

Pre- versus post-intervention results from the 12-month field study:

SubjectIntentgSoV pregSoV postLift
Travel (subjective)Recommendation4.2%28.5%+578%
TravelInformational8.1%14.3%+76%
Food (factual)Regulatory2.5%38.1%+1424%
FoodComparative12.0%41.2%+243%

SERP rankings remained static: ρ(ΔgSoV, ΔRank) = −0.14, p = 0.31. AEO operated as a distinct optimization vector from traditional search rank.

Staggered deployment provided temporal evidence: the food consultancy inflected at month 6 (SAA deployment); the travel startup at month 8 (community signal). No cross-contamination between subjects.

Source substitution: community vs static vs authoritative

To isolate placement from content quality, the researchers created content-identical pages for 6 synthetic entities across three conditions:

ConditionDescription
CommunityReddit posts, 15–30 upvotes, 3–5 comments
StaticHTML pages, DA ≈ 25
AuthoritativeInstitutional pages, DA ≈ 60

Same claims, tone, word count, and HTML structure. Only source domain differed. N = 600 queries, K = 10 runs, across ChatGPT and Perplexity.

Subjective queries

Source conditiongSoV95% CI
Community26.4%[23.1, 29.7]
Static8.2%[6.4, 10.0]
Authoritative12.1%[9.8, 14.4]

Community placement adds ΔgSoV = 18.2 pp over static (p < 0.001) — independent of content quality. This aligns with earned media driving 84% of AI citations and Reddit citations concentrating at bottom-of-funnel.

For subjective queries, logistic regression on citation probability yielded α = 0.82 for community signal intensity versus β = 0.18 for Domain Authority.

Factual queries

Source conditiongSoV95% CI
Authoritative18.3%—
Community14.8%—
Static11.2%—

For factual queries, structured architecture mattered more than community placement. Authority still helps — but SAA closed most of the gap between static and authoritative for regulatory content.

Cross-platform agreement is low

Jaccard similarity of top-5 cited entities across engines:

Engine pairJaccard
ChatGPT–Perplexity0.24
ChatGPT–Gemini0.31
Perplexity–Gemini0.28

Different agentic mediators traverse the same hypertext graph and arrive at radically different compositions. Multi-engine evaluation is not optional — consistent with cross-engine divergence research.

What to implement: SAA checklist

Based on the paper's Structured Answer Architecture:

Do

  1. Pair every heading with a direct answer in < 60 words immediately below.
  2. Use definition lists (<dl>) for term-definition content — not bold text in paragraphs.
  3. Lead with the answer at DOM top (inverse pyramid); context and nuance follow.
  4. Mark component boundaries with semantic wrappers (<section>, <article>, <aside>).
  5. Publish in HTML, not PDF, for content you want cited. PDF SRR was 89% in the study.
  6. Keep answer units atomic — one question, one answer, one citation opportunity.

Avoid

  1. Dense unstructured paragraphs without heading anchors.
  2. PDF-only regulatory or product content.
  3. FAQ blocks that bury answers in accordion markup without semantic structure.
  4. Assuming community placement substitutes for structure on factual queries.
  5. Treating SERP rank as a proxy for generative citation.

Validate

Run 10 repeated inference runs per query on your target engine. Compare gSoV before and after structural changes. Single runs are insufficient — see volatility research.

How this interacts with authority and exposure

SAA optimizes conditional citation — what happens after a page enters the retrieval context. It does not substitute for:

Sequence the work:

  1. Challenger brands: Earn exposure through distribution and third-party mentions first. Then apply SAA so retrieved pages survive synthesis.
  2. High-authority brands: SAA is high-ROI immediately — you are already in the retrieval pool.
  3. All brands: Publish decision-stage content as structured HTML answer units, not PDFs or unstructured blog posts.

How Obsurfable fits

Obsurfable measures whether structural changes move citation and mention rates on real buyer prompts — not synthetic benchmarks.

After deploying SAA on product pages, run your fixed prompt panel and compare citation rate and mention rate over 4–8 weeks. If exposure is flat, structure is not the bottleneck. If exposure is high but citation rate is low, SAA targets the right stage.

The public corpus lets you see which pages AI engines actually cite in your category — and whether competitors earn citations through structure, community placement, or authority.

FAQ

Is this just "add schema"?

No. The paper's strongest differentiator is semantic HTML anchor structure — heading-answer pairs, definition lists, component boundaries — not JSON-LD schema. Schema may help consistency; SAA targets synthesis rejection directly.

Should I move all content off PDF?

For pages you want AI to cite, yes. PDF SRR was 89% in the study. Keep PDFs for download; publish the same content as structured HTML for retrieval.

Does community seeding mean astroturfing Reddit?

The paper explicitly flags ethical concerns and presents community signal as a phenomenon to detect, not a tactic to deploy at scale. Authentic participation on surfaces AI already cites is the defensible version.

How does gSoV relate to mention rate?

gSoV measures citation attribution — explicit hyperlinked or named source credit. Mention without citation is a separate signal. Track both.

Bottom line

Structured HTML yields 2.6× higher citation rates and 4.0× higher extraction fidelity than unstructured content in controlled conditions. Community placement adds 18.2 percentage points for subjective queries. SERP rank did not move — generative citation is a distinct optimization vector.

Ship answer-first HTML with semantic anchors. Stop publishing citation-critical content as PDF. Measure gSoV with repeated runs, not single checks.

The hypertext structure you author determines whether machine mediators can extract, synthesize, and attribute your content — or silently reject it.