Obsurfable

The citation gap: same prompt, different answer anatomy

Obsurfable

Short answer: When the same buyer-style prompt is run through ChatGPT, Gemini, and Claude, the answers differ in structure as much as in content. In Obsurfable's research corpus, 91% of ChatGPT answers on matched three-way prompts include zero cited sources. Claude cites four or more on 58% of the same prompts. Gemini sits between them — and when Gemini's web search is active, average citation count jumps from 0.2 to 5.6 per answer on matched ChatGPT–Gemini pairs.

This is not a ranking disagreement. It is a citation anatomy gap: some engines answer from training memory with brand lists but no links; others attach a source bundle on the same question.


What we analysed

We used Obsurfable Explorer's research corpus: buyer-style prompts tracked across multiple AI engines. For each prompt we keep the latest observation per platform family — ChatGPT, Gemini, Claude, Perplexity, and Grok — preferring web-app captures over API runs when both exist.

ScopeValue
Corpus window29 June – 15 September 2026
Active research prompts31,166
Total observations (research subset)35,807
Prompts with 2+ platform families900
Three-way panel (ChatGPT + Gemini + Claude)531
Four-way panel (+ Perplexity)196
ChatGPT + Grok matched subset265

Primary metrics: source count (URLs attached to the answer) and brand count (named products extracted from response text). Categories in the three-way panel skew toward B2B software: data infrastructure (202 prompts), developer tools (187), and developer media (140).

Browse live examples on Obsurfable Explorer.


What we found

1. ChatGPT answers are overwhelmingly citation-free on matched prompts

On the 531-prompt three-way panel, source attachment diverges sharply:

Cited sources in answerChatGPTGeminiClaude
0 sources483 (91.0%)287 (54.0%)192 (36.2%)
1–3 sources5 (0.9%)7 (1.3%)30 (5.6%)
4+ sources43 (8.1%)237 (44.6%)309 (58.2%)

Average sources per answer on the same prompts:

EngineAvg sourcesAvg brands named
ChatGPT0.709.29
Gemini2.6410.33
Claude2.9810.17

ChatGPT still names brands on most answers — but without links. On 38.8% of ChatGPT–Gemini matched pairs (236/609), ChatGPT cites nothing while Gemini attaches at least one source on the identical prompt.

2. Search-enabled engines pack citations; memory-first engines do not

The gap tracks whether the engine ran a live web search:

ConditionnAvg sourcesAvg brands
Gemini with web search (matched ChatGPT+Gemini prompts)2735.6412.4
Gemini without web search (same panel)320.1610.4
Perplexity (search-native, four-way panel)1964.2012.9
Grok (ChatGPT+Grok matched subset)2658.298.9

When Gemini's search is off, its citation profile collapses toward ChatGPT's. When search is on, citation density approaches Claude and Grok — even though brand counts stay high in both modes.

Across the full research corpus (not just matched panels), the platform-level split is starker:

EngineObservationsAvg sources% answers with zero sources
ChatGPT31,1580.0299.8%
Gemini6162.5456.3%
Claude5792.9936.4%
Perplexity3195.7248.0%
Grok2668.270.8%

Most ChatGPT observations in the corpus come from API runs (which do not attach citations). UI captures show a higher citation rate — but the matched-prompt panel above already controls for question identity, and the citation gap persists.

3. Brand lists inflate slightly when citations appear — but the bigger shift is structural

On the three-way panel, empty-brand rates are low for search-heavy engines:

EngineAnswers with zero brandsAnswers with 6+ brands
ChatGPT10.2%67.4%
Gemini0.8%85.7%
Claude1.1%82.5%

Gemini names more brands than ChatGPT on 60.8% of matched ChatGPT–Gemini pairs (avg delta +1.6 brands). That is modest compared with the source delta (+1.9 sources on average, and +4.9 when Gemini's web search is active).

The practical split: ChatGPT can return a vendor shortlist with no traceable URLs; Claude and Grok typically return a shortlist and a link bundle on the same prompt.

4. Category context: data-infrastructure prompts show the widest citation split

On matched ChatGPT–Gemini prompts with at least 20 observations per category:

CategorynChatGPT avg sourcesGemini avg sourcesGemini cites more (share)
Technology · data infrastructure2221.644.0358%
Technology · developer tools2160.052.6450%
Media · developer media1410.180.194%

Data-infrastructure and developer-tool prompts — proxy providers, SERP APIs, scraping stacks — trigger the largest citation gaps. Developer-media prompts (newsletters, blogs) show near-parity: both engines often answer from memory without attaching links.


Worked examples (same prompt, different citation anatomy)

These are live corpus records.

How can companies source data for AI development?explorer.obsurfable.com/prompts/how-can-companies-source-data-for-ai-development-dbb0014b

EngineCited sourcesBrands named
ChatGPT08
Gemini1412
Claude511

ChatGPT names vendors; Gemini and Claude attach link lists on the same question.

Can you recommend a LinkedIn scraper that bypasses restrictions?explorer.obsurfable.com/prompts/can-you-recommend-a-linkedin-scraper-that-bypasses-restrictions-65a767de

EngineCited sourcesBrands named
ChatGPT07
Gemini1013
Claude710

What are the best email APIs for transactional messages?explorer.obsurfable.com/prompts/what-are-the-best-email-apis-for-transactional-messages-bc33cb33

EngineCited sourcesBrands named
ChatGPT09
Grok128

On the ChatGPT–Grok subset, Grok averages 8.3 sources per answer vs 0.7 for ChatGPT on the same prompts.

Category hubs with the densest multi-engine panels: data infrastructure, developer tools.


Why this is surprising

Domain-citation leaderboards (Ahrefs Brand Radar, third-party trackers) already show that which sites get linked varies by engine. This finding is different: on identical buyer prompts, the presence and density of citations varies just as much as the brand shortlist — and the two dimensions are only loosely coupled.

Three implications for measurement:

  1. Brand mention tracking ≠ citation tracking. A prompt can yield eight named vendors on ChatGPT with zero URLs, and twelve vendors plus fourteen links on Gemini. A program that counts mentions will score both engines as "active"; a program that counts citations will score only one.

  2. "Visible in the answer" is not the same as "visible in the evidence." ChatGPT's brand lists are often untraceable to a source bundle in our corpus. Claude and Grok answers, by contrast, routinely attach four or more links — making downstream domain analysis possible on the same prompt.

  3. Web search is the mechanism, not the whole story. Gemini without search looks like ChatGPT (0.2 avg sources on matched prompts). Gemini with search looks like Claude. Perplexity and Grok are search-native and sit at the high-citation end. The anatomy of an AI answer is largely determined by whether the engine retrieved the web — not just which model weights sit behind it.

This complements our cross-platform brand shortlist analysis: engines disagree on who to recommend and on whether to show their work.


Limitations

  • Panel size: Multi-platform coverage is sampled — 900 prompts with 2+ engines, not all 31k. Technology and developer-media verticals dominate early corpus work.
  • ChatGPT API weight: Most ChatGPT observations are API runs, which do not attach citations. The matched-prompt panel controls for question identity but reflects how we capture each engine today.
  • Latest observation only: We compare the most recent capture per engine, not time-series drift on the same prompt.
  • Source extraction: Counts reflect URLs parsed from model outputs. Inline references, footnotes, or post-hoc UI links may be undercounted.
  • Brand extraction: Lists reflect Obsurfable's entity extraction. Generic terms (e.g. "REST APIs") may appear instead of vendors.
  • Perplexity bimodality: On the four-way panel, Perplexity shows a high zero-source rate (66%) alongside a high average (4.2) — it often attaches many links or none. Reported averages mask that split.

Methodology

Corpus filter: Active research prompts only — buyer-style questions not scoped to a single customer website.

Platform families: Raw platform values mapped to public families per Explorer taxonomy — e.g. OpenAI API and ChatGPT web app → ChatGPT; Gemini web app and Gemini API → Gemini; Claude web app and Claude API → Claude.

Deduplication: For each (prompt, platform family) pair, keep the latest observation; prefer web-app captures over API runs at equal timestamps.

Metrics: Source count = number of URLs attached to the answer. Brand count = entries in the extracted "brands mentioned" list. Matched panels require the same prompt ID observed on each engine in the comparison.

Exclusions: Observations with unmapped platform keys omitted from family-level joins. Customer-scoped prompts excluded.

Reproduce or explore individual prompts at explorer.obsurfable.com. Signed-in users can view observation history per prompt.


Conclusion

On matched buyer prompts in Obsurfable's corpus, AI engines ship fundamentally different answer shapes. ChatGPT cites nothing on nine out of ten three-way panel answers; Claude attaches four or more sources on six out of ten. Gemini's citation profile depends almost entirely on whether web search ran. If your visibility program tracks brand mentions alone — or citations from a single engine — you are measuring one layer of a multi-layer answer. Track both who gets named and whether the engine shows its sources, per platform, on matched prompts.