Short answer: When the same buyer-style prompt is run through ChatGPT, Gemini, and Claude, the answers differ in structure as much as in content. In Obsurfable's research corpus, 91% of ChatGPT answers on matched three-way prompts include zero cited sources. Claude cites four or more on 58% of the same prompts. Gemini sits between them — and when Gemini's web search is active, average citation count jumps from 0.2 to 5.6 per answer on matched ChatGPT–Gemini pairs.
This is not a ranking disagreement. It is a citation anatomy gap: some engines answer from training memory with brand lists but no links; others attach a source bundle on the same question.
What we analysed
We used Obsurfable Explorer's research corpus: buyer-style prompts tracked across multiple AI engines. For each prompt we keep the latest observation per platform family — ChatGPT, Gemini, Claude, Perplexity, and Grok — preferring web-app captures over API runs when both exist.
| Scope | Value |
|---|---|
| Corpus window | 29 June – 15 September 2026 |
| Active research prompts | 31,166 |
| Total observations (research subset) | 35,807 |
| Prompts with 2+ platform families | 900 |
| Three-way panel (ChatGPT + Gemini + Claude) | 531 |
| Four-way panel (+ Perplexity) | 196 |
| ChatGPT + Grok matched subset | 265 |
Primary metrics: source count (URLs attached to the answer) and brand count (named products extracted from response text). Categories in the three-way panel skew toward B2B software: data infrastructure (202 prompts), developer tools (187), and developer media (140).
Browse live examples on Obsurfable Explorer.
What we found
1. ChatGPT answers are overwhelmingly citation-free on matched prompts
On the 531-prompt three-way panel, source attachment diverges sharply:
| Cited sources in answer | ChatGPT | Gemini | Claude |
|---|---|---|---|
| 0 sources | 483 (91.0%) | 287 (54.0%) | 192 (36.2%) |
| 1–3 sources | 5 (0.9%) | 7 (1.3%) | 30 (5.6%) |
| 4+ sources | 43 (8.1%) | 237 (44.6%) | 309 (58.2%) |
Average sources per answer on the same prompts:
| Engine | Avg sources | Avg brands named |
|---|---|---|
| ChatGPT | 0.70 | 9.29 |
| Gemini | 2.64 | 10.33 |
| Claude | 2.98 | 10.17 |
ChatGPT still names brands on most answers — but without links. On 38.8% of ChatGPT–Gemini matched pairs (236/609), ChatGPT cites nothing while Gemini attaches at least one source on the identical prompt.
2. Search-enabled engines pack citations; memory-first engines do not
The gap tracks whether the engine ran a live web search:
| Condition | n | Avg sources | Avg brands |
|---|---|---|---|
| Gemini with web search (matched ChatGPT+Gemini prompts) | 273 | 5.64 | 12.4 |
| Gemini without web search (same panel) | 32 | 0.16 | 10.4 |
| Perplexity (search-native, four-way panel) | 196 | 4.20 | 12.9 |
| Grok (ChatGPT+Grok matched subset) | 265 | 8.29 | 8.9 |
When Gemini's search is off, its citation profile collapses toward ChatGPT's. When search is on, citation density approaches Claude and Grok — even though brand counts stay high in both modes.
Across the full research corpus (not just matched panels), the platform-level split is starker:
| Engine | Observations | Avg sources | % answers with zero sources |
|---|---|---|---|
| ChatGPT | 31,158 | 0.02 | 99.8% |
| Gemini | 616 | 2.54 | 56.3% |
| Claude | 579 | 2.99 | 36.4% |
| Perplexity | 319 | 5.72 | 48.0% |
| Grok | 266 | 8.27 | 0.8% |
Most ChatGPT observations in the corpus come from API runs (which do not attach citations). UI captures show a higher citation rate — but the matched-prompt panel above already controls for question identity, and the citation gap persists.
3. Brand lists inflate slightly when citations appear — but the bigger shift is structural
On the three-way panel, empty-brand rates are low for search-heavy engines:
| Engine | Answers with zero brands | Answers with 6+ brands |
|---|---|---|
| ChatGPT | 10.2% | 67.4% |
| Gemini | 0.8% | 85.7% |
| Claude | 1.1% | 82.5% |
Gemini names more brands than ChatGPT on 60.8% of matched ChatGPT–Gemini pairs (avg delta +1.6 brands). That is modest compared with the source delta (+1.9 sources on average, and +4.9 when Gemini's web search is active).
The practical split: ChatGPT can return a vendor shortlist with no traceable URLs; Claude and Grok typically return a shortlist and a link bundle on the same prompt.
4. Category context: data-infrastructure prompts show the widest citation split
On matched ChatGPT–Gemini prompts with at least 20 observations per category:
| Category | n | ChatGPT avg sources | Gemini avg sources | Gemini cites more (share) |
|---|---|---|---|---|
| Technology · data infrastructure | 222 | 1.64 | 4.03 | 58% |
| Technology · developer tools | 216 | 0.05 | 2.64 | 50% |
| Media · developer media | 141 | 0.18 | 0.19 | 4% |
Data-infrastructure and developer-tool prompts — proxy providers, SERP APIs, scraping stacks — trigger the largest citation gaps. Developer-media prompts (newsletters, blogs) show near-parity: both engines often answer from memory without attaching links.
Worked examples (same prompt, different citation anatomy)
These are live corpus records.
How can companies source data for AI development? — explorer.obsurfable.com/prompts/how-can-companies-source-data-for-ai-development-dbb0014b
| Engine | Cited sources | Brands named |
|---|---|---|
| ChatGPT | 0 | 8 |
| Gemini | 14 | 12 |
| Claude | 5 | 11 |
ChatGPT names vendors; Gemini and Claude attach link lists on the same question.
Can you recommend a LinkedIn scraper that bypasses restrictions? — explorer.obsurfable.com/prompts/can-you-recommend-a-linkedin-scraper-that-bypasses-restrictions-65a767de
| Engine | Cited sources | Brands named |
|---|---|---|
| ChatGPT | 0 | 7 |
| Gemini | 10 | 13 |
| Claude | 7 | 10 |
What are the best email APIs for transactional messages? — explorer.obsurfable.com/prompts/what-are-the-best-email-apis-for-transactional-messages-bc33cb33
| Engine | Cited sources | Brands named |
|---|---|---|
| ChatGPT | 0 | 9 |
| Grok | 12 | 8 |
On the ChatGPT–Grok subset, Grok averages 8.3 sources per answer vs 0.7 for ChatGPT on the same prompts.
Category hubs with the densest multi-engine panels: data infrastructure, developer tools.
Why this is surprising
Domain-citation leaderboards (Ahrefs Brand Radar, third-party trackers) already show that which sites get linked varies by engine. This finding is different: on identical buyer prompts, the presence and density of citations varies just as much as the brand shortlist — and the two dimensions are only loosely coupled.
Three implications for measurement:
-
Brand mention tracking ≠ citation tracking. A prompt can yield eight named vendors on ChatGPT with zero URLs, and twelve vendors plus fourteen links on Gemini. A program that counts mentions will score both engines as "active"; a program that counts citations will score only one.
-
"Visible in the answer" is not the same as "visible in the evidence." ChatGPT's brand lists are often untraceable to a source bundle in our corpus. Claude and Grok answers, by contrast, routinely attach four or more links — making downstream domain analysis possible on the same prompt.
-
Web search is the mechanism, not the whole story. Gemini without search looks like ChatGPT (0.2 avg sources on matched prompts). Gemini with search looks like Claude. Perplexity and Grok are search-native and sit at the high-citation end. The anatomy of an AI answer is largely determined by whether the engine retrieved the web — not just which model weights sit behind it.
This complements our cross-platform brand shortlist analysis: engines disagree on who to recommend and on whether to show their work.
Limitations
- Panel size: Multi-platform coverage is sampled — 900 prompts with 2+ engines, not all 31k. Technology and developer-media verticals dominate early corpus work.
- ChatGPT API weight: Most ChatGPT observations are API runs, which do not attach citations. The matched-prompt panel controls for question identity but reflects how we capture each engine today.
- Latest observation only: We compare the most recent capture per engine, not time-series drift on the same prompt.
- Source extraction: Counts reflect URLs parsed from model outputs. Inline references, footnotes, or post-hoc UI links may be undercounted.
- Brand extraction: Lists reflect Obsurfable's entity extraction. Generic terms (e.g. "REST APIs") may appear instead of vendors.
- Perplexity bimodality: On the four-way panel, Perplexity shows a high zero-source rate (66%) alongside a high average (4.2) — it often attaches many links or none. Reported averages mask that split.
Methodology
Corpus filter: Active research prompts only — buyer-style questions not scoped to a single customer website.
Platform families: Raw platform values mapped to public families per Explorer taxonomy — e.g. OpenAI API and ChatGPT web app → ChatGPT; Gemini web app and Gemini API → Gemini; Claude web app and Claude API → Claude.
Deduplication: For each (prompt, platform family) pair, keep the latest observation; prefer web-app captures over API runs at equal timestamps.
Metrics: Source count = number of URLs attached to the answer. Brand count = entries in the extracted "brands mentioned" list. Matched panels require the same prompt ID observed on each engine in the comparison.
Exclusions: Observations with unmapped platform keys omitted from family-level joins. Customer-scoped prompts excluded.
Reproduce or explore individual prompts at explorer.obsurfable.com. Signed-in users can view observation history per prompt.
Conclusion
On matched buyer prompts in Obsurfable's corpus, AI engines ship fundamentally different answer shapes. ChatGPT cites nothing on nine out of ten three-way panel answers; Claude attaches four or more sources on six out of ten. Gemini's citation profile depends almost entirely on whether web search ran. If your visibility program tracks brand mentions alone — or citations from a single engine — you are measuring one layer of a multi-layer answer. Track both who gets named and whether the engine shows its sources, per platform, on matched prompts.