Obsurfable

When AI engines disagree on brand shortlists

Obsurfable

Short answer: On the same buyer-style prompts, ChatGPT, Gemini, and Claude usually recommend different brand shortlists. In Obsurfable's research corpus, 46% of prompts observed on all three engines share zero brands across their top-five lists. ChatGPT and Gemini agree on the #1 pick on only 34.5% of matched prompts — and 33.5% of the time their top-five lists do not overlap at all.

That is not citation noise or domain-ranking drift. It is product-level disagreement on who to recommend when the question is identical.


What we analysed

We used Obsurfable Explorer's research corpus: buyer-style prompts. Each prompt records the latest observation per platform family — ChatGPT / OpenAI API, Gemini / Gemini API, Claude / Claude API, Perplexity, and Grok — preferring UI captures over API runs when both exist.

ScopeValue
Corpus window29 June – 15 September 2026
Active research prompts31,166
Total observations (research subset)35,794
Prompts with 2+ platform families900
Three-way panel (ChatGPT + Gemini + Claude)531
Four-way panel (+ Perplexity)190

Categories in the three-way panel skew toward B2B software and developer media: data infrastructure (202 prompts), developer tools (187), and developer media (140).

Primary comparison metric: top-five brand overlap — how many brands appear in both engines' first-five 'brands mentioned' lists on the same prompt. We also track #1 agreement (same first brand named) and triple-zero (no brand shared across all three top-fives).

Browse live examples on Obsurfable Explorer.


What we found

1. Most matched prompts produce different shortlists

On the 531-prompt three-way panel, brand shortlists diverge sharply:

Shared brands across all three top-5 listsPromptsShare
024446.0%
113224.9%
210820.3%
3397.3%
471.3%
5 (identical shortlists)10.2%

32.0% of prompts (170/531) produce a different #1 brand on each of ChatGPT, Gemini, and Claude.

On the tighter 190-prompt four-way panel (adds Perplexity), 55.8% share zero brands across all four top-fives, and 22.6% have four different #1 picks.

2. Pairwise disagreement is the norm, not the exception

PairMatched prompts (n)Zero top-5 overlapSame #1 pickAvg brands shared
ChatGPT vs Gemini60933.5%34.5%1.47
ChatGPT vs Claude57833.0%31.3%1.38
ChatGPT vs Perplexity31325.9%30.4%1.71
Gemini vs Claude53122.2%35.8%1.69
ChatGPT vs Grok26522.6%28.7%1.64

Even the "closest" major pair — Gemini vs Claude — still shows 22.2% zero overlap. Full five-brand agreement between ChatGPT and Gemini happens on 0.7% of matched prompts (4/609).

3. Category context matters — but disagreement persists everywhere

Three-way panel by category (minimum 15 prompts):

CategorynZero shared top-5 (all three)All different #1
Technology · data infrastructure20253%29%
Technology · developer tools18743%28%
Media · developer media14040%41%

ChatGPT vs Gemini zero-overlap rates on multi-platform prompts:

CategorynZero overlapSame #1
Technology · SEO/AEO tools2065%25%
Technology · data infrastructure22243%31%
Technology · developer tools21627%42%
Media · developer media14126%30%

Developer-tools prompts show relative convergence (27% zero overlap ChatGPT–Gemini), but the three-way panel still yields 43% triple-zero — far too high for a single "AI visibility score."

4. ChatGPT often names no brands where rivals do

On prompts observed on both ChatGPT and Gemini, 13.5% (82/609) return an empty brand list on ChatGPT while Gemini names at least one brand. Across all multi-platform prompts, ChatGPT's empty-brand rate is 14.5% vs 1.1% for Gemini, 1.9% for Claude, and 2.6% for Perplexity.

That gap is structural: the same prompt can yield a tool-stack answer (REST APIs, webhooks) on one engine and a vendor shortlist on another — or brands on one engine and silence on ChatGPT.


Worked examples (same prompt, different winners)

These are live corpus records — not hypotheticals.

Best platforms for devtool exposureexplorer.obsurfable.com/prompts/best-platforms-for-devtool-exposure-02ec7236

EngineTop picks
ChatGPTGitHub, Product Hunt, Hacker News
GeminiHacker News, Daily.dev, Reddit
ClaudeDevHunt, Show HN, Hacker News

Three different #1 answers on one prompt.

How do I set up automated web scraping for competitor price monitoring?explorer.obsurfable.com/prompts/how-do-i-set-up-automated-web-scraping-for-competitor-price-monitoring-2c3e386b

EngineTop picks
ChatGPTPlaywright, Selenium, Puppeteer
GeminiOctoparse, Apify, Browse.ai
ClaudeAmazon, Prisync, Price2Spy

ChatGPT recommends open-source automation stacks; Gemini names no-code scrapers; Claude jumps to retail price-monitoring SaaS — zero shared brands.

Best digest-style dev blogsexplorer.obsurfable.com/prompts/best-digest-style-dev-blogs-1705bdb6

EngineTop picks
ChatGPTMicrosoft Dev Blogs, Google Developers Blog, AWS Architecture Blog
GeminiTLDR, Bytes.dev, The Pragmatic Engineer
ClaudeTLDR, Programming Digest, The Pragmatic Engineer

ChatGPT favours official vendor blogs; Gemini and Claude favour newsletter-style digests — only partial overlap at the long tail.

Category hubs for the densest panels: developer tools, data infrastructure.


Why this is surprising

Citation-leaderboard studies (Ahrefs Brand Radar, Peec, BotRank) already show that domains differ by engine. This finding is sharper: on identical buyer prompts, the named products in the answer shortlist diverge just as much.

Three implications for measurement:

  1. A single-engine brand rank is a partial photograph. ChatGPT's top brands in developer tools (Typeform, Jotform, Google Forms) differ from Gemini/Claude leaders (SurveyJS, Formbricks) on the same underlying prompt set — see the ChatGPT vs Gemini category splits.

  2. "Visible on ChatGPT" ≠ "visible in AI." One engine in seven with an empty brand list — on prompts where competitors name five — will undercount or overcount depending which panel you track.

  3. Cross-engine consensus is rare. Only 0.2% of three-way prompts (1/531) produced identical top-five shortlists. Treat "AI recommendation share" like portfolio exposure: measure per engine, then compare — do not aggregate into one vanity number.


Limitations

  • Panel size: Multi-platform coverage is intentionally sampled — 900 prompts with 2+ engines, not all 31k. Early corpus work focused on technology and developer-media verticals.
  • Latest observation only: We compare the most recent capture per engine, not full time-series drift on the same prompt.
  • Brand extraction: Lists reflect Obsurfable's entity extraction on model outputs. Generic terms (e.g. "REST APIs") may appear instead of vendors on some engines — visible in data-infrastructure prompts.
  • Platform mix: Most three-way panels use UI captures (ChatGPT, Gemini, Claude web apps) plus API runs where noted; Perplexity and Grok appear in smaller matched subsets (313 and 265 prompts respectively vs ChatGPT).
  • Not citation analysis: This report measures brand mentions in answers, not URL citations. A brand can be cited without being named, and vice versa.

Methodology

Corpus filter: active research prompts only.

Platform families: Raw 'platform' values mapped to public families per Explorer taxonomy — e.g. 'openai' and 'chatgpt_ui' → ChatGPT; 'gemini_ui' and 'google' → Gemini; 'claude_ui' and 'anthropic' → Claude.

Deduplication: For each '(prompt, platform family)' pair, keep the latest 'observed_at' observation; prefer UI platform keys over API keys at equal timestamps.

Overlap math: Top-five overlap = |A ∩ B| where A and B are the first five entries in 'brands mentioned' (empty list if none). Triple-zero = |A ∩ B ∩ C| = 0 for ChatGPT, Gemini, Claude top-fives.

Exclusions: Observations with unmapped platform keys omitted from family-level joins. Customer-scoped prompts excluded.

Reproduce or explore individual prompts at explorer.obsurfable.com. Signed-in users can view observation history per prompt.


Bottom line

On matched buyer prompts in Obsurfable's corpus, AI brand recommendations do not converge. Nearly half of ChatGPT–Gemini–Claude panels share no top-five brands at all; one in three gives each engine a different #1 pick. If your visibility program tracks one engine — or one blended score — you are measuring disagreement, not market reality. Start with matched prompts, per-engine shortlists, and the overlap math above.