Short answer: When AI engines attach cited sources to buyer-style prompts, they overwhelmingly link to vendor-owned domains — not third-party review aggregators. In Obsurfable's research corpus, review sites (G2, Capterra, Trustpilot, Gartner) account for 0.3% of all citation URLs across 18,600 linked sources. Vendor and product domains capture 85.5%. Even in B2B software categories where review platforms dominate traditional SEO, review sites appear in under 1% of AI citations.
The citation landscape is also fragmented: 4,785 unique hosts appear across the corpus, and the single most-cited domain — plainenglish.io — holds just 1.9% of all links.
What we analysed
We classified every cited URL in Obsurfable Explorer's research corpus: buyer-style prompts tracked across ChatGPT, Gemini, Claude, Perplexity, and Grok. For each prompt we keep the latest observation per platform family, preferring web-app captures over API runs when both exist.
Each cited URL was assigned a host taxonomy:
| Host type | Examples |
|---|---|
| Vendor / product | Brand homepages, product docs on vendor domains, comparison blogs on vendor sites |
| Publisher | Medium, Substack, dev.to, Plain English, Daily.dev |
| Documentation | docs., learn., developers.google.com |
| Developer | github.com, npmjs.com, pypi.org |
| Review | g2.com, capterra.com, trustpilot.com, gartner.com |
| Community | reddit.com, stackoverflow.com, Hacker News |
| Government / encyclopedia / video | .gov domains, wikipedia.org, youtube.com |
| Scope | Value |
|---|---|
| Corpus window | 11 March – 7 October 2026 |
| Active research prompts | 37,296 |
| Total observations | 44,792 |
| Deduplicated observations (latest per prompt × platform) | 39,724 |
| Observations with at least one cited source | 1,867 (4.7%) |
| Total citation URLs | 18,600 |
| Unique cited hosts | 4,785 |
Primary metrics: citation host share (% of all linked URLs by domain and taxonomy class) and host concentration (top-domain share and HHI within categories).
Browse live category data on Obsurfable Explorer.
What we found
1. Review sites barely appear in AI citations
Across the full corpus, citation URLs break down by host type:
| Host taxonomy | Citations | Share |
|---|---|---|
| Vendor / product | 15,895 | 85.5% |
| Publisher | 1,338 | 7.2% |
| Documentation | 765 | 4.1% |
| Developer (GitHub, npm) | 375 | 2.0% |
| Community (Reddit, Stack Overflow) | 78 | 0.4% |
| Review (G2, Capterra, Gartner, etc.) | 56 | 0.3% |
| Government | 43 | 0.2% |
| Video | 43 | 0.2% |
| Encyclopedia | 7 | 0.0% |
56 review-site citations out of 18,600 total. For context, brands are named on roughly 76% of all observations in the corpus — but when engines attach links, they point to vendor properties, not review aggregators.
2. B2B software categories show the same pattern
In the four highest-volume B2B technology categories, review sites remain negligible:
| Category | Citations | Review sites | Publisher | Vendor / product |
|---|---|---|---|---|
| Technology · API platforms | 5,158 | 0.3% | 5.4% | 86.7% |
| Technology · data infrastructure | 3,928 | 0.2% | 3.4% | 88.5% |
| Technology · developer tools | 2,714 | 0.7% | 3.6% | 84.3% |
| Technology · SEO/AEO tools | 2,499 | 0.1% | 0.8% | 95.8% |
In API platforms — a category where G2 and Capterra pages rank prominently in traditional search — the top cited hosts are vendor domains: notify.cx (5.2%), resend.com (5.1%), postmarkapp.com (3.7%). No review aggregator appears in the top 20.
3. Citation hosts are more fragmented than brand mentions
Brand mentions in the same categories are already fragmented (see our fragmentation analysis). Citation hosts are even more dispersed:
| Category | #1 brand share | #1 host share | Unique brands | Unique hosts |
|---|---|---|---|---|
| API platforms | Twilio 6.9% | notify.cx 5.2% | 1,745 | 994 |
| Data infrastructure | Bright Data 4.9% | brightdata.com 5.1% | 2,383 | 998 |
| Developer tools | SurveyJS 2.2% | github.com 5.9% | 2,193 | 1,020 |
| SEO/AEO tools | ChatGPT 6.3% | semrush.com 2.2% | 984 | 1,123 |
In SEO/AEO tools, the top cited host (semrush.com) holds just 2.2% of citations while 1,123 unique hosts compete for links — more hosts than named brands.
Corpus-wide, the top 20 cited domains account for only 16.4% of all links combined:
| Rank | Host | Citations | Share |
|---|---|---|---|
| 1 | plainenglish.io | 349 | 1.9% |
| 2 | dev.to | 343 | 1.8% |
| 3 | github.com | 306 | 1.6% |
| 4 | resend.com | 280 | 1.5% |
| 5 | notify.cx | 275 | 1.5% |
| 6 | resources.plainenglish.io | 250 | 1.3% |
| 7 | brightdata.com | 202 | 1.1% |
| 8 | postmarkapp.com | 202 | 1.1% |
| 9 | sequenzy.com | 165 | 0.9% |
| 10 | developers.google.com | 124 | 0.7% |
No single domain commands even 2% of AI citations.
4. Publisher citations concentrate in media categories
The one category where non-vendor hosts matter is developer media:
| Host taxonomy | Share of citations |
|---|---|
| Vendor / product | 65.5% |
| Publisher | 29.9% |
| Documentation | 2.2% |
| Community | 1.0% |
| Review | 0.5% |
resources.plainenglish.io alone accounts for 10.0% of citations in this category — the highest single-host concentration in any category with 100+ citations. Publisher-heavy categories are the exception; B2B software citations remain vendor-dominated.
5. Citation attachment varies sharply by platform
Most corpus observations come from ChatGPT API runs, which rarely attach sources. Among platform families that do cite:
| Platform | Observations | % with sources | Avg citations/answer | Unique hosts |
|---|---|---|---|---|
| ChatGPT | 37,286 | 0.6% | 0.09 | 846 |
| Gemini | 760 | 53.0% | 2.64 | 814 |
| Claude | 721 | 57.3% | 2.78 | 830 |
| Perplexity | 511 | 79.5% | 7.60 | 1,448 |
| Grok | 446 | 98.2% | 16.45 | 2,548 |
Grok attaches sources on 98.2% of answers and links to 2,548 distinct hosts — the broadest citation surface in the corpus. Perplexity cites an average of 7.6 URLs per answer. ChatGPT's citation rate in the corpus is near zero because most observations are API captures without source attachment.
When search-native engines cite, they spread links across hundreds of vendor domains — not a handful of review aggregators.
Why this is surprising
Three assumptions in AI visibility strategy do not match the citation data:
-
"Get listed on G2/Capterra to win AI citations." Review aggregators appear in 0.3% of linked URLs. Vendor product pages, docs, and vendor-published content dominate.
-
"A few authority domains control AI citations." The top host holds 1.9% corpus-wide. Even category leaders like notify.cx in API platforms reach only 5.2% — and that is a vendor domain, not a third-party reviewer.
-
"Citations and brand mentions follow the same concentration pattern." Brand mentions are fragmented (median category leader at 2.9%), but citation hosts are more fragmented — with higher unique-host counts than unique-brand counts in several categories.
The gap between naming a brand (76% of observations) and linking to a source (4.7%) also matters. Engines frequently recommend products by name without attaching any URL. Citation strategy and mention strategy are separate games.
Limitations
- Low overall citation rate. Only 4.7% of deduplicated observations include cited sources, driven by ChatGPT API captures that do not attach links. Findings describe the citation subset — not all AI answers.
- Host taxonomy is heuristic. Domains like aimultiple.com or itechguides.com are classified as vendor/other; some function as comparison publishers. Manual review of borderline hosts would shift taxonomy shares slightly but would not change the review-site finding (56 citations total).
- Corpus skew. B2B technology and developer-media categories dominate prompt volume. Consumer categories with stronger review-site presence in traditional search are underrepresented.
- Point-in-time captures. Host rankings reflect the corpus window (March–October 2026) and may shift as engines update retrieval behavior.
Methodology
Corpus. Active research prompts in Obsurfable Explorer — buyer-style vendor-selection questions across 325+ categories. Customer-scoped prompts excluded.
Deduplication. Latest observation per prompt per platform family (ChatGPT, Gemini, Claude, Perplexity, Grok). When both UI and API captures exist for the same family, UI preferred.
Citation extraction. Each URL in the sources array attached to an observation. Hostname normalized (www. stripped). Multiple URLs from the same host in one answer each counted.
Host taxonomy. Rule-based classification on hostname patterns (see table above). Vendor/product is the residual class for brand domains, product blogs, and comparison sites without a review-aggregator or publisher pattern.
Concentration metrics. Top-host share = citations to #1 domain ÷ total citations in scope. HHI = sum of squared host shares (lower = more fragmented).
Platforms. ChatGPT (UI + API), Gemini (UI), Claude (UI), Perplexity, Grok (UI). Copilot, DeepSeek, Mistral, and Meta AI are not yet represented at sufficient volume for category-level citation analysis.
Data and live examples: explorer.obsurfable.com.
Conclusion
AI engines that attach sources link to vendor-owned domains, not review aggregators. Across 18,600 citations in Obsurfable's corpus, G2, Capterra, Trustpilot, and Gartner combined account for 0.3% of linked URLs. Vendor and product sites capture 85.5%. The citation landscape is a long tail of 4,785 hosts with no domain above 2% corpus-wide.
If your AI visibility program optimizes for review-site presence, the citation data suggests a different target: your own product pages, documentation, and publishable content — because that is what engines actually link to when they cite at all.