Obsurfable

How to Measure AI Visibility Without Fooling Yourself

Obsurfable

A single check of whether ChatGPT mentions your brand tells you close to nothing.

That is not a hot take. It is the opening line of LLM Pulse's methodology guide, published 27 September 2026 — and it reflects a growing consensus in the measurement literature. An AI assistant's answer is one sample from a distribution, conditioned on the prompt, whatever was retrieved, model version, and sampling configuration. Ask the same question twice and you can get different brand lists, different orderings, and different sources.

The GEO industry has a measurement problem. Vendors publish visibility percentages without the design behind them — and the design decides whether the number means anything. This article unpacks the sampling discipline LLM Pulse advocates, connects it to recent research on volatility and cross-engine divergence, and gives you a framework for building defensible AI visibility metrics.

Why single snapshots fail

Three independent lines of evidence converge on the same conclusion:

1. Citation sets rotate quickly

Trakkr Research found 73.5% of citation URLs appear exactly once across 10 months of daily tracking on 8 models. Writesonic tracked 23 million cited sources and reported median citation half-lives of 4–5 weeks. Scrunch/Stacker measured 51.8% typical week-over-week citation count swings across 3.5 million events.

2. One run misses most of the brand universe

A September 2026 study covering 50 questions, six engines, and 15 repeated runs found that one run captured only 62%–77% of the brands observed across five runs. Single checks systematically undercount.

3. Cross-engine divergence is extreme

Tannenbaum's September 2026 audit of four live engines on 15 matched prompts found mean cross-engine URL overlap of 0.0079 Jaccard — 96.4% of URLs appeared in only one engine.

Volatility is not noise to smooth away. It is a property of the system you are measuring. Rank-style tracking — one snapshot, compare last week — inherits assumptions from SEO that do not hold in generative search.

The LLM Pulse sampling design

LLM Pulse's methodology rests on five principles:

1. Fixed prompt sets, not ad hoc queries

Build a library of 20–50 prompts derived from actual buyer language — how customers ask about your category, not how marketers keyword-stuff. Run the same set on a schedule. Changing the prompt panel changes the denominator; you cannot compare periods with different panels.

Prompt classification matters. LLM Pulse separates brand-focused prompts from category prompts so competitor comparisons are not trivially biased by prompts that name your brand.

2. Repeated observations per engine and locale

Run each prompt multiple times per engine. Separately per model and per locale. Keep every answer: text, brands mentioned, mention position, sentiment, and every source cited.

The history stays in your account so the baseline you compare against is your own — not a vendor's opaque benchmark.

3. Explicit denominators

"38% visibility" means nothing until it says "of 412 answers." A rate over 12 answers and a rate over 1,200 do not belong in the same chart. Always report:

  • Number of prompts
  • Number of runs per prompt per engine
  • Total answer count
  • Time window

4. Separate signals per answer

Track three distinct outcomes per response:

SignalDefinitionWhy it matters
MentionBrand named in answer textAwareness and recommendation
CitationEngine links your domain as a sourceTraffic and verifiability
RecommendationBrand explicitly suggested as a choiceDecision-stage influence

GEOly's June 2026 US data illustrates the gaps: 14% of brand mentions in ChatGPT shopping answers had no buyable card attached. Brands won the argument and lost the checkout. Monitoring mentions alone misses citation and commerce signals.

5. Engine-specific read schedules

EngineSuggested cadenceRationale
ChatGPT / PerplexityWeeklyBrowsing mode pulls live web data
GeminiBi-weekly to monthlyTracks closer to Google's update rhythm
ClaudeMonthly to quarterlyTraining-data-driven answers shift slowest

Blending engines into one score hides platform-specific movement. Each engine reads a different search index.

Tremor: measuring background volatility

LLM Pulse publishes Tremor, a public AI search volatility tracker updated every Tuesday. It quantifies how much AI answers change on their own week to week — so nobody has to take a vendor's word for how unstable the background is.

Tremor compares prompts that ran in both the current and previous week, per AI model:

SignalWhat it measuresWeight in headline
Citation churnJaccard distance on cited domains week-over-weekHighest
Mention churnBrands appearing, disappearing, or moving positionMedium
Sentiment variationTone shifts toward tracked brandsLower

Each signal is scored against the model's own trailing eight-week baseline (robust z-score), then blended into a seismic magnitude from 0 (unusually calm) to 10 (major upheaval). A normal week reads around 2.

Use Tremor to separate model updates from real visibility changes. When Tremor spikes and your mention rate drops, the engine moved — not necessarily your content.

Six questions to ask before you believe a visibility study

LLM Pulse's guide gives six questions for evaluating vendor-published research. Adapted and extended:

1. How many observations?

Two hundred prompts run once is 200 observations. The same 200 run weekly for three months is 2,600+. The difference is not cosmetic — it determines whether a percentage is a rate or a coin flip.

2. Are prompts fixed or drifting?

If the vendor changes prompts between periods, movement may reflect panel composition, not engine behavior.

3. Is the denominator explicit?

Demand mention rate, citation rate, and recommendation rate as separate metrics with answer counts.

4. Does the study query live engines?

Engine-agnostic page scores estimate fit — not exposure. Studies based only on page inspection cannot measure citation probability.

5. How many engines?

Single-engine studies explain at most a fraction of the multi-engine citation universe. Only 2.7% of domains are cited by all five engines.

6. Does the methodology account for the stage model?

Recent research decomposes visibility into match, exposure, selection, and prior (stage model paper). A study that reports only end-to-end mention rate cannot tell you which lever to pull.

Building your measurement stack

Tier 1: Minimum viable monitoring (week 1)

  1. Select 20–30 buyer-intent prompts from sales calls, support tickets, and search query data.
  2. Run each prompt 3–5 times on ChatGPT and one secondary engine (Perplexity or Gemini).
  3. Record mentions, citations, and competitor presence in a spreadsheet.
  4. Calculate mention rate with explicit denominator: mentions / total answers.

This is exploratory. Do not make budget decisions on 60–150 observations.

Tier 2: Operational monitoring (month 1–3)

  1. Expand to 40–50 prompts across intent stages (awareness, comparison, decision).
  2. Run weekly per engine on a fixed schedule.
  3. Track three rates separately: mention rate, citation rate, recommendation rate.
  4. Add 2–3 named competitors for share-of-voice.
  5. Store full answer text for audit and dispute resolution.

Tier 3: Program-grade measurement (quarter 1+)

  1. Integrate with GA4 or server logs to compare AI referral traffic against visibility metrics.
  2. Decompose by stage: exposure rate (domain in sources), selection rate (mention given citation).
  3. Monitor Tremor or equivalent volatility baseline before attributing movement.
  4. Cross-reference with Search Console generative AI performance report (impressions only — no clicks or position for AI Overviews and AI Mode).
  5. Run paraphrase variants of top prompts to test phrasing sensitivity.

Erlin's client data shows AI traffic converting at 3–6× the rate of organic — a gap that only appears when you track both channels together.

Common measurement mistakes

Mistake 1: Treating visibility like rank

SEO rank is position-based and relatively stable day to day. AI visibility is binary per answer and probabilistic across runs. A brand can be "rank 1" in 40% of answers and absent in 60% — there is no single position to track.

Mistake 2: Blending mentions and citations

A mention without a citation means the model recalled your brand from memory or third-party evidence. A citation means your content was used as a verifiable source. They require different remediation. See ChatGPT's Citation Gap.

ChatGPT generates roughly 91% of AI referral sessions per Erlin's analysis — but citation behavior on Perplexity, Gemini, and Google AI Overviews diverges sharply. Optimizing for one engine may not transfer.

Mistake 4: Ignoring Google's measurement limits

Search Console's generative AI performance report gives impressions only — no clicks, no CTR, no position — for AI Overviews and AI Mode. You cannot reconcile payouts (like Google's AI Contribution pilot) against usage from this report alone.

Mistake 5: Vendor score comparison

The industry has not adopted a universal visibility formula. Scores from different methodologies should not be compared. Some report simple mention rate; others add position weighting, sentiment, consistency, or source authority.

HubSpot, LLM Pulse, and the tooling landscape

September 2026 brought renewed attention to measurement tooling:

  • HubSpot AEO (September 2026 roundup) tracks citations and mentions across ChatGPT, Perplexity, and Gemini with daily prompt runs. Standalone at $50/month for 25 prompts.
  • LLM Pulse covers ChatGPT, Perplexity, Gemini, Google AI Overviews, and AI Mode on weekly schedules, with Tremor for public volatility tracking.
  • Erlin, GEOly, Frase, Semrush AI Toolkit offer overlapping capabilities with different engine coverage, cadence, and pricing models.

Tool choice matters less than methodology discipline. A expensive platform with a drifting prompt panel produces less signal than a spreadsheet with fixed prompts and explicit denominators.

How Obsurfable fits

Obsurfable is built on the observation model this article advocates. Each response becomes an observation in a shared public corpus: prompts, answers, brands mentioned, and citations — free for anyone to browse.

That structure supports the sampling design directly:

  • Fixed prompts, repeated runs — not one-off checks.
  • Separate mention and citation signals — not a blended vanity score.
  • Public corpus — anyone can audit the underlying answers, not just trust a dashboard number.
  • Cross-engine comparison — browse how ChatGPT and Gemini answered the same prompt differently.

The Visibility Director uses your observation history to flag prompt-level gaps and draft content aimed at the retrieval layer that is failing — exposure, selection, or match.

For the white-paper depth on KPI frameworks, see Measuring AI Visibility: A Framework for 2026.

FAQ

How many runs per prompt do I need?

Start with 3–5 for exploration. For program-grade rates, 10+ runs per prompt per engine reduce confidence interval width materially. The September 2026 repeated-query study found one run misses 23%–38% of observed brands.

Should I weight mentions by position?

Optional. First-mentioned brands carry more decision influence than fifth in a list. LLM Pulse's AI Visibility Score weights by position. Raw mention rate is simpler and still useful — just be consistent.

Can I use Search Console instead of prompt monitoring?

Search Console measures impressions on Google AI surfaces — not brand mentions, not competitor share, not ChatGPT or Perplexity. It complements prompt monitoring; it does not replace it.

What if my visibility score drops but Tremor is calm?

Investigate content, competitive, and distribution changes first. A calm Tremor background makes movement more likely to reflect real shifts in your footprint or competitor activity.

Bottom line

AI visibility is a sampled distribution, not a rank you look up. Fixed prompt sets, repeated observations, explicit denominators, separate mention/citation/recommendation signals, and engine-specific cadences turn coin flips into defensible metrics.

Before you believe any visibility number — including ours — ask how many observations, whether prompts were fixed, and which stage of the pipeline was measured. The design decides whether the number means anything.

Measure the background volatility. Report the denominator. Sequence optimization to the stage that is failing.