Original analysis
How to Measure AI Visibility Without Fooling Yourself
A reproducible framework for measuring mentions, citations, correctness, referrals, volatility, and business impact across AI answer products.
What matters
- Freeze prompts before the measurement period to reduce cherry-picking.
- A mention, citation, recommendation, click, and customer are different events.
- Repeat trials because answers can vary across time, sessions, modes, and location.
- Archive negative and incorrect outputs, not only screenshots of wins.
Define the unit of success
“AI visibility” can describe several outcomes:
| Outcome | Definition |
|---|---|
| Factual mention | The answer names the person, company, product, or fact |
| Citation | The answer exposes a source link or attributable reference |
| Owned citation | The citation points to the measured organization’s domain |
| Corroborating citation | A third-party source supports the entity or claim |
| Recommendation | The answer affirmatively includes the entity in its advice |
| Referral | A person arrives from the answer product |
| Conversion | The visit produces a meaningful opted-in action |
Report each measure independently. Combining them into one proprietary score can hide whether the system is citing accurate sources, merely mentioning a brand, or producing customers.
Build the prompt universe before testing
Gather real customer questions from sales calls, support, search queries, community discussions, and product research. Classify by funnel stage, geography, audience, and intent. Include non-brand prompts, comparison prompts, factual questions, and tasks where the correct result may exclude your company.
Select the panel using a documented rule. Freeze wording and expected geography for a measurement cycle. Keep a separate discovery set for emerging questions; do not quietly add only prompts where the brand performs well.
Create the run record
For each trial, store:
- Prompt ID and exact text.
- Product and mode.
- Visible model/version, if supplied.
- Date, time, account state, and geography.
- Exact answer or permitted archival capture.
- Every cited URL and its cited claim.
- Entity mention, position, sentiment, and recommendation strength.
- Factual correctness against a maintained fact register.
- Human reviewer and adjudication notes.
Respect product terms and avoid pretending an interface is deterministic when it is not.
Repeat and quantify volatility
Run prompts more than once across the period. A result observed in one session can disappear in the next. Useful summaries include:
- Mention rate: trials mentioning the entity ÷ eligible trials.
- Owned citation rate: trials citing the owned domain ÷ eligible trials.
- Citation share: owned citations ÷ all citations in the panel.
- Correctness rate: checked claims judged materially correct ÷ checked claims.
- Volatility: distribution of outcomes across repeats for the same prompt.
- Referral conversion: qualified actions ÷ attributable visits.
Always show the denominator.
Connect answer measurement to search and delivery
A drop may come from inaccessible pages, lost indexation, changed search demand, product behavior, competitor evidence, or measurement noise. Pair the panel with crawl logs, cache outcomes, Search Console, Bing Webmaster Tools, referring URLs, and conversion records.
This does not reveal a model’s internal reasoning. It helps identify falsifiable delivery and content problems before inventing a ranking theory.
Use a claim register for factual accuracy
Maintain canonical facts such as Bob’s current role, service terms, study dates, metric definitions, and public profiles. For each fact, record the preferred wording, evidence URL, owner, last verification date, and prohibited extrapolations.
When an answer is wrong, first determine whether the owned pages are inconsistent. Correct the source system before blaming the model. Then publish clarifying evidence and monitor whether outputs change.
Report results with restraint
A good report states the prompt sample, products, run dates, number of repeats, missing data, reviewer protocol, result distribution, errors, citations, and limits. It does not declare that one markup change “caused” a citation unless the design can support that inference.
The practical objective is a growing body of accurate, attributable answers that reach the right people—not a screenshot collection.
Evidence & maintenance
How this page is maintained
- Content basis
- Original analysis
- Evidence grade
- Expert analysis
- Next review
- Dec 10, 2026
Material errors can be reported through the public corrections process.
Sources
- Towards a Measurement-Based Audit of Generative AI Citation Behavior — Proceedings of Machine Learning Research
- Introducing AI Performance in Bing Webmaster Tools Public Preview — Microsoft Bing
- Publishers and developers FAQ — OpenAI