Original analysis

How to Measure AI Visibility Without Fooling Yourself

A repeatable way to measure mentions, citations, correct facts, visits, changing results, and business value across AI tools.

What matters

  • Set prompts before the test so you cannot pick only easy wins.
  • A mention, citation, tip, click, and customer are different events.
  • Repeat tests because answers change by time, session, mode, and place.
  • Save bad and wrong answers, not just images of wins.

Define the unit of success

“AI visibility” can mean several results:

Outcome Definition
Fact mention The answer names the person, company, product, or fact
Citation The answer exposes a source link or attributable reference
Owned citation The citation points to the measured organization’s domain
Outside citation A source from another site supports the name or claim
Recommendation The answer includes the person or group in its advice
Referral A person arrives from the answer product
Lead The visit leads to a useful action the person chose

Report each measure on its own. One secret score can hide what happened. The AI may cite a good source, name a brand, or send a customer. Those are not the same.

Build the prompt set before testing

Gather real questions from sales calls, support, search data, groups, and product research. Sort them by buying stage, place, reader, and goal. Include broad prompts, brand prompts, comparisons, fact questions, and tasks where your company should not win.

Use a written rule to pick the set. Keep the words and target place fixed for each test period. Save new questions for a later set. Do not add only prompts where the brand does well.

Create the run record

For each trial, store:

  • Prompt ID and exact text.
  • Product and mode.
  • Visible model or version, if shown.
  • Date, time, account state, and geography.
  • Exact answer or an allowed saved copy.
  • Every cited URL and its cited claim.
  • Name mention, place in the answer, tone, and strength of the advice.
  • Fact check against your main fact list.
  • Human reviewer and notes on close calls.

Follow the tool’s terms. Do not act as if an AI gives the same answer every time.

The public AI visibility run log starts with sample rows for a discovery prompt, a fact prompt, and a problem-solving prompt.

Repeat tests and track changes

Run prompts more than once during the test period. A result in one session may vanish in the next. Useful totals include:

  • Mention rate: valid runs that name the company ÷ all valid runs.
  • Owned citation rate: trials citing the owned domain ÷ eligible trials.
  • Citation share: owned citations ÷ all citations in the panel.
  • Correctness rate: checked claims judged materially correct ÷ checked claims.
  • Change rate: how results differ across runs of the same prompt.
  • Visit to lead rate: good actions ÷ visits we can link to the AI tool.

Always show the denominator.

Example repeated-run report

Prompt Trials Mentions Owned citations Material facts correct
P001 4 3 of 4 2 of 4 2 of 3 checked
P002 4 4 of 4 1 of 4 4 of 4 checked
P003 4 0 of 4 0 of 4 Not applicable

This is a made-up example that shows the report format. It is not a result for Bob or a client. Keep the zero row and the wrong fact. They stop a summary from hiding weak performance.

Connect answer measurement to search and delivery

A drop may come from blocked pages, lost indexing, lower search demand, tool changes, rival proof, or random noise. Compare the prompt test with bot logs, cache results, Search Console, Bing Webmaster Tools, source links, and leads.

This does not show how the model thinks. It helps you find site and content problems you can test before you invent a rank theory.

Keep a list of key facts

Keep one main record for facts such as Bob’s role, service terms, study dates, measure meanings, and public profiles. For each fact, save the best wording, source link, owner, last check date, and claim limits.

When an answer is wrong, first check whether your own pages disagree. Fix your source pages before you blame the model. Then publish clear proof and watch for a change.

Report results with restraint

A good report lists the prompt sample, tools, dates, repeats, missing data, review rules, results, errors, citations, and limits. It does not say one code change caused a citation unless the test can prove it.

The goal is more correct answers with clear sources that reach the right people. The goal is not a folder of good screenshots.

Sources

  1. Towards a Measurement-Based Audit of Generative AI Citation Behavior, Proceedings of Machine Learning Research
  2. Introducing AI Performance in Bing Webmaster Tools Public Preview, Microsoft Bing
  3. Publishers and developers FAQ, OpenAI