Mixed evidence

Why AI Citations Vary From One Test to the Next

AI answers change across tools, modes, dates, prompts, sessions, and places. One image cannot prove a stable citation rank.

What matters

  • Archive exact prompts, dates, modes, answers, and cited URLs.
  • One citation shows what can happen, not a stable rank.
  • Track citations, correct facts, and advice on their own.

The answer is an observation, not a position

Search ranks already change by place, device, time, and result type. AI tools also choose and join sources in ways that can change on each run.

So “we rank first in ChatGPT” means little without details. Name the product, mode, prompts, dates, place, sample, repeat plan, and score rule.

Record the distribution

Run each fixed prompt many times on a set plan. Record if the source was named, cited, shown correctly, or advised. Save all source links. Report the share of all valid runs, not the best answer.

For example, one prompt can produce this small record:

Trial Brand named Owned source cited Key facts correct
1 Yes Yes Yes
2 Yes No Yes
3 No No Not applicable
4 Yes Yes No

The owned citation rate is 2 of 4, not “we rank in this tool.” The wrong fact in trial 4 also matters. This example explains the math and is not a live result for Bob or a client.

When results change, check simple causes first. Look at page access, robots rules, indexing, fact conflicts, rival sources, and tool updates. Do not claim a hidden penalty from a small sample.

The full protocol is in How to Measure AI Visibility Without Fooling Yourself.

Download the AI visibility run log to save the prompt, run conditions, exact result, citations, and fact check.

Sources

  1. Towards a Measurement-Based Audit of Generative AI Citation Behavior, Proceedings of Machine Learning Research