Mixed evidence
Why AI Citations Vary From One Test to the Next
AI answers are variable across products, modes, dates, prompts, sessions, and locations, so one screenshot cannot establish a stable citation ranking.
What matters
- Archive exact prompts, dates, modes, answers, and cited URLs.
- A citation win in one session is evidence of possibility, not stable rank.
- Separate citation rate from factual correctness and recommendation language.
The answer is an observation, not a position
Classic rank tracking already varies by location, device, time, and result type. Generative products add retrieval and synthesis choices that can produce different source sets or language across repeated trials.
That makes “we rank number one in ChatGPT” an ill-defined statement unless the speaker provides the product, mode, prompt set, dates, geography, sample, repeat protocol, and scoring rule.
Record the distribution
For a fixed prompt, run repeated trials according to a predeclared schedule. Record whether the source was mentioned, cited, represented correctly, or recommended. Keep all cited URLs. Summarize the fraction of eligible trials, not the best-looking output.
When a result changes, check accessible explanations first: page availability, robots behavior, indexing, factual consistency, competitor sources, and product updates. Do not infer an invisible “penalty” from a small sample.
The full protocol is in How to Measure AI Visibility Without Fooling Yourself.
Evidence & maintenance
How this page is maintained
- Content basis
- Mixed evidence
- Evidence grade
- Peer-reviewed research
- Next review
- Dec 10, 2026
Material errors can be reported through the public corrections process.
Sources
- Towards a Measurement-Based Audit of Generative AI Citation Behavior — Proceedings of Machine Learning Research