Ranking, retrieval, mention, citation, recommendation and conversion are six different outcomes, and one does not guarantee the next. A page can rank without being retrieved by an AI assistant, be retrieved without being cited, and be cited without anyone clicking. Measuring them as one “visibility” number hides where the chain actually breaks.
What does each outcome mean?
Each one describes a different moment between a question being asked and a customer acting on the answer.
| Outcome | What happened |
|---|---|
| Ranking | Your page appeared at a position in a list of results |
| Retrieval | A system pulled your content into what it used to build an answer |
| Mention | The answer named your brand |
| Citation | The answer linked your page as a source |
| Recommendation | The answer suggested you as the choice |
| Conversion | Someone arrived and did what you wanted them to do |
Which outcomes can actually be measured?
Rankings and conversions can be measured directly; mentions and citations can be observed by testing; retrieval mostly has to be inferred.
| Outcome | How | Certainty |
|---|---|---|
| Ranking | Search Console, rank trackers | Measured |
| Retrieval | Rarely exposed by platforms | Inferred |
| Mention / citation | Repeated, logged prompt tests | Observed |
| Recommendation | Decision-stage prompt tests | Observed |
| Referral | Analytics referrer data, partly | Partly measured |
| Conversion | GA4 or CRM events | Measured |
How do you test AI answers properly?
Treat every prompt test as an experiment: fix the conditions, repeat it, log everything and compare like with like.
- Write prompt families, not single prompts: informational, comparison and decision-stage questions
- Record the conditions: platform, model if shown, date, region, signed in or not, search mode
- Separate cold and contextual tests: a fresh chat behaves differently from a long conversation
- Repeat each prompt several times, because answers vary run to run
- Log mentions, citations and recommendations as separate columns
- Re-run monthly and date every result, since platforms change without notice
What are the common measurement mistakes?
Treating one screenshot as a trend, and treating a correlation as a cause.
- Reporting a single AI answer as proof of visibility
- Adding mentions and citations together into one score
- Crediting a change for a rise that also happened to competitors
- Comparing results taken on different platforms or dates
- Hiding tests that came back negative or inconclusive
The GEO Visibility Index in the Lab applies this routine to citation tracking.