AI visibility is often reduced to a single number. A brand scores 31%, up from 27%, and that looks like a simple improvement. But that score can hide several very different outcomes.
A brand might be mentioned only to warn people away from it. It might be cited as a useful source while the answer recommends a competitor. It might be recommended without citing anything from its own site. Or one of its pages might appear in the sources an engine considered, but never make it into the final answer.
Those are different events, with different causes and different fixes. Averaging them into one visibility score can make the number move without telling you what actually changed.
Retrieved, cited, mentioned and recommended answer different questions
Retrieved. Did a page from the company’s domain enter the observable source set for the answer?
Cited. Did the finished answer visibly rely on the company’s domain as evidence?
Mentioned. Did the brand appear in the answer at all?
Recommended. Did the answer present the vendor as an option it favoured?
Retrieval and citation tell you which sources an answer used. Mention and recommendation tell you what the answer actually said about the brand. Combining all four into one number hides that distinction.
Retrieval is also different in one important respect: it cannot always be measured, because not every engine shows which sources it considered.
Why it matters
A vendor can be cited without being recommended
A vendor might publish the best explanation of its category. AI answers use that content as a source, but recommend a competitor instead.
A citation-based report would show this as a strong result. But the vendor’s content is being trusted as a source; its product is not being preferred. Those are different outcomes and usually require different work.
A vendor can be recommended without being cited
A well-known vendor might be recommended in the answer, while the citations point to review sites and editorial articles.
If you only track citations to the vendor’s own website, you miss the recommendation entirely, even though it is the outcome most closely tied to a potential sale.
A page can be retrieved without being cited
A page might appear in the sources an engine considered, but not make it into the final answer.
That is different from the page never being considered at all. Neither tells you exactly why it happened, but each points to a different problem to investigate.
AI visibility is not preference
Mention share tells you how often a brand appears in an answer, but not whether the answer is positive.
An answer might say: avoid Vendor A because it charges per page and puts exports behind a more expensive plan. Mention share still counts that as visibility. A recommendation measure shows what actually happened: the brand was rejected in favour of someone else.
Brands can earn lots of mentions because they are well known, frequently compared or often criticised. Treating every mention as a positive result can therefore make a weak position look stronger than it is.
This happens in real measurement
In one B2B SaaS assessment run in June 2026, the difference was visible in the data. Across 765 non-branded answers, one competitor was mentioned 205 times, giving it a 26.8% mention share. It was the lead recommendation in none of them, and in 90 of those 205 answers it was the vendor being argued against.
The same separation appeared between citation and recommendation. Of 160 answers that recommended the vendor being assessed, 40 did so without retrieving or citing a page from its domain. Conversely, 35 of the 140 answers that cited the vendor’s domain did not recommend it.
These are not different ways of measuring the same outcome. A brand can be mentioned without being preferred, cited without being recommended, and recommended without its own content appearing anywhere in the visible source chain.
Separate measurement improves diagnosis
The value of separating retrieval, citation, mention and recommendation becomes clearer when they move differently. None of the measures tells you what caused the change, but the gaps between them tell you where to investigate next.
If mention share rises while recommendation falls, the brand may simply be appearing more often as a negative comparison. A headline visibility score can show improvement while preference is getting worse.
If citations rise while recommendation stays flat, the company’s content may be becoming more useful as evidence without making the product more likely to be chosen. That points towards checking whether the cited material explains the category well but does little to establish the product as an option.
The reverse can happen too. Recommendation may improve while citations to the company’s own domain stay flat. A citation-based dashboard would miss that movement entirely. Third-party sources are then one obvious place to investigate.
Retrieval creates another useful distinction. A page that never enters the observable source set presents a different problem from one that is repeatedly retrieved but never cited.
The measures do not give you the fix. They tell you which question to ask next.
Keep the evidence behind every AI visibility metric
For every answer, store the full response, citations, retrieved pages where the engine exposes them, and your classification results. The raw evidence lets you check or change the measurement later.
If you save citations but not retrieval data, you cannot tell whether a page was never retrieved or was retrieved and then left out of the final citations.
The same applies to mentions and recommendations. If you save only a score, rather than the answer itself, you cannot later tell whether the brand was recommended, mentioned neutrally or criticised.
The underlying answers should therefore be treated as the source data. Your metrics can then be recalculated as definitions improve, without having to run the prompts again.
Measure AI visibility metrics separately
Retrieval, citation, mention and recommendation measure different things, and combining them into one score removes the information needed to understand what changed. Measure them separately wherever the data allows, and show the denominator for each.
Presence is not preference, and citation is not recommendation. A useful AI visibility report should show which of these outcomes changed, so you can investigate the right problem rather than reacting to a single blended score.