A category is forming around measuring AI visibility, and it is filling with numbers that look like metrics. Share of voice in AI answers. Citation scores. Visibility indices. They are presented with the confidence of Search Console data and they are not that.

It is worth being precise about what can and cannot currently be known, because the gap is where a lot of budget is about to go.

What Google gives you

The generative AI performance report in Search Console shows impressions in AI Overviews and AI Mode, by page, country, device and date. No clicks. No click-through rate. No average position. No query dimension. It is first-party and trustworthy, and it tells you that a page was surfaced and nothing about what happened next.

It is also still rolling out to a subset of properties, and it can be empty for reasons that have nothing to do with performance.

What your analytics gives you

GA4 has an AI Assistant channel covering third-party assistants, and it explicitly excludes AI Overviews and AI Mode, which are counted as Organic Search. So the analytics view and the Search Console view do not describe the same thing, and cannot be added together.

On top of that, sessions arriving without a referrer land in Direct, which understates every AI bucket by an unknown amount.

What the tools give you

Third-party AI visibility tools work by running prompts against models on a schedule and recording which brands and URLs appear. That is a legitimate method and it produces real information.

What it does not produce is a population statistic. The prompt set is chosen, not sampled from actual user behaviour, which nobody outside the platforms can observe. Model responses vary between runs, between accounts, between regions, and between one week and the next. A number built on that describes the prompt set, on those days, in that configuration.

Google states the boundary plainly in its own guidance: be wary of third-party tools that promise ranking success or claim to use internal Google metrics, because no third-party tool has access to its internal ranking or AI systems. That is worth quoting because it is the platform ruling out a claim its own ecosystem keeps making.

What we do instead

Say what each number is. An impression count is a surfacing signal. A prompt test is a sample. A referral session is one visit that kept its referrer. None of those is visibility, and combining them into one figure produces a number with no definition.

Run prompt tests, but run them as experiments with a written question, a fixed method, a stated date and published limitations, so a result can be repeated and disagreed with. A finding that cannot be replicated is an anecdote with a chart.

Then anchor the whole thing on the outcome that has not changed. Qualified leads and revenue are still measurable, still first-party, and still the only figures that decide whether any of this worked. If AI visibility work is real, it shows up there eventually. If it only ever shows up in a vendor's index, it did not happen.

This is a temporary position rather than a permanent one. Google has said it expects to add metrics to the generative AI report over time. When clicks or queries arrive, the honest answer changes, and we will say so.

What this means for an operator

Do not set a target on an AI visibility score. Track the three partial sources separately, label each for what it is, and hold the programme accountable to qualified leads.