Does AI recommend your brand — and would you know if it did?
The conversation about AI visibility has split into two tribes. One counts citations — how often ChatGPT or Google's AI Overview links to your site — and optimises the number. The other waves the counting away, says it is all about "narrative" and "cultural relevance", and declares the whole thing unmeasurable by design. Both are half-right, and both are missing the same thing.
Being cited is not being recommended. Ask an AI assistant a question in your category and one of three things is true: you are absent, you are one of the sources the answer was quietly built from, or you are the brand the answer actually names. Most AI-visibility tools report a single number — a citation count, a share — that cannot tell those three apart. The counters measure the first tribe's rung and call it visibility; the narrative crowd is right that it is thin, but wrong that the answer is to stop measuring.
The honest answer to "does AI recommend my brand" is not a percentage. It is a position on a ladder — and measuring the ladder is exactly how you reconcile the two tribes.
Why "AI citation rate" is a vanity metric
I spent more than a decade running eight-figure marketing budgets, and the one law that never broke is Goodhart's: the moment a metric becomes the target, someone games it. A raw AI-citation count is the newest target, and it has three problems on top of that — each fatal on its own.
It is gameable. A citation means a model consulted your page. That is pushed by the same tactics that have always chased links and rankings. We measure it, and we tell clients out loud that it is the soft rung — precisely so the numbers above it carry the weight.
It moves when nothing changes. Ask the same questions twice on the same morning, everything identical, and in our capture sets 15 to 20% of them flip on whether an AI answer even appears. Google decides at serve time. A tracker that reports "AI visibility up 12%" this week is, more often than not, reporting the weather.
It usually undercounts itself without telling you. Google now loads its AI Overviews after the page, so a normal capture stores an empty shell and records "no AI answer here". In a camera-and-imaging market, our capture logs put true answer-layer presence at 43% while a standard capture saw 17% — barely 40% of what was actually there. The rest read as silence, with no error and no warning.
None of this means citations do not matter. It means a citation rate, alone, is sand. You need the rungs above it.
How is AI visibility measured? The five-rung ladder
We measure AI visibility as a ladder because a single number collapses questions that have different answers and different difficulty. Each rung asks a sharper question than the one below, and — the point — each is harder to fake. This is what generative engine optimization (GEO) is actually optimising: not a citation tally, but a climb.
| Rung | The question | What it really measures | Can spam earn it? |
|---|---|---|---|
| 0 | Is the answer layer even there? | Whether AI answers this kind of question at all — captured correctly, per surface | It's the instrument, not the score |
| 1 | Are you in the room? | Answers where you are used as a source, on each surface | Yes — and we say so |
| 2 | Are you the recommendation? | Whether the answer text names you as the option — and how prominently, and whether it speaks well of you | Harder — needs preference, not reference |
| 3 | Will it hold? | Whether the position is supported by search rank, owned questions and topic authority | No |
| 4 | Does it matter? | The real search demand behind the questions you win | No |
Read top to bottom, it is a diagnosis, not a scoreboard. A brand strong on rung 1 and absent from rung 2 is being read by the models and recommended by none of them. That is the commonest and most expensive failure, and a citation count hides it completely. A brand that wins rung 2 in one category and not another learns exactly where its authority is real — and where it is merely a source.
And the ladder defends itself. A sceptic who attacks rung 1 — "citations are gameable" — is right, and we agree out loud. Then we point at rung 2: being named as the recommendation is not something spam earns. The framework survives its own harshest reading, which is the only kind worth putting in front of a board.
The gap most brands miss: the source versus the name
Here is the finding that changes how a marketing leader should think about this. You can be the source an AI answer is built from — quoted, linked, load-bearing — and never be the brand the reader sees.
The model reads your page, uses your facts, assembles a recommendation, and names someone else. To every dashboard that counts citations, you are winning. To the human reading the answer, you do not exist. Citation counts presence; being named is preference. A first-position mention, framed as the one to choose, is a completely different asset from a footnote three-quarters of the way down — and only an instrument that reads the answer, rather than tallying the references under it, can tell them apart.
So "get cited more" is the wrong brief. The work is to move from reference to preference — from a source the machine trusts to the name it hands the reader. That is the rung the citation-counters cannot see and the narrative tribe cannot measure. It is measurable. It is just not a tally.
What no AI-visibility metric can tell you
The most useful thing a measurement framework does is state its own edges, so we will.
No rung measures conversion. There is no clickthrough, no revenue attribution, from an AI answer — not for us, not for anyone, today. Rung 4 weights your visibility by the real demand behind each question, which is exposure, not outcome. Anyone selling you an "AI visibility ROI" is inferring it.
The surfaces are not one surface. Google AI Overviews, ChatGPT and Gemini cite different worlds, and pooling them reverses the answer. In our capture sets YouTube is the single most-cited domain in Google's AI Overviews — and close to invisible inside the chatbots; Reddit supplies well over half of ChatGPT's community citations and a fraction of the others. A channel that looks dead on one surface leads on another, so we never report an AI-visibility number without naming the surface it came from.
Movement inside the noise band is not movement. Given the 15-to-20% same-day flip, we do not call a sub-20% swing in presence a change. Domain-level standing is far steadier — it averages over many answers — so that is what we track for trend, with presence as context, never headline.
None of these are caveats bolted on at the end. They are the reason the numbers above them can be trusted. A framework armed to reject a bad metric is one you can believe when it reports a good one.
Some take-aways
Stop asking "how often is my brand cited by AI". Start asking the ladder's questions, in order: is the answer even there, am I in the room, am I the recommendation, will it hold, and does the demand behind it justify the fight. The rung you are stuck on tells you what to do next — and it is almost never "get cited more".
This also settles the two tribes. The counters are right that presence is real and measurable; the narrative tribe is right that a citation tally is not the point. The ladder holds both: measure everything from present to recommended, and be honest about the rung that actually moves a buyer.
We run this ladder for our own market and for clients across health, consumer electronics and retail, on the same skeleton every time — because repeatability is part of what is being sold. The rungs do not change; only the numbers in the column do. For the whole-market view behind it, see why market research is four silos that don't talk and what market intelligence actually is, or explore the platform.
Method note: figures are Theia measurements across live consumer markets (a UK private-healthcare market and a global camera-and-imaging brand's category), captured with Google's asynchronous answer layer requested explicitly, and reported as a rate over repeated captures. Surface-level results are never pooled. We publish the result and the reasoning, never the method that produces them.
Frequently asked
- Does being cited by AI mean the AI recommends my brand?
- No. A citation means the model used your page as a source while building its answer. Being recommended means the answer names you as the option a reader should choose. They are different rungs — a brand can be the source an answer is built from and never the brand the reader sees.
- Is 'AI citation rate' a reliable metric?
- On its own, no. It is the most gameable rung on the ladder, and it moves 15 to 20% between two identical captures on the same day. Read it beside the rungs that spam cannot earn: whether you are named as the recommendation, and whether that naming holds.
- What is a good AI visibility score?
- There isn't a single trustworthy one. A one-number 'AI visibility score' collapses being present, being used as a source and being named as the recommendation into a figure that also moves 15 to 20% between two identical same-day captures. Read the ladder instead — the rung you are stuck on is the useful number, not the average.
- Can I measure whether AI recommendations convert to sales?
- Not today. There is no clickthrough or revenue attribution from an AI answer available to anyone. Any 'AI visibility ROI' figure is inferred, not measured. We weight visibility by the demand behind each question — that is exposure, not outcome, and we say so.