The short version
- Five numbers are enough: share of voice, citation rate, average position, answers lost, and sentiment.
- Measure per prompt, not per day. A daily average tells you the weather; the prompt tells you which deal you lost.
- None of these are website traffic. There is no impression to log and usually no click — if a tool reports “AI referrals” as clicks, ask precisely how it attributes them.
- Cadence that works: weekly prompt tracking, monthly trend review, quarterly benchmark against competitors.
Why your analytics can’t see this#
A buyer asks an assistant which tool they should use. Three names come back. Yours isn’t one of them. Nothing about that event reaches your analytics — no impression, no session, no bounce. The loss is invisible, which means it never gets a budget line or an owner.
So measurement has to be active rather than passive. Instead of waiting for traffic to describe demand, you ask the engines the questions your buyers ask, on a schedule, and record what comes back. That is the whole method. Everything below is what to record.
The five numbers#
1. Share of voice
Of all the brand mentions across your tracked prompts, what fraction are you? If ten prompts return five names each, that’s fifty mentions — nine of them yours is 18% share of voice.
This is the headline number because it is inherently competitive. Your score going up while a rival’s goes up faster is not a win, and an absolute score can’t tell you that. Pair it with your rank in the category and the multiple between you and the leader.
2. Citation rate
What share of answers link to you as a source? Distinct from being mentioned: an assistant can recommend you without citing you, and can cite you while recommending somebody else.
Citation rate is the most actionable of the five, because the sources behind an answer are a literal to-do list of places to go get mentioned. Track which domains recur, and how many of them you appear on.
3. Average position
When you are named, where in the answer? First is not fourth. Buyers read generated answers top-down and the first name carries most of the intent, so presence alone hides the difference between leading a category and being the also-ran at the end of a list.
Weight your score by position, or you will celebrate a scan where you moved from first to fourth in every answer.
4. Answers lost
The count of prompts where a competitor was named and you weren’t. This is the number to take to a board, because it is the closest thing in the set to countable missed revenue.
It also localises the problem. Ten lost answers spread across ten topics is a brand-awareness issue; ten concentrated on one topic is a content gap you can close this month.
5. Sentiment
Being mentioned is not the same as being mentioned well. Engines summarise the tone of what they’ve read, so a brand can appear in every answer as the expensive option, the complicated one, or the one people leave.
Track the tone, and track the themes behind it with quotes attached. The quote is the useful part: it tells you which review, thread or article taught the model to say that.
What not to measure#
Clicks out of AI answers. Some assistants pass a referrer, many don’t, and a lot of the influence never produces a click at all — the buyer reads three names and searches for one directly a day later. Any number labelled “AI referrals” deserves a hard question about how it was attributed, including when it comes from us.
A single blended score, on its own. One number going up is reassuring and directs no work. Keep the score for the trend line, and act on the five components.
Anything measured once. Generated answers vary between runs. A single scan is an anecdote — you need the same prompts repeated on a schedule before a change means anything.
Measure per prompt, and think per topic#
Record at the prompt level, because that is the resolution at which the outcome is decided and the only resolution at which a fix is obvious. “Visibility fell 4 points” starts an argument; “we lost best X for small teams to a rival who got written up on Reddit” starts work.
But roll up to topics when you plan. A study of 50,000 brands in ChatGPT found visibility behaves as a topic-level property rather than a keyword-level one [1] — so a cluster of prompts covering one subject is the meaningful grouping, not any individual phrasing.
A cadence that survives contact with a real week#
| Rhythm | What you look at | Decision it drives |
|---|---|---|
| Weekly | Prompt-level movement, answers lost | Which single gap to work on next |
| Monthly | Share of voice trend, citation rate, sentiment shift | Whether last month’s work landed |
| Quarterly | Rank against the competitive set, topic coverage | Where to point the content budget |
Cadence and metric definitions cross-checked against public sources, August 2026: Sanbi, HubSpot, Page One Power, topic-level findings from Semrush.
How StayFound reports these#
For transparency about what you’d actually see: the Overview tab shows a position-weighted score out of 100 with a plain-language verdict, your share of voice with rank in the category, how many engines mention you, and how many cited sources you own out of the total found. Competitors gives share-of-voice bars and the gap in percentage points to each rival. Citations lists the source domains with whether you’re on each one. Analytics covers tone and the themes behind it, with quotes.
Deliberately absent: any click or referral figure. We measure what assistants say and cite, not traffic, and we’d rather the gap be stated than implied.
If you want the numbers for your own domain without setting anything up, we’ll run the report and email it. If you’d rather compare the tools first, we priced the whole field — and if the terminology is still slippery, GEO vs AEO vs SEO sorts it out.