How can I track my brand’s visibility in AI search?
Generated answers have no positions, vary between identical runs, and are archived nowhere by default. Measuring them needs a different method from rank tracking, not a modified one.
Rank tracking does not transfer to AI search, and trying to adapt it produces numbers that look precise and mean nothing. The surface is different enough to need a different method.
Why rank tracking does not transfer
Three properties of generated answers break the ranking model completely.
There are no positions. An answer names three companies or it does not. There is no position four, no page two, and no numbered list to occupy a slot in.
Answers vary between identical runs. Generating text involves sampling, so the same prompt produces different outputs. A single measurement is a sample of a distribution, not a state.
Nothing is archived. Search results can be re-checked; a generated answer that has passed is gone unless you stored it. Without a stored baseline, every change is unfalsifiable.
You are not measuring a position. You are sampling a distribution, and you have to keep the samples.
The four metrics
Mention rate. Across your prompt set, in what proportion of answers is your brand named at all? Coarsest measure, usually moves first, easiest to explain.
Citation rate. In what proportion is one of your URLs actually cited as a source? Stricter and more valuable — being named is recognition, being cited is traffic and authority.
Share of voice. Of all brands named across your prompt set, what proportion are you? This is the competitive measure and the one that survives an executive audience, because it is immediately comparable.
Accuracy. When you are mentioned, is what is said correct? Frequently overlooked and occasionally the most urgent finding — an assistant confidently quoting pricing you abandoned two years ago is worse than not being mentioned.
Track all four. They move independently, and a programme that raises mention rate while citation rate falls has made the brand more discussed and the website less useful.
Building the prompt set
The prompt set determines everything downstream. A bad one produces a confident, worthless number.
Write questions, not keywords. “CRM software” is a keyword. “What’s the best CRM for a 12-person nonprofit that needs Xero integration?” is a prompt. The second is how people actually talk to assistants and the one you can realistically win.
Cover the buying journey. Category discovery (“what tools exist for X”), comparison (“X vs Y”), constraint-driven (“best X for [specific situation]”), problem-led (“how do I solve Z”), and brand-direct (“is [your brand] any good”). Each behaves differently.
Include the uncomfortable ones. “[Your brand] alternatives” and “problems with [your brand]” are asked by real buyers and frequently produce the most actionable findings.
Then freeze it. A set that changes when results are disappointing is worthless for trend analysis and destroys credibility the first time someone notices. Add prompts to a clearly-labelled second cohort rather than editing the original.
Running it properly
- Three engines minimum — ChatGPT, Gemini and Perplexity. They behave differently enough that one tells you little about the others.
- Two to three runs per prompt per engine. Non-negotiable given the variance. One run is a coin flip.
- Clean sessions. No conversation history, no personalisation, no memory of previous questions. Otherwise you are measuring your own account rather than the default answer.
- Record the full answer text, not just a yes/no on mention. The wording, the ordering, and the cited URLs are where the actionable detail lives.
- Record the citations separately. The cited domains are a ranked map of what influences your category, and it is the most useful single output of the whole exercise.
- Same day, same cadence. Weekly, consistently, so the trend is not confounded by timing.
Skip the spreadsheet
Klepha runs your prompts across three engines on a schedule and stores every answer, citation and competitor mention automatically.
Run my free scanReading the results
Read trends, never single runs. Four to eight weeks before drawing any conclusion. A week-on-week change smaller than your observed variance is noise, and you will only know your variance because you ran each prompt multiple times.
Segment by engine. Perplexity reflects changes within days and will improve first. Gemini follows Google’s recrawl cadence. ChatGPT lags both, because it depends on third-party sources catching up. A programme that has moved Perplexity and not ChatGPT is normal at six weeks and a concern at six months.
Segment by prompt type. Broad category questions favour the largest brands. Constrained, specific questions are winnable. If you are absent from the first group and present in the second, you are correctly positioned rather than invisible.
Watch the citation list more than your own score. The domains being cited tell you where the corroboration work belongs, and that is usually the highest-value finding in the whole report.
Doing it manually
A credible baseline costs an afternoon and nothing else.
Twenty prompts, three engines, two runs each is 120 queries. Paste each into a fresh session, copy the answer into a spreadsheet with columns for date, engine, prompt, run number, brands named in order, URLs cited, and any factual error about you.
That is genuinely enough to establish where you stand and to produce the artefact that wins an internal argument. What it is not is sustainable — repeating 120 queries every week by hand is where manual tracking dies, usually around week three.
What the first four weeks look like
Week 1. Write and freeze the prompt set. Run it fully — three engines, two runs each — and record everything including citations. This is your baseline and it is the only week that feels like a project.
Week 2. Run it again unchanged. Compare against week one. The difference between the two runs is your variance, and knowing it is what stops you over-reading every subsequent movement.
Week 3. Run it again, and start the citation map — a simple list of every third-party domain cited across all answers, ranked by frequency. This is usually the most valuable artefact produced in the whole first month.
Week 4. Run it again and write the first trend read. Four data points is enough to distinguish a direction from noise, and it is the point at which the exercise starts informing decisions rather than just describing a state.
What to do when the numbers are bad
The first honest baseline is usually worse than expected, and the reaction to it determines whether the programme survives. Two rules help.
Present the baseline as a starting position rather than a performance result, and do it before anyone else discovers it. A number nobody has seen before is information; the same number presented after a stakeholder found it themselves is a failure.
And pair it immediately with the diagnosis. “We are named in two of twenty answers, and the reason is that our product pages render client-side so no AI crawler has ever read them” is a fixable problem with an owner. The same number without a cause is just bad news.
Five measurement pitfalls
Testing in a logged-in session. Conversation history and personalisation shape the answer. You end up measuring your own account rather than what a stranger sees, and the results flatter you.
Single runs. Given the variance, a one-run measurement is a coin flip presented as data. Every conclusion drawn from it is unsafe.
Prompts phrased as keywords. “Best CRM software” invites the model to name the three largest players. Almost nobody outside the market leaders can win that phrasing, and treating it as the benchmark produces a permanently depressing chart that measures nothing you can change.
Changing the prompt set. Every edit resets your history. The temptation arrives precisely when results are disappointing, which is the worst possible moment to lose comparability.
Not recording citations. Teams record whether they were mentioned and discard the source list. The source list is the most actionable output in the entire exercise — it is a ranked map of what influences your category.
When to buy tooling
Three conditions make tooling worth paying for.
You need weekly cadence sustained. The manual method works once and collapses under repetition.
You need stored history you can query. Comparing this week to eleven weeks ago, per prompt, per engine, is where the real insight is — and a spreadsheet of 120 rows a week becomes unusable quickly.
You need to know why, not just whether. Monitoring tells you that you were not mentioned. Diagnosing why requires checking whether AI crawlers can read your pages at all, which no amount of prompt sampling reveals. This is the distinction between a monitor and a platform, covered in our guide to what a GEO platform does.
Whatever you use, insist on one thing: the ability to open the raw answer behind any number. A score you cannot trace back to a stored response is a claim rather than a measurement, and in a surface this variable that distinction is the whole basis of trusting the data.
Sources
- Google — 2.5 billion monthly active users for AI Overviews and 1 billion for AI Mode, stated 19 May 2026.
- SERP feature tracking, 2026 — AI Overview presence averaging ~32.5% of monitored queries within a 31.02%–34.40% band.
- Vercel and MERJ — analysis of 500M+ GPTBot fetches finding zero JavaScript execution; same behaviour for ClaudeBot and PerplexityBot.
Frequently asked questions
How can I track my brand’s visibility in AI search?
Build a fixed set of twenty to sixty buyer questions, run them against ChatGPT, Gemini and Perplexity on a schedule, and record four things per run: whether your brand was mentioned, whether your URLs were cited, which competitors appeared, and your share of the total answer space. Store the raw answers, because responses vary between identical runs.
Why can I not just track positions like in SEO?
Because generated answers have no numbered results. There is no position one. What exists instead is whether you were named, where in the answer, and which sources were cited. The metric has to change to match the surface, and applying a ranking model to it produces numbers that do not mean anything.
How many prompts should I track?
Twenty is a workable minimum for a focused business; sixty covers a broader category or several product lines. Realism matters more than count — prompts phrased the way buyers actually ask. Keep the set fixed, because changing it resets your history and makes trends meaningless.
How often should I run AI visibility scans?
Weekly is enough for most businesses and gives a readable trend within four to eight weeks. Daily is worth it during an active optimisation programme when you want to see whether a change landed. Because answers vary between runs, more frequent measurement mostly buys you a clearer view of the variance rather than faster signal.
Why do my AI visibility results change between runs?
Generating text involves sampling, so identical prompts produce different outputs by design, and the retrieved candidate set can also differ. This is inherent rather than a fault. It is the single strongest argument for running each prompt several times and storing every response.
Can I track AI visibility for free?
You can do a credible manual baseline for free — twenty questions, three engines, two runs each, recorded in a spreadsheet takes an afternoon. Free scans also exist. What you cannot do manually is sustain it weekly at scale with stored history, which is where tooling earns its cost.