Measurement

How can I track my brand’s visibility in AI search?

Generated answers have no positions, vary between identical runs, and are archived nowhere by default. Measuring them needs a different method from rank tracking, not a modified one.

Klepha insight panel showing tracked prompts, brand mentions and cited sources across AI engines.

Rank tracking does not transfer to AI search, and trying to adapt it produces numbers that look precise and mean nothing. The surface is different enough to need a different method.

Why rank tracking does not transfer

Three properties of generated answers break the ranking model completely.

There are no positions. An answer names three companies or it does not. There is no position four, no page two, and no numbered list to occupy a slot in.

Answers vary between identical runs. Generating text involves sampling, so the same prompt produces different outputs. A single measurement is a sample of a distribution, not a state.

Nothing is archived. Search results can be re-checked; a generated answer that has passed is gone unless you stored it. Without a stored baseline, every change is unfalsifiable.

You are not measuring a position. You are sampling a distribution, and you have to keep the samples.

The four metrics

Mention rate. Across your prompt set, in what proportion of answers is your brand named at all? Coarsest measure, usually moves first, easiest to explain.

Citation rate. In what proportion is one of your URLs actually cited as a source? Stricter and more valuable — being named is recognition, being cited is traffic and authority.

Share of voice. Of all brands named across your prompt set, what proportion are you? This is the competitive measure and the one that survives an executive audience, because it is immediately comparable.

Accuracy. When you are mentioned, is what is said correct? Frequently overlooked and occasionally the most urgent finding — an assistant confidently quoting pricing you abandoned two years ago is worse than not being mentioned.

Track all four. They move independently, and a programme that raises mention rate while citation rate falls has made the brand more discussed and the website less useful.

Building the prompt set

The prompt set determines everything downstream. A bad one produces a confident, worthless number.

Write questions, not keywords. “CRM software” is a keyword. “What’s the best CRM for a 12-person nonprofit that needs Xero integration?” is a prompt. The second is how people actually talk to assistants and the one you can realistically win.

Cover the buying journey. Category discovery (“what tools exist for X”), comparison (“X vs Y”), constraint-driven (“best X for [specific situation]”), problem-led (“how do I solve Z”), and brand-direct (“is [your brand] any good”). Each behaves differently.

Include the uncomfortable ones. “[Your brand] alternatives” and “problems with [your brand]” are asked by real buyers and frequently produce the most actionable findings.

Then freeze it. A set that changes when results are disappointing is worthless for trend analysis and destroys credibility the first time someone notices. Add prompts to a clearly-labelled second cohort rather than editing the original.

Running it properly

  1. Three engines minimum — ChatGPT, Gemini and Perplexity. They behave differently enough that one tells you little about the others.
  2. Two to three runs per prompt per engine. Non-negotiable given the variance. One run is a coin flip.
  3. Clean sessions. No conversation history, no personalisation, no memory of previous questions. Otherwise you are measuring your own account rather than the default answer.
  4. Record the full answer text, not just a yes/no on mention. The wording, the ordering, and the cited URLs are where the actionable detail lives.
  5. Record the citations separately. The cited domains are a ranked map of what influences your category, and it is the most useful single output of the whole exercise.
  6. Same day, same cadence. Weekly, consistently, so the trend is not confounded by timing.

Skip the spreadsheet

Klepha runs your prompts across three engines on a schedule and stores every answer, citation and competitor mention automatically.

Run my free scan

Reading the results

Read trends, never single runs. Four to eight weeks before drawing any conclusion. A week-on-week change smaller than your observed variance is noise, and you will only know your variance because you ran each prompt multiple times.

Segment by engine. Perplexity reflects changes within days and will improve first. Gemini follows Google’s recrawl cadence. ChatGPT lags both, because it depends on third-party sources catching up. A programme that has moved Perplexity and not ChatGPT is normal at six weeks and a concern at six months.

Segment by prompt type. Broad category questions favour the largest brands. Constrained, specific questions are winnable. If you are absent from the first group and present in the second, you are correctly positioned rather than invisible.

Watch the citation list more than your own score. The domains being cited tell you where the corroboration work belongs, and that is usually the highest-value finding in the whole report.

Doing it manually

A credible baseline costs an afternoon and nothing else.

Twenty prompts, three engines, two runs each is 120 queries. Paste each into a fresh session, copy the answer into a spreadsheet with columns for date, engine, prompt, run number, brands named in order, URLs cited, and any factual error about you.

That is genuinely enough to establish where you stand and to produce the artefact that wins an internal argument. What it is not is sustainable — repeating 120 queries every week by hand is where manual tracking dies, usually around week three.

What the first four weeks look like

Week 1. Write and freeze the prompt set. Run it fully — three engines, two runs each — and record everything including citations. This is your baseline and it is the only week that feels like a project.

Week 2. Run it again unchanged. Compare against week one. The difference between the two runs is your variance, and knowing it is what stops you over-reading every subsequent movement.

Week 3. Run it again, and start the citation map — a simple list of every third-party domain cited across all answers, ranked by frequency. This is usually the most valuable artefact produced in the whole first month.

Week 4. Run it again and write the first trend read. Four data points is enough to distinguish a direction from noise, and it is the point at which the exercise starts informing decisions rather than just describing a state.

What to do when the numbers are bad

The first honest baseline is usually worse than expected, and the reaction to it determines whether the programme survives. Two rules help.

Present the baseline as a starting position rather than a performance result, and do it before anyone else discovers it. A number nobody has seen before is information; the same number presented after a stakeholder found it themselves is a failure.

And pair it immediately with the diagnosis. “We are named in two of twenty answers, and the reason is that our product pages render client-side so no AI crawler has ever read them” is a fixable problem with an owner. The same number without a cause is just bad news.

Five measurement pitfalls

Testing in a logged-in session. Conversation history and personalisation shape the answer. You end up measuring your own account rather than what a stranger sees, and the results flatter you.

Single runs. Given the variance, a one-run measurement is a coin flip presented as data. Every conclusion drawn from it is unsafe.

Prompts phrased as keywords. “Best CRM software” invites the model to name the three largest players. Almost nobody outside the market leaders can win that phrasing, and treating it as the benchmark produces a permanently depressing chart that measures nothing you can change.

Changing the prompt set. Every edit resets your history. The temptation arrives precisely when results are disappointing, which is the worst possible moment to lose comparability.

Not recording citations. Teams record whether they were mentioned and discard the source list. The source list is the most actionable output in the entire exercise — it is a ranked map of what influences your category.

When to buy tooling

Three conditions make tooling worth paying for.

You need weekly cadence sustained. The manual method works once and collapses under repetition.

You need stored history you can query. Comparing this week to eleven weeks ago, per prompt, per engine, is where the real insight is — and a spreadsheet of 120 rows a week becomes unusable quickly.

You need to know why, not just whether. Monitoring tells you that you were not mentioned. Diagnosing why requires checking whether AI crawlers can read your pages at all, which no amount of prompt sampling reveals. This is the distinction between a monitor and a platform, covered in our guide to what a GEO platform does.

Whatever you use, insist on one thing: the ability to open the raw answer behind any number. A score you cannot trace back to a stored response is a claim rather than a measurement, and in a surface this variable that distinction is the whole basis of trusting the data.

Sources

Frequently asked questions

How can I track my brand’s visibility in AI search?

Build a fixed set of twenty to sixty buyer questions, run them against ChatGPT, Gemini and Perplexity on a schedule, and record four things per run: whether your brand was mentioned, whether your URLs were cited, which competitors appeared, and your share of the total answer space. Store the raw answers, because responses vary between identical runs.

Why can I not just track positions like in SEO?

Because generated answers have no numbered results. There is no position one. What exists instead is whether you were named, where in the answer, and which sources were cited. The metric has to change to match the surface, and applying a ranking model to it produces numbers that do not mean anything.

How many prompts should I track?

Twenty is a workable minimum for a focused business; sixty covers a broader category or several product lines. Realism matters more than count — prompts phrased the way buyers actually ask. Keep the set fixed, because changing it resets your history and makes trends meaningless.

How often should I run AI visibility scans?

Weekly is enough for most businesses and gives a readable trend within four to eight weeks. Daily is worth it during an active optimisation programme when you want to see whether a change landed. Because answers vary between runs, more frequent measurement mostly buys you a clearer view of the variance rather than faster signal.

Why do my AI visibility results change between runs?

Generating text involves sampling, so identical prompts produce different outputs by design, and the retrieved candidate set can also differ. This is inherent rather than a fault. It is the single strongest argument for running each prompt several times and storing every response.

Can I track AI visibility for free?

You can do a credible manual baseline for free — twenty questions, three engines, two runs each, recorded in a spreadsheet takes an afternoon. Free scans also exist. What you cannot do manually is sustain it weekly at scale with stored history, which is where tooling earns its cost.

Garry Charter

SEO Specialist · Klepha

Twelve years in search, covering technical SEO, keyword research and — since generative search arrived — answer engine and generative engine optimization. Writes Klepha's guides on ranking in Google and being cited by AI assistants.