How can I monitor AI mentions of my brand?
Tracking visibility asks whether you appear. Monitoring asks what is said when you do — and the second question produces more urgent findings than the first.
Most teams that start measuring AI search begin with visibility: do we appear? That is the right first question and it is not the most urgent one.
The more urgent question is what gets said when you do appear. An assistant that names you and states pricing you abandoned two years ago is actively costing you deals in a way that absence does not.
Monitoring vs tracking
These are different jobs with different prompt sets and different owners.
Visibility tracking asks: across the category questions our buyers ask, how often are we named and cited? It is a marketing metric, it moves slowly, and it is measured against competitors.
Monitoring asks: when someone asks about us specifically, what does the assistant say, and is it true? It is a reputational and communications concern, it can change abruptly, and it is measured against reality.
A programme that only does the first will report a healthy trend while an assistant tells every prospect your product lacks a feature you shipped last year.
What to monitor
Five things per answer, recorded every run.
Presence. Were you named at all, and where in the answer.
Framing. How were you characterised — the budget option, the enterprise choice, the one for beginners? Framing sticks, propagates between sources, and is much harder to change than a factual error.
Factual claims. Every specific statement about pricing, features, company size, availability, integrations. These are what buyers act on.
Comparisons. Who you were compared to and how you fared. This tells you which competitive set the model places you in, which is frequently not the one you think.
Sources. Which URLs the answer cited. When something is wrong, this is where the wrong thing came from, and it is the only actionable route to fixing it.
The reputational prompt set
Different from your visibility set, and deliberately uncomfortable.
- “What is [brand]?” — the baseline description.
- “Is [brand] any good?” — the evaluative summary.
- “How much does [brand] cost?” — the most commonly wrong answer, and the most commercially damaging.
- “[Brand] alternatives” — who you are being replaced by.
- “Problems with [brand]” — the one nobody wants to run, and frequently the most useful.
- “[Brand] vs [competitor]” — for each of your two or three main competitors.
- “Is [brand] legitimate / safe / reliable?” — asked more often than most companies assume.
- “Who owns [brand]?” and similar factual questions where stale information persists.
Run each against all three engines, at least twice, in clean sessions with no conversation history.
See what is being said about you
A free scan returns the real answers ChatGPT, Gemini and Perplexity give about your brand — and the sources behind them.
Run my free scanAuditing for accuracy
Take each factual claim in the stored answers and label it: correct, outdated, or wrong.
Outdated is the most common category by a distance. Pricing that changed, a feature that shipped, a positioning you moved on from, a founder who left. The information was true once and the web has not caught up.
Wrong is rarer and more serious. A confusion with a similarly-named company, a misattributed incident, an invented feature limitation.
Then trace each one to its source using the citation list. In most cases the origin is identifiable within minutes — an old page of yours that still ranks, a third-party listing nobody has updated since 2024, or a review written against a previous version of the product.
That traceability is the entire value of storing raw answers with their citations. Without it you know something is wrong and have no route to changing it.
When it says something wrong
There is no takedown process for organic AI answers. What changes the output is changing the inputs, in this order.
1. Publish the correct fact yourself, on a crawlable, server-rendered page, stated plainly enough to be extracted. A pricing page that renders client-side cannot correct a pricing error, because no AI crawler will ever read it. Make the correction explicit — a dated statement is more useful than a silent edit.
2. Fix your own stale pages. Old blog posts, outdated documentation, superseded landing pages. If your own site still carries the wrong version, you are the source of the error.
3. Correct the third-party sources that are actually cited. Prioritise by the citation list rather than by reach. Five directory listings and one review site carrying the old price will keep reproducing it indefinitely.
4. Add structured data reflecting the correct facts — Organization, Product, Offer — matching the visible text exactly.
5. Rescan on a schedule. Correction propagates at the speed of recrawling and of third-party updates, so expect weeks rather than days, and longer on ChatGPT than on Perplexity.
Monitoring competitors too
Often more informative than monitoring yourself, and rarely done.
Run the same reputational set against your two or three main competitors. Three things fall out.
The category’s dimensions. How the model characterises each player tells you which axes it thinks matter — price, ease, scale, specialisation. If your positioning is not one of those axes, that is why it does not register.
The source map. Which sites shape the view of each competitor. That is your corroboration target list, derived rather than guessed.
Direct comparisons. Whether an assistant recommends a competitor over you and on what grounds. This is the most commercially specific finding available anywhere in AI search measurement, and it frequently identifies a weakness you can address directly.
Who should own this
Monitoring sits awkwardly between functions, which is why it frequently belongs to nobody.
The measurement is a marketing or SEO task — same tooling, same prompt discipline as visibility tracking. But the findings are communications findings: a wrong price, a stale feature list, a competitor being recommended over you on specific grounds. Those need someone who can commission a correction, contact a review site, or brief a press response.
The workable arrangement is that whoever owns AI visibility runs the scans, and anything in the “wrong” or “damaging framing” category is escalated to communications with the stored answer and the source attached. What does not work is a monthly report circulated to nobody in particular, which is the default state at most companies currently doing this at all.
Building the monitoring record
Keep one row per run with the date, engine, prompt, the full answer text, the brands named in order, the URLs cited, and a flag for any factual claim about you. That is enough structure to answer every question that will be asked later, and it is simple enough to survive being maintained by someone under time pressure.
The field that matters most in hindsight is the citation list, because when something wrong appears the first question is where it came from. A record of presence without sources tells you a problem exists and gives you no route to fixing it.
What to escalate immediately
Three findings justify interrupting someone rather than waiting for the monthly report: a materially wrong price, a claim that your product lacks something it has, and any suggestion that your company is unreliable, insolvent or unsafe. Each of those is actively costing deals for as long as it persists, and each traces to a specific source you can go and correct.
Changing how you are framed
Factual errors are correctable. Framing is harder and matters more over time.
If assistants consistently describe you as “the budget option” or “best for beginners” and that is not your positioning, no single page fixes it. Framing comes from the aggregate of how independent sources characterise you, and it propagates — one influential review describing you a certain way gets echoed by later writers, which reinforces it in every subsequent answer.
Three things shift it, all slowly. Publish material that demonstrates the positioning you want rather than asserting it, since a technical depth claim is only credible alongside technical depth. Target the specific sources that carry the current framing, because correcting the two or three most-cited ones does more than a hundred new mentions. And be patient — framing changes over quarters, and the earliest visible signal is usually a shift in which competitors you get compared against rather than a change in the adjectives used.
Cadence and alerting
Monthly is sufficient for a stable brand. Move to weekly in three situations: after a pricing or positioning change, while actively correcting an error, and during a period of press attention where new sources are being created quickly.
Alerting is harder than it sounds because of variance. Generated answers differ between identical runs, so a naive alert on “brand not mentioned” will fire constantly and be ignored within a fortnight. Two rules make alerting usable: alert on factual claims changing rather than on presence, and require the change to persist across at least two consecutive runs before it fires.
One final discipline: keep the history permanently. The question “when did they start saying that?” is unanswerable without it, and it is the first question anyone senior asks when a wrong claim surfaces.
Sources
- Vercel and MERJ — analysis of 500M+ GPTBot fetches finding zero JavaScript execution; same behaviour for ClaudeBot and PerplexityBot.
- Google — 2.5 billion monthly active users for AI Overviews and 1 billion for AI Mode, stated 19 May 2026.
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande — “GEO: Generative Engine Optimization”, ACM SIGKDD 2024 (arXiv:2311.09735).
Frequently asked questions
How can I monitor AI mentions of my brand?
Build a reputational prompt set — questions about your brand directly, alternatives to you, problems with you, and your pricing — and run them against ChatGPT, Gemini and Perplexity on a schedule. Record not just whether you are named but what is said, whether it is accurate, and which sources the claim came from. Store every answer, because they vary between runs.
What is the difference between monitoring and visibility tracking?
Visibility tracking asks whether you appear across category questions. Monitoring asks what is said about you when you do appear, and whether it is correct. They use different prompt sets and produce different urgencies — an assistant confidently stating pricing you abandoned two years ago is more urgent than a few points of mention rate.
What do I do if ChatGPT says something wrong about my company?
Publish the correct fact on a crawlable, structured page of your own, stated plainly enough to be extracted. Then update or request correction on the third-party sources carrying the stale version, prioritising the ones the assistants actually cite. Then rescan on a schedule, because correction propagates at the speed of recrawling rather than immediately.
Can I get an AI assistant to remove false information about my business?
There is no removal request process for organic answers in the way there is for search results. What changes the output is changing the inputs — publishing the correct information in a retrievable form and correcting the third-party sources the model draws on. For genuinely defamatory content, the underlying source is the appropriate target rather than the assistant.
How often should I monitor AI mentions?
Monthly is sufficient for most brands, weekly if you have recently changed pricing, positioning or product, or if you are actively correcting an error. Because answers vary between runs, run each prompt more than once and read the trend rather than reacting to a single response.
Should I monitor what AI says about competitors?
Yes, and it is often more informative than monitoring yourself. How competitors are described tells you what the model considers the category’s dimensions, which sources shape that view, and where your positioning is failing to register. It also catches the case where an assistant recommends a competitor by directly comparing them to you.