How can I improve my AI visibility?
AI visibility is a number, and most teams try to improve it before deciding what it measures. Definition first, baseline second, interventions third — in that order it is tractable.
“Improve our AI visibility” is a goal in the same way “increase traffic” is a goal. Before it becomes work, someone has to decide what number is being improved and what it currently is.
Most teams skip both steps and go straight to interventions, which is why so many AI visibility programmes cannot say afterwards whether they worked.
Define the metric first
AI visibility is not one number. It is four, and they behave differently.
Mention rate. Across your prompt set, how often is your brand named at all? This is the coarsest measure and usually the one that moves first.
Citation rate. How often is one of your URLs actually cited as a source? This is stricter and more valuable — being named is recognition, being cited is traffic and authority.
Share of voice. Of all the brands named across your prompt set, what proportion are you? This is the competitive measure and the one that survives contact with an executive audience.
Accuracy. When you are mentioned, is what is said correct? Frequently overlooked and occasionally the most urgent — an assistant confidently stating pricing you abandoned two years ago is worse than absence.
Pick which one you are managing before you start. They can move in opposite directions, and a programme that raises mention rate while citation rate falls has usually made the brand more famous and the website less useful.
Building an honest baseline
Four rules, and skipping any of them produces a number you cannot trust later.
Fix the prompt set. Twenty to sixty questions phrased the way a buyer would actually ask, written once and then left alone. Changing the set resets your history, and a set that quietly changes when results are disappointing is worthless.
Run each prompt multiple times. Generated answers vary between identical runs. A baseline built from single runs is a snapshot of randomness. Two or three runs per prompt per engine is the practical minimum.
Cover more than one engine. Being strong in Perplexity tells you very little about ChatGPT. They differ in how they retrieve, how openly they cite, and how heavily they weight third-party description.
Store the raw answers. This is the step most often skipped and the one that makes everything else auditable. A score without the response behind it cannot be checked, cannot be trended honestly, and cannot settle a disagreement about whether something improved.
A visibility score you cannot trace back to a stored answer is a claim, not a measurement.
Interventions by leverage
Once you have a baseline, the interventions sort cleanly into two groups by how fast they move the number.
| Intervention | Speed | Ceiling |
|---|---|---|
| Fix JavaScript-only content | Days to weeks | Very high if it was broken |
| Unblock AI crawlers | Days to weeks | Very high if it was blocked |
| Rewrite pages answer-first | 2–6 weeks | High |
| Add original data and sources | 3–8 weeks | High and durable |
| Entity clarity and schema | 4–8 weeks | Moderate |
| Third-party corroboration | 1–2 quarters | Highest, hardest |
What moves it in weeks
Retrievability. Load your key pages with JavaScript disabled. If the answer is not in the raw HTML, AI crawlers never saw it — Vercel and MERJ found zero JavaScript execution across more than 500 million GPTBot fetches, with the same behaviour from ClaudeBot and PerplexityBot. Where this is the problem, fixing it produces the largest single improvement available, because you go from invisible to visible rather than from adequate to good.
Crawler access. Check robots.txt, then the CDN or WAF bot-management settings where the block usually lives. Same logic: binary, cheap, occasionally the whole problem.
Answer-first rewriting. Take the pages that should be answering the prompts where you are absent, and move the direct answer to the first two sentences of each section. Models lift passages; a self-contained chunk is far easier to quote than one that builds across half a page.
Specificity. Replace vague claims with dated, sourced, attributable ones. The Princeton and IIT Delhi GEO study found Statistics Addition, Quotation Addition and Cite Sources were the three highest-performing methods tested, delivering roughly 30–40 percent relative visibility improvements across around 10,000 queries.
Baseline it properly, free
Klepha runs your prompts across ChatGPT, Gemini and Perplexity and stores every answer — so the number always traces back to evidence.
Run my free scanWhat moves it in quarters
Corroboration is the highest ceiling and the longest timeline. Models weigh what independent sources say about you far more heavily than your own copy, and unlinked mentions count.
The efficient version of this work is not broad PR. It is targeted at the specific domains the assistants already cite when answering your category’s questions. Those citation lists are a ranked map of what influences your category, and being accurately described on five of those domains is worth more than fifty mentions elsewhere.
Entity clarity is slower than it looks because it requires consistency across everything, not a single page edit. One plain sentence — what category you are in, for whom, against which alternatives — stated identically on your site, in your schema, and in every third-party profile you control.
Original data is the most durable asset available. A benchmark, a survey, an analysis of your own operations. Nobody else has it, everyone writing about the topic has to cite it, and it earns both AI citations and links passively for years.
Reporting it upward
AI visibility is unusually hard to report because there is no position, no impression count, and no established benchmark. Three practices make it credible to people outside the team.
Lead with share of voice against named competitors. “We are named in 31% of category answers, our closest competitor in 58%” is immediately legible and immediately motivating. An abstract score out of 100 is neither.
Show the raw answer. One screenshot of an assistant recommending three competitors and not you does more to secure a budget than any chart. It is also honest — it is the actual product, not a derived metric.
Never report a single scan as a result. Answers vary between identical runs, and a stakeholder who learns that a number moved eight points because of sampling will discount everything else you present. Report the trend, state the variance, and say explicitly how many runs it is based on.
The one chart worth building
A single stacked line per engine: your mention rate, and the mention rate of your two closest competitors, plotted weekly across the same fixed prompt set. Three lines, three engines, one page.
It answers every question anyone will ask — are we improving, are we gaining on them, and does it differ by engine — and it is impossible to misread in the way an aggregate score is. It also makes the variance visible, which quietly teaches everyone looking at it not to over-interpret a single week.
Connecting it to revenue
The question that eventually arrives is what AI visibility is worth. There is no clean attribution model yet, and pretending otherwise damages credibility.
Three honest partial answers. Referral traffic from assistants is increasingly visible as its own channel in analytics, so some of it is directly measurable. Branded search volume tends to rise when assistants start naming you, which is measurable in Search Console. And in categories where buyers routinely arrive having already shortlisted, sales teams can usually tell you whether prospects mention having asked an assistant — anecdotal, but the earliest available signal.
When the number stops moving
Most programmes hit a plateau at some point. Three diagnoses cover nearly all of them.
You have exhausted the mechanical fixes. Retrievability and extractability are one-time gains. Once your content is readable and answer-first, doing more of that does not compound. The remaining work is corroboration, which is slower and less comfortable.
Your prompt set is too broad. If your prompts are phrased like keywords — “best CRM software” — they invite the model to name the largest players and nobody else. Check whether you appear on the constrained, realistic versions of those questions. Frequently you do, and the plateau is an artefact of testing questions you were never going to win.
The ceiling is real. On genuinely broad category questions, a small company may not be nameable yet regardless of execution. That is a positioning conclusion rather than an optimisation failure, and the response is to own the narrow questions completely and widen from there.
One thing worth checking before concluding any of the above: whether the accuracy metric has moved while the others stalled. Correcting a wrong fact an assistant repeats about you is often worth more commercially than a few points of mention rate, and it is invisible unless you were tracking it separately.
Sources
- Vercel and MERJ — analysis of 500M+ GPTBot fetches finding zero JavaScript execution; same behaviour for ClaudeBot and PerplexityBot.
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande — “GEO: Generative Engine Optimization”, ACM SIGKDD 2024 (arXiv:2311.09735); ~30–40% relative visibility improvement across ~10,000 queries.
- Google — 2.5 billion monthly active users for AI Overviews and 1 billion for AI Mode, stated 19 May 2026.
Frequently asked questions
How can I improve my AI visibility?
Define what you are measuring, baseline it honestly, then work interventions by leverage. Fast interventions are mechanical — fixing JavaScript-only content, unblocking AI crawlers, and rewriting pages to answer one question in the first two sentences. Slow interventions are corroboration: being described consistently across the third-party sources assistants already cite in your category.
What is an AI visibility score?
A composite of four things measured across a fixed prompt set: how often your brand is mentioned, how often your URLs are cited, which competitors appear alongside you, and what share of the total answer space you occupy. Scores are only comparable within one methodology, so a number from one tool cannot be compared with a number from another.
What is a good AI visibility score?
There is no universal benchmark, because scores depend entirely on the prompt set chosen. The meaningful comparisons are against your own trend over time and against named competitors on the same prompts. Any vendor quoting an industry-standard target is describing their own scale rather than a property of the market.
How many prompts do I need for a reliable baseline?
Twenty is a workable minimum for a focused business; sixty covers a broader category or several product lines. More important than count is realism and stability — prompts phrased the way buyers actually ask, and kept fixed so the trend means something. Changing the prompt set resets your history.
Why does my AI visibility fluctuate week to week?
Because generated answers vary between identical runs. That is inherent to sampling from a language model, not a fault in measurement. It is why a single scan is noise, why storing raw answers matters, and why you should read a trend across four to eight weeks rather than reacting to any one result.
Can I improve AI visibility without improving SEO?
Only to a limited extent. Assistants retrieve candidate pages before writing an answer, and retrieval leans on organic signals. You can fix retrievability and extractability independently, and those are often the biggest wins, but a site with no organic presence is not in the candidate set for a model to select from.