Is my website optimized for AI search?
Six checks with objective pass conditions, runnable in twenty minutes without buying anything. Most sites fail the first one and have never tested it.
Six checks, objective pass conditions, twenty minutes, no tooling required for five of them. Run them in order — the first two gate everything below.
Check 1: does your content exist in raw HTML?
How: disable JavaScript in your browser settings and load your five most important pages.
Pass: the answer content is fully present and readable.
Fail: you see an empty shell, a loading state, or navigation without body content.
Why it matters: AI crawlers do not execute JavaScript. An analysis by Vercel and MERJ covering more than 500 million GPTBot fetches recorded zero JavaScript execution — GPTBot downloaded JavaScript files roughly 11.5% of the time and never ran them, with the same behaviour from ClaudeBot and PerplexityBot. Googlebot renders, so this failure is invisible to every conventional SEO audit and to your rankings.
This is the most common failure on modern sites and the one nobody tests. A page can rank first on Google and show GPTBot nothing at all.
Check 2: are AI crawlers allowed?
How: read yourdomain.com/robots.txt and look for rules affecting GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot. Then check your CDN or WAF bot-management settings. Then request your homepage from outside your own network with an AI crawler user agent.
Pass: 200 response, no disallow rules on the crawlers you want.
Fail: a 403, or disallow rules nobody remembers adding.
Why it matters: many platforms ship bot-management presets that deny AI user agents by default, so the block is inherited from a security configuration rather than chosen as a policy. It does not appear in robots.txt and most audits never look.
Check 3: does each page answer one question?
How: take your ten most important pages and write down, in one sentence each, the single question it answers.
Pass: you can do it for at least seven of them without using “and”.
Fail: most pages need a list of three or four things to describe.
Why it matters: models extract passages that answer a specific prompt. A page covering six topics answers none of them cleanly enough to be preferred over a page devoted to one. This inverts the content-hub instinct and is a genuine strategic change rather than a tweak.
Check 4: do passages survive extraction?
How: copy three consecutive sentences from the middle of each key page. Paste them into a blank document and read them cold.
Pass: they still answer something without the surrounding page.
Fail: they open with “This means…”, depend on an unresolved pronoun, or reference a chart above them.
Why it matters: a model assembling an answer lifts two or three sentence chunks. If nothing on your page survives that treatment, it uses someone else’s paragraph and you never see the near-miss.
Have the assessment run for you
A free Klepha scan performs these checks automatically and adds what assistants actually say about your category.
Run my free scanCheck 5: is there anything worth attributing?
How: on each key page, find the most specific factual claim. Ask whether a model could reproduce that claim without crediting you.
Pass: at least one dated statistic, named source, or original data point per important page.
Fail: everything is “many companies find”, “research suggests”, or “can help improve”.
Why it matters: vague claims are paraphrased and lose attribution. The Princeton and IIT Delhi GEO study found Statistics Addition, Quotation Addition and Cite Sources were the three highest-performing methods tested, delivering roughly 30–40 percent relative visibility improvements.
Check 6: do assistants actually name you?
How: write ten questions a real buyer would ask before choosing a supplier in your category. Ask each to ChatGPT and Perplexity in a fresh session with no history. Record whether you are named and who is named instead.
Pass: named in at least three of ten, consistently across runs.
Fail: named in zero or one, with competitors named consistently.
Why it matters: this is the outcome the other five checks exist to produce. It is also the check that tells you whether your problem is on-site or external — if you pass checks one to five and still fail this, the gap is corroboration rather than optimisation.
When to re-run this
Quarterly as a routine, and immediately after four specific events.
Any redesign or replatform. The most common way check one starts failing, because a framework migration can move content behind JavaScript without anyone noticing that anything changed.
Any CDN or security configuration change. Bot-management presets get updated, and AI crawlers are frequently in the default deny list.
Any significant content restructure. Merging or splitting pages changes which passages exist and whether they stand alone.
Any time a competitor starts appearing where you do not. That is a signal something changed on one side or the other, and the assessment tells you which.
The failures this catches are silent by nature. Nothing errors, no alert fires, and rankings are unaffected — which is precisely why a scheduled check is the only thing that finds them.
Three optional checks
If the six core checks pass, these three separate a good site from an unusually well-prepared one.
Check 7: server response time under load. AI crawlers operate under tight timeouts, often one to five seconds. A page that takes four seconds to return its first byte is at risk of not being crawled at all, with no error anywhere to tell you. Pass: consistent sub-second time to first byte on your key pages.
Check 8: consistency of description. Read your homepage, your about page, your top three third-party listings and your Organization schema. Do they describe the same company in the same category using the same vocabulary? Pass: a stranger reading all five would produce one description. Most companies fail this and have never looked at the five side by side.
Check 9: multimodal retrievability. If meaningful information lives only in video, images or audio, is there a transcript, descriptive alt text or an on-page summary? Pass: every substantive non-text asset has a text equivalent. A ten-minute video with no transcript is invisible to a text retrieval pipeline regardless of its quality.
Running it on a large site
The manual version works for five to ten pages. On a large site, two adjustments make it tractable.
Test by template, not by page. Checks one and two are almost always template-level: if one product page renders client-side, all of them do. Pick one page per template and the coverage is effectively complete.
Sample by traffic and by intent. For checks three to five, test your ten highest-traffic informational pages rather than a random sample. Those are the pages assistants would plausibly cite, and the commercial pages that dominate a random sample are not citation candidates anyway.
That reduces a 4,000-page site to about fifteen checks, which is an afternoon rather than a project.
What this assessment deliberately does not check
It does not check page speed scores, meta description length, heading level ordering, keyword placement or text-to-HTML ratio. Those appear on most audit templates and none of them determines whether an AI engine can read, extract or trust your page.
That omission is the point. The reason most sites fail check one while passing conventional audits is that conventional audits were designed for a different problem, and adding AI-era checks to them matters more than tightening the existing ones.
Scoring it
| Result | What it means |
|---|---|
| Fail 1 or 2 | Invisible. Nothing else matters until fixed. Days of engineering work. |
| Pass 1–2, fail 3–5 | Readable but unquotable. Editorial work on ten to twenty pages. |
| Pass 1–5, fail 6 | Your site is fine; the web does not describe you. Corroboration work, quarters. |
| Pass all six | Ahead of most. Focus on original data and defending the position. |
The value of this scoring is that each outcome points at a different owner. Failing one and two is an engineering ticket. Failing three to five is an editorial project. Failing only six is a communications brief. Teams routinely apply the wrong one of those three to their actual problem.
What to do with each result
If you failed checks 1 or 2: stop everything else. Server-render the answer content and unblock the crawlers. This is days of work and it moves you from invisible to visible, which is a larger change than any editorial improvement can produce. Re-check within two weeks.
If you failed checks 3–5: pick the ten pages that should be answering your category’s most valuable questions. Split anything covering six topics, move the answer to the first two sentences, make passages self-contained, and add one concrete claim to each. Four to eight weeks before a readable trend.
If you only failed check 6: your site is not the problem. Look at which URLs assistants cite when answering your category’s questions — those domains are your target list. Being accurately described on five of them is worth more than any further on-site work.
If you passed everything: the remaining upside is original data and maintenance. Publish something you are the only source for, and re-run this assessment quarterly, because checks one and two break silently with every deploy.
Sources
- Vercel and MERJ — analysis of 500M+ GPTBot fetches finding zero JavaScript execution; GPTBot downloaded JS files ~11.5% of the time without running them; same behaviour for ClaudeBot and PerplexityBot.
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande — “GEO: Generative Engine Optimization”, ACM SIGKDD 2024 (arXiv:2311.09735); ~30–40% relative visibility improvement across ~10,000 queries.
- OpenAI, Anthropic and Perplexity documentation — crawler user agents and robots.txt directives.
Frequently asked questions
Is my website optimized for AI search?
Run six checks. Does your content appear with JavaScript disabled? Are AI crawlers unblocked at robots.txt and your CDN? Does each important page answer one question rather than six? Do passages survive being copied out of context? Is there a specific dated claim worth attributing? And do assistants actually name you for your category’s questions? Most sites fail the first check and have never tested it.
How do I test if AI can read my website?
Disable JavaScript in your browser settings and load your most important pages. Whatever text remains is what AI crawlers see, because they do not execute JavaScript. If your answer content disappears, that is a definitive failure and the highest-priority fix available.
How do I check if AI crawlers are blocked?
Read robots.txt for rules affecting GPTBot, ClaudeBot, PerplexityBot and Google-Extended. Then check your CDN or WAF bot-management settings, which is where blocks more commonly live and where they are usually inherited from a preset rather than chosen. Verify by requesting your homepage from outside your own network.
What is a good AI readiness score?
Passing checks one and two is the minimum viable state — below that, nothing else can help. Passing four of six puts you ahead of most sites. Passing all six means your remaining work is corroboration, which is external rather than on-site and takes quarters rather than weeks.
Do I need a tool to assess AI readiness?
Not for the first five checks, which are manual and take twenty minutes. Check six — whether assistants name you — is manual in principle but tedious to sustain, since answers vary between runs and a credible sample means several runs across several engines. That is where tooling earns its cost.
How often should I re-run this assessment?
Quarterly, and immediately after any redesign, replatform or CDN change. Checks one and two break silently with deploys — a template change that moves content behind JavaScript will not show up in any conventional monitoring, and the damage compounds until someone thinks to test it.