How do I perform an SEO audit?
Most SEO audits produce a document nobody acts on, because they list everything wrong instead of ranking what matters. This one is built to end in a prioritised fix list.
Most SEO audits fail as documents. They arrive as a two-hundred-row spreadsheet colour-coded by severity, everyone agrees it is thorough, and eleven months later four items have been fixed.
The problem is not effort. It is that a catalogue is not a diagnosis. An audit should end with a short ranked list of things that are actually suppressing performance, in the order they should be fixed — and to produce that, you have to work in layers where each one gates the next.
Key takeaways
- Six layers, in order: access, indexation, structure, content, authority, AI retrievability.
- Stop and fix rather than continuing to catalogue. A blocked layer makes everything below it unmeasurable.
- Severity beats completeness. Three fixed items beat two hundred logged ones.
- Layer six is missing from almost every audit template and is now frequently the largest single gap.
- The deliverable is a ranked fix list, not a report.
The organising principle
Search is sequential. A page must be reachable before it can be indexed, indexed before it can rank, and ranked before content quality matters to anyone. That sequence gives you both the audit order and the severity ranking for free.
It also tells you when to stop. If layer one is broken, findings at layer four are noise — you cannot assess whether content is underperforming when the pages are not being served properly. Fix the blocking layer, wait for a recrawl, then continue.
Layer 1: access
The question: can search engines and AI crawlers reach the pages at all?
The checks. Fetch ten representative URLs as plain HTTP requests and confirm 200 responses. Read robots.txt line by line and confirm nothing important is disallowed — staging rules that survived launch are the classic finding. Check your CDN or WAF bot-management settings, which is where crawler blocks usually live rather than in robots.txt. Confirm your XML sitemap exists, returns 200, and contains only canonical, indexable URLs. Follow the redirect chains on your top pages and confirm none exceeds two hops.
Severity: catastrophic. Anything found here outranks everything else in the report.
Layer 2: indexation
The question: of the pages engines can reach, which have they chosen to store?
The checks. Open the Pages report in Search Console and read the excluded reasons rather than the summary count. Compare indexed pages against the number you believe you have — a large gap is the finding. Then interpret the specific exclusions: “Crawled — currently not indexed” is a quality or duplication judgement, not a fault, and resubmitting will not change it. “Duplicate without user-selected canonical” means several pages compete to answer the same thing. “Discovered — currently not indexed” usually signals weak internal linking.
A large crawled-not-indexed bucket is not a technical problem. It is Google telling you the site contains more pages than it contains value.
Severity: high. The fix is usually consolidation and deletion, which teams resist and which works.
Layer 3: structure
The question: does the architecture help engines and readers understand what the site is about?
The checks. Confirm every important page is reachable within three clicks of the homepage. Find orphaned pages — live URLs with no internal links pointing at them — using a crawler compared against your sitemap. Check that internal anchor text is descriptive rather than “click here”. Confirm canonical tags point where you intend. Look for topic clusters: pages on the same subject should link to each other, and if your best related pages never reference each other, that is free authority being left unrouted.
Severity: medium, but unusually cheap to fix. Internal linking is entirely within your control and consistently underused.
Layer 4: content
The question: do the pages answer the queries they target, better than what currently ranks?
The checks. Export Search Console queries by page. For each significant page, compare its actual queries against what you intended it to target — the gap between those two is often the finding. Check for cannibalisation with site:yourdomain.com [keyword]. For pages ranking 8–20, list the H2s of the top three results and identify the subtopics you are missing. Check whether each page answers its core question in the first two sentences or buries it under preamble.
Severity: medium to high, and this is where most of the sustainable upside lives once layers one and two are clean.
Audit the layer your template is missing
A free Klepha scan checks AI retrievability and shows what ChatGPT, Gemini and Perplexity say about your category. Two minutes, no card.
Run my free scanLayer 5: authority
The question: do other credible sources vouch for you on this topic?
The checks. Compare your referring domain count and quality against the sites ranking above you for your priority terms. Look at the anchor text distribution — heavily commercial anchors from low-quality sources is a risk signal. Check for lost links on pages that used to rank. Look for unlinked brand mentions, which are both a link-building opportunity and, increasingly, a direct AI visibility signal in their own right.
Severity: high for competitive commercial terms, low for long-tail. Judge it per term rather than site-wide, because a site can be perfectly authoritative for its niche and hopeless against a national publisher.
Layer 6: AI retrievability
The question: can AI crawlers read your answers, and do assistants name you?
This layer is absent from virtually every audit template in circulation, because the templates predate it. It is now frequently the largest single gap on an otherwise healthy site.
The checks. Disable JavaScript in your browser and load your ten most important pages. If the answer text is not there, AI crawlers do not see it either — an analysis by Vercel and MERJ covering more than 500 million GPTBot fetches recorded zero JavaScript execution, with the same behaviour for ClaudeBot and PerplexityBot. Then confirm those crawlers are not blocked in robots.txt or at the CDN. Then check whether your structured data matches the visible text, since a mismatch is worse than no markup. Finally, run ten of your category’s real buyer questions through ChatGPT and Perplexity and record whether you are named and who is named instead.
Severity: potentially catastrophic and completely invisible to every other layer. A page can pass layers one through five and be entirely absent from the surface where your buyer’s shortlist is formed. Our AI-readiness assessment covers this layer in full.
What the output should look like
Not a spreadsheet of everything. A single page with three sections:
Blocking. Things preventing crawling, indexation or rendering. Usually zero to three items. These get fixed this week, regardless of what else is scheduled.
Suppressing. Things holding back pages that otherwise work — cannibalisation, missing subtopics, poor titles on high-impression pages. Usually five to fifteen items, ranked by the traffic at stake rather than by how wrong they are.
Logged. Everything else. Real findings, genuinely low priority, recorded so nobody re-discovers them next quarter and treats them as urgent.
Each item needs three fields to be actionable: what is wrong, what it costs, and what specifically to change. A finding without an estimated cost cannot be prioritised, and a finding without a specific change is a complaint rather than a task.
How often to run one
A full six-layer audit annually. A layer one and two check quarterly, because access and indexation break silently and cheaply. An immediate audit after any migration, replatform, CMS upgrade or unexplained traffic change — those are the moments when layer one breaks, and the cost of finding out three months late is enormous.
Between audits, monitor rather than re-audit. Search Console coverage alerts, a monthly indexed-page count, and a periodic AI visibility scan will surface anything that matters faster than a scheduled review would.
Five findings that are usually misdiagnosed
Certain audit findings get the wrong treatment reliably enough to be worth naming.
“Crawled — currently not indexed” treated as a technical bug. It is not. Google reached the page and declined to store it, which is a quality or duplication judgement. Resubmitting, rebuilding the sitemap and requesting indexing will all fail. The fix is consolidation, or making the page substantially better than what already exists.
Duplicate content treated as a penalty risk. There is no duplicate content penalty. What actually happens is signal splitting: several pages compete for one query and none accumulates enough to rank. The consequence is real, the mechanism is not punitive, and the fix is merging rather than rewriting.
Slow pages treated as the main ranking problem. Speed is a tiebreaker. A site at position fifteen is not there because of load time. Fix speed for conversion, for user experience, and because AI crawlers time out — not in the expectation of a ranking jump.
Thin content judged by word count. A 400-word page that answers its question completely is not thin. A 3,000-word page that restates the first page of results is. The test is whether the page contains anything the current top result does not, and length is a poor proxy for it.
Missing schema treated as urgent. Structured data helps engines understand and attribute a page, and it matters more for AI citation than it used to. But adding schema to a page that is not indexed changes nothing, and schema that contradicts the visible text is worse than none at all. It belongs after layers one and two, not before them.
What a good audit feels like
A finished audit should be uncomfortable rather than reassuring. If it lists two hundred issues and none of them explains why traffic is where it is, it has catalogued rather than diagnosed. If it names three things and one of them makes someone say “oh”, it has worked.
The other marker is that it should shorten the roadmap rather than lengthen it. A good audit frequently removes work — pages to delete rather than improve, findings to log rather than fix, terms to abandon rather than pursue. An audit that only ever adds to the backlog has not prioritised anything.
Sources
- Vercel and MERJ — analysis of 500M+ GPTBot fetches finding zero JavaScript execution; same behaviour observed for ClaudeBot and PerplexityBot.
- Google Search Central — documentation on the Pages report, index coverage states and canonicalisation.
- SERP feature tracking, 2026 — AI Overview presence averaging ~32.5% of monitored queries within a 31.02%–34.40% band.
Frequently asked questions
How do I perform an SEO audit?
Work six layers in order: access, whether engines can reach the site; indexation, whether pages are actually stored; structure, whether the architecture and internal linking make sense; content, whether pages answer their queries; authority, whether other sites vouch for you; and AI retrievability, whether AI crawlers can read your answers without JavaScript. Each layer gates the next, so you stop and fix rather than continuing to catalogue.
How long does an SEO audit take?
A focused audit on a small site takes half a day. A thorough audit on a large or complex site takes one to two weeks, most of which is interpretation rather than data collection. If it is taking longer than that, the audit has usually turned into a catalogue instead of a diagnosis.
What tools do I need for an SEO audit?
Google Search Console for indexation and query data, a crawler such as Screaming Frog for site-level technical issues, and a browser with JavaScript disabled for the AI retrievability check. That covers the majority of findings on most sites. Paid suites add competitor and backlink context, which matters at the authority layer.
How often should I audit my site?
A full audit annually, a lighter technical check quarterly, and an immediate audit after any migration, replatform or unexplained traffic change. Auditing more often than that tends to generate findings faster than anyone can fix them, which is its own failure mode.
What is the most commonly missed item in an SEO audit?
AI retrievability. Nearly every audit template predates AI crawlers, and none of them check whether the page content exists in raw HTML before JavaScript runs. Since AI crawlers do not execute JavaScript, a client-side-rendered page can pass a conventional audit completely and still be invisible to ChatGPT.
Should I fix everything an audit finds?
No. Most audit findings are cosmetic and fixing them consumes capacity that should go to the two or three items actually suppressing performance. Rank findings by severity — does this block crawling, indexation, or a specific ranking — and fix in that order. A finding that cannot be tied to an outcome should be logged, not scheduled.