Technical SEO audit checklist: what to check for Google and AI crawlers
This is the order I run a technical SEO audit in: seven steps, each with the free tool to use and the official source behind the check. The AI crawler step is the new one. A site can be fine for Google and blocked for ChatGPT search without anyone touching the site, because a CDN or security setting changed underneath it. If you'd rather I ran it for you, that's my technical SEO service.
What a technical SEO audit is (and what it isn't)
A technical SEO audit checks whether search engines and AI crawlers can reach, render and index your pages, and whether they understand them once they do. It isn't a content review or a backlink audit, and it isn't the score from an automated checker. Tools list issues. The audit decides which of them actually cost you visibility, and in what order to fix them.
If you're wondering how this differs from SEO in general: technical SEO is the foundation layer, and content and authority sit on top of it. The difference is explained in full on the service page. If a term in this checklist is new to you, my glossary of technical SEO terms defines it in plain English.
Before you start, you need four things: access to Google Search Console for the site, the live site itself (not a staging copy), a desktop site crawler, and some uninterrupted time. For a small business site, one to three hours is realistic. A large online store takes longer, mostly because there are more URL patterns to understand.
Step 1: Can Google and AI crawlers reach the site?
Open your live robots.txt in a browser, at yoursite/robots.txt, and read every line. Then check that nothing in front of the site, such as a CDN, firewall or bot protection, is blocking crawlers the file allows. Robots.txt is a request, not a lock, and a firewall rule can turn away a crawler your file welcomes. Both layers need to agree.
- Live file vs the file on your server. Cloudflare's managed robots.txt prepends its own section, starting "# BEGIN Cloudflare Managed content", with Disallow rules for GPTBot, ClaudeBot, Google-Extended and other AI crawlers. Editing your file won't remove it. I explain how I found this on my own site on the service page.
- Search-type crawlers allowed. Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot and PerplexityBot should not be blocked in the file, at the CDN or by the firewall, by user agent or by IP range.
- Training crawlers set deliberately. GPTBot, ClaudeBot and Google-Extended are about model training. Allowing or blocking them is your call; what matters is that someone chose.
- No blocked CSS or JavaScript. If robots.txt blocks files a page needs to render, Google sees a broken page.
- Server errors and timeouts. Search Console's Crawl Stats report shows 5xx errors, host problems and how fast your server responds to Googlebot.
- What Google fetched. Search Console's robots.txt report shows the version Google last fetched and any problems it found.
Robots.txt does not keep a page out of Google. A blocked URL can still be indexed if other pages link to it. Use a noindex tag instead, and don't block that page in robots.txt, or Google never sees the noindex.
Step 2: Are the right pages indexed, and only those?
Open Search Console's Page indexing report and compare the indexed count with the number of pages you actually want in Google. Too few means something is blocked, noindexed or judged a duplicate. Too many means parameter URLs, tag pages or staging copies are being indexed. Read the reasons Google gives for each group, then fix the biggest group first.
- Stray noindex. Check both the robots meta tag and the X-Robots-Tag HTTP header. A noindex left over from staging is one of the most common launch mistakes.
- Canonicals. Each canonical tag should point at the preferred URL. Google treats redirects and rel=canonical as strong signals and sitemap inclusion as a weak one, so a sitemap alone won't fix duplicate content.
- One version of every URL. HTTP and HTTPS, www and non-www, with and without a trailing slash: all but one should redirect.
- XML sitemap. It should list only indexable pages that return a 200 status. Google says a small, well-linked site of about 500 pages or fewer may not need one, but it rarely hurts.
- Soft 404s, chains and broken links. Look for "not found" pages that return 200, redirect chains and loops, and internal links that lead to a 404.
For the pages that matter most, run the URL Inspection tool. It shows whether Google indexed the page, which canonical it chose, and when it last crawled it.
Step 3: Does the content exist before JavaScript runs?
Compare the raw HTML of a key page with what Google renders. If your headings, body text, links or structured data only appear after JavaScript runs, you're relying on every crawler executing it. Google renders JavaScript, but Google itself says not all bots can run JavaScript, and recommends server-side rendering or pre-rendering. Put the content that matters in the HTML.
- View Source vs the rendered page. Compare your browser's View Source with the URL Inspection tool's "View crawled page" in Search Console.
- Real links. Navigation and internal links should be real <a href> elements, not click handlers on other elements.
- Lazy-loaded content. Anything loaded as the visitor scrolls should still appear in the rendered HTML.
- Meaningful status codes. A missing page should return a 404, not a 200 with a "not found" message drawn by JavaScript, which is common on single-page apps.
Step 4: Is the site structure easy to follow?
Every important page should be reachable through normal links within a few clicks of the home page, with descriptive anchor text. Pages with no internal links pointing at them, called orphan pages, are found late, crawled rarely and look unimportant. Breadcrumbs, clear navigation and links between related pages fix most structure problems without a redesign.
- Orphan pages. Compare your crawler's list with the sitemap and your analytics. Pages that only exist in the sitemap need internal links.
- Click depth. Pages you want to rank should not be buried five or six clicks deep.
- One page per topic. Two pages targeting the same search compete with each other. Merge them or give each a distinct job.
- Breadcrumbs. They help people and crawlers see where a page sits.
- Hreflang. Only needed if you publish separate pages for different countries or languages.
Step 5: Does your structured data describe the business accurately?
Structured data tells search engines and AI systems what a page is and who is behind it. Use JSON-LD, describe only what is visible on the page, and define your business once, then reference it from every other page. Validate it, but remember Google says no special schema is needed to appear in AI Overviews. Accuracy matters more than volume.
- JSON-LD. Google recommends JSON-LD where your setup allows it.
- One entity graph. Give the business, the person and each service a stable @id and reference them, rather than redefining the business on every page.
- Markup matches the page. No FAQs, reviews or prices in the schema that visitors can't see.
- The right type. Use schema.org's ProfessionalService or LocalBusiness only when it fits the business.
- Validate. Use Google's Rich Results Test for Google features and the Schema Markup Validator for schema.org in general.
- FAQ markup is for context now. Google shows FAQ rich results only for well-known, authoritative government and health sites, so FAQ markup no longer earns stars in the results for a business.
Structured data also plays a part beyond Google, and I explain how structured data supports AI search visibility on my AI search page. Local business schema is covered on my local SEO page.
Step 6: Do your Core Web Vitals pass on mobile?
Check the Core Web Vitals report in Search Console, which uses data from real visitors. Google's "good" thresholds are Largest Contentful Paint (LCP) within 2.5 seconds, Interaction to Next Paint (INP) under 200 milliseconds and Cumulative Layout Shift (CLS) under 0.1. Fix the templates that fail rather than individual pages, and remember Google indexes the mobile version of your site first.
- Field data first. Search Console and PageSpeed Insights field data reflect real visitors. Lab scores are for diagnosing.
- The largest element. The hero image or heading font is usually what decides LCP.
- Layout shifts. Late-loading banners, embeds and images without dimensions push content around.
- Mobile parity. Under mobile-first indexing, the mobile page needs the same content, links and structured data as desktop.
- HTTPS everywhere. No mixed content warnings on any page.
Be realistic about it: good Core Web Vitals help, but they won't outrank a more useful page on their own. Google says they align with what its ranking systems seek to reward, and no more than that.
Step 7: AI-readiness checks beyond the basics
Once Google can crawl and index you, check three things for AI assistants: whether their crawlers are actually visiting, whether your robots.txt states how your content may be used, and whether your key facts are in plain text on the page. An llms.txt file is optional and unproven: fine to add, not a fix for anything.
- Are AI crawlers visiting? Look for AI crawler hits in your server logs or Cloudflare's AI Crawl Control. User agents can be faked, so verify against the IP lists OpenAI and Perplexity publish.
- Content signals. If you want to state preferences, content signals in robots.txt (search, ai-input, ai-train) say how your content may be used. Like robots.txt itself, they are preferences, not enforcement.
- Facts in text. Services, pricing policy, locations and contact details should be written on the page, not only in images or PDFs.
- llms.txt. A 2024 proposal, unproven. Google says you don't need AI text files for its AI features, and none of the major AI crawler documentation mentions it. I publish one on this site because it's cheap to maintain.
This step matters because AI answers are now part of most commercial searches here: in my NZ AI Overviews study, 74% of 1,000 NZ commercial searches showed an AI Overview, and those answers link to pages Google could crawl and index.
How to prioritise what the audit finds
Fix blockers first: anything that stops important pages being crawled or indexed. Then fix waste, such as duplicates, redirect chains and junk URLs in the index. Then improve experience and markup: Core Web Vitals and structured data. A 200-item report from a tool usually contains a handful of fixes that matter, and the rest can wait or be ignored.
| Priority | Examples | Why |
|---|---|---|
| 1. Blockers | Crawlers blocked in robots.txt or at the CDN, noindex on live pages, server errors, broken redirects on key pages | Nothing else counts until the pages can be crawled and indexed |
| 2. Waste | Duplicate URLs, redirect chains, soft 404s, parameter and tag pages in the index, orphan pages | Crawlers spend time on the wrong URLs and signals split across copies |
| 3. Experience and markup | Failing Core Web Vitals templates, missing or inaccurate structured data, weak internal links | Improves how pages are understood and experienced once they are indexed |
Technical SEO audit tools: what you actually need
For most small and mid-sized sites, three free or low-cost tools cover it: Google Search Console for indexing and Core Web Vitals, PageSpeed Insights for page-level speed, and a desktop site crawler for status codes, redirects and internal links. Automated checkers are useful for spotting issues. They can't tell you which ones matter for your business.
- Google Search Console: the Page indexing report, Crawl Stats report, robots.txt report, Core Web Vitals report and URL Inspection tool.
- PageSpeed Insights: field and lab data for a single URL.
- Rich Results Test and Schema Markup Validator: for structured data.
- Bing Webmaster Tools: Bing's equivalent of Search Console, worth connecting too.
- A desktop crawler: to list every URL, its status code, canonical and internal links.
- An all-in-one SEO suite: optional, useful for tracking issues over time on larger sites.
Platforms like Shopify and WordPress handle some of this by default, such as sitemaps and canonical tags. Check anyway: themes, apps and plugins change the defaults.
The technical SEO audit checklist
Here is the whole audit on one page, grouped by the seven steps. Work down it in order, note anything that fails, then sort the failures using the priority table above. Every threshold in it comes from Google's own documentation, and each "pass" describes what a healthy site looks like, not a perfect score in a tool.
| Check | Where to look | Pass looks like |
|---|---|---|
| 1. Crawl access | ||
| Live robots.txt matches your file | yoursite/robots.txt in a browser | No unexpected Disallow rules, no managed block you didn't choose |
| Search crawlers allowed | robots.txt, CDN and firewall rules | Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot and PerplexityBot not blocked |
| Training crawlers set on purpose | robots.txt, CDN AI settings | GPTBot, ClaudeBot and Google-Extended allowed or blocked by decision |
| CSS and JavaScript crawlable | robots.txt, URL Inspection | No resources the page needs are blocked |
| Server health | Search Console › Crawl Stats | No sustained 5xx errors or host problems |
| Google's copy of robots.txt | Search Console › robots.txt report | Fetched successfully, no errors |
| 2. Indexing | ||
| Indexed pages = pages you want | Search Console › Page indexing | No important URL excluded without a reason you chose |
| No stray noindex | Page source, HTTP headers | Noindex only on pages you want out of Google |
| Canonicals | Crawler, URL Inspection | Each page's canonical points at the preferred URL |
| One version of each URL | Type the variants into a browser | HTTP, non-preferred host and slash variants redirect once |
| XML sitemap | yoursite/sitemap.xml, Search Console › Sitemaps | Only indexable URLs that return 200 |
| Redirects and 404s | Crawler | No chains, loops or internal links to 404 pages |
| 3. JavaScript and rendering | ||
| Content in the HTML | View Source vs URL Inspection | Headings, text, links and structured data present before JavaScript runs |
| Real links | Page source | Navigation uses <a href> elements |
| Honest status codes | Crawler, a made-up URL | Missing pages return 404 |
| 4. Site structure | ||
| No orphan pages | Crawler vs sitemap vs analytics | Every page you want found has internal links |
| Click depth | Crawler | Important pages a few clicks from the home page |
| One page per topic | Your page list and target searches | No two pages chasing the same search |
| Hreflang (if needed) | Page source, crawler | Country or language versions reference each other |
| 5. Structured data | ||
| Valid JSON-LD | Rich Results Test, Schema Markup Validator | Parses with no errors |
| Business defined once | Page source across the site | One definition, referenced by @id elsewhere |
| Markup matches the page | Page vs its schema | Nothing in the markup that visitors can't see |
| 6. Core Web Vitals and mobile | ||
| LCP / INP / CLS | Search Console › Core Web Vitals | Within 2.5 s / under 200 ms / under 0.1 |
| Mobile parity | Phone vs desktop | Same content, links and structured data |
| HTTPS | Browser, crawler | Every page on HTTPS, no mixed content |
| 7. AI-readiness | ||
| AI crawler visits | Server logs or CDN bot analytics | Visits from the search crawlers you allow, verified by IP |
| Key facts in text | Your service, pricing and contact pages | Services, areas and contact details readable as text |
| Content signals and llms.txt | robots.txt, yoursite/llms.txt | Present only if you chose them; optional |
Technical SEO audit: questions people ask
Can ChatGPT do an SEO audit?
Partly. It can explain issues and review a robots.txt file or a page's HTML that you paste in. It can't see your Search Console data, crawl your whole site or measure real-visitor Core Web Vitals, and it can be confidently wrong about your setup. Use it as an assistant alongside the data from the tools above, not as the audit.
How much does a technical SEO audit cost?
It depends on the size of the site and how complex it is, and the checklist on this page is free to run yourself. If you'd rather pay for one, my SEO and AI search audit covers technical SEO at a fixed published price, and typical New Zealand audit prices are on SEO costs in New Zealand.
How often should I run a technical SEO audit?
Run a full audit once a year, and again after any redesign, migration, platform change or CDN change, because those are when things break. In between, look at Search Console's Page indexing and Core Web Vitals reports once a month, and fetch your live robots.txt now and then. Most problems show up there first.
Is a free technical SEO checker enough?
A free checker is good for finding issues, not for judging them. It can't see your Search Console data or your CDN and firewall settings, so it misses some of the problems that matter most, and it flags plenty that don't. Use one to build the list, then work through this checklist to decide what's worth fixing.
Can you give me an example of a technical SEO problem?
A CDN adding AI crawler blocks is a good one. Cloudflare's managed robots.txt can place Disallow rules for GPTBot, ClaudeBot and Google-Extended above your own file, so the file on your server looks fine while the live one blocks them. Nobody notices, because nothing changes for visitors. More on this is on my technical SEO service page.
Sources
- Google Search Central: Introduction to robots.txt and Block Search indexing with noindex
- Google Search Central: AI features and your website
- Google Search Central: How to specify a canonical URL and Learn about sitemaps
- Google Search Central: JavaScript SEO basics
- Google Search Central: Introduction to structured data markup, structured data policies and FAQPage
- Google Search Central: Understanding Core Web Vitals and mobile-first indexing
- Search Console Help: Page indexing report, Crawl Stats report, robots.txt report and URL Inspection tool
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Anthropic's crawlers and how to block them
- Perplexity: Perplexity crawlers
- Cloudflare: Managed robots.txt
- The llms.txt proposal and Content Signals
Want me to run this on your site?
Start with the free AI Visibility Audit. It shows where you rank in Google today, whether ChatGPT, Google AI Overviews and Perplexity mention you, and whether anything is stopping their crawlers reading your pages.
Book Your Free Audit