Technical SEO that lets Google and AI crawlers read your whole site
Technical problems are invisible to visitors, which is why they last for years. I find what stops Googlebot, GPTBot, ClaudeBot and PerplexityBot reading your pages, fix it or brief your developer, and prove it's fixed.
What technical SEO is, and what it does
Technical SEO is the work that lets search engines and AI systems find, crawl, render and index your pages, then understand what they are about. It covers robots.txt, XML sitemaps, status codes, redirects, canonicals, JavaScript rendering, structured data, page speed and the mobile experience. Content and links decide whether you deserve to rank. Technical SEO decides whether you are in the running at all.
Google describes how Search works in three stages: crawling, where Googlebot fetches your pages; indexing, where it works out what each page is and stores it; and serving, where it chooses results for a search. A problem at the first stage breaks everything after it. So technical SEO does four jobs:
- Access. Crawlers can reach every page you want found, over HTTPS, without errors or blocks.
- Indexing. The right pages are indexed once each, with no noindex left on by accident and no duplicate content competing with itself.
- Understanding. Structured data in JSON-LD and a clear site architecture tell search engines and AI models what each page and the business behind it are.
- Experience. Core Web Vitals and mobile-first indexing, because Google judges the mobile version of your site.
The checks I run, step by step, are in my technical SEO audit checklist.
Technical SEO now has to include AI crawlers
ChatGPT, Claude, Perplexity and Google's Gemini each have their own crawlers, and each one reads its own line in robots.txt. A site can rank well in Google and still be invisible to ChatGPT search because OAI-SearchBot is blocked, often by a CDN setting rather than a decision anyone made. Checking AI crawler access is now part of every technical audit I run.
| Crawler | Company | What it's for | If you block it |
|---|---|---|---|
| Googlebot | Google Search, including AI Overviews and AI Mode | You drop out of Google Search and its AI features | |
| Google-Extended | Whether Google may use your content for its Gemini models | No effect on inclusion or ranking in Google Search | |
| OAI-SearchBot | OpenAI | ChatGPT search results | Your site won't be shown in ChatGPT search answers |
| GPTBot | OpenAI | Training OpenAI's models | Your content isn't used for training |
| ClaudeBot / Claude-SearchBot | Anthropic | Training / search quality | Excluded from training / possibly less visible in Claude's search results |
| PerplexityBot | Perplexity | Perplexity search results (not training) | Perplexity says to allow it if you want to appear in its results |
| ChatGPT-User, Claude-User, Perplexity-User | All three | Fetching a page a person asked about | OpenAI says robots.txt may not apply; Perplexity says its user agent generally ignores it |
The decision splits in two. Search-type crawlers are how you get found and cited, so blocking them by accident is expensive. Training crawlers are a business choice: some owners want their content in future models, others don't. I help you make that choice deliberately, crawler by crawler, rather than inheriting whatever a default decided.
It matters in New Zealand: AI Overviews appeared on 74% of 1,000 NZ commercial searches in my NZ AI Overviews study. For Google itself, there is nothing extra to build. Google says its AI features need no special files or markup, just a page that is indexed and eligible to show with a snippet, with crawling allowed in robots.txt and by any CDN or hosting infrastructure.
The Cloudflare setting that quietly blocks AI crawlers
If your site runs through Cloudflare, check your live robots.txt, not the file on your server. Cloudflare's managed robots.txt adds its own block above yours that disallows GPTBot, ClaudeBot, Google-Extended and other AI crawlers. Your file can say "allow" while the live version says "disallow", and nothing on the page looks any different to a visitor or to you.
I found this on my own site. Cloudflare's managed robots.txt setting was placing Disallow rules for GPTBot, ClaudeBot, Google-Extended and other AI crawlers above my own robots.txt, on a site whose whole purpose is to be read by AI. Turning off the individual AI crawler toggles didn't remove it; the separate "Managed robots.txt" switch did. The lesson I took: fetch the live robots.txt over HTTP rather than trusting the file on disk, and that is now a check I run on every site.
The injected section starts with a "# BEGIN Cloudflare Managed content" line. Because Cloudflare adds it at the edge, editing your own file does nothing to it. My robots.txt, with comments explaining all of this, is public if you want to see how it is set up now.
Robots.txt is a request, not a lock: Cloudflare's own documentation says compliance is voluntary. Cloudflare's AI Crawl Control and AI bot policies are different, because they can block crawlers at the edge, and Cloudflare has announced new defaults for new domains from 15 September 2026. So check your settings in both places rather than assuming. If a firewall rule turns out to be the blocker, Perplexity publishes Cloudflare WAF rules for allowing its crawlers.
What's included in technical SEO work
Every engagement starts with a crawl of the whole site, your Google Search Console data and a live check of what each major crawler is allowed to fetch. From that I write a ranked fix list: blockers first, then indexing waste, then speed and structured data. I implement the fixes myself where I have access, or brief your developer precisely.
Crawl access
Robots.txt, CDN and firewall rules, and server errors, checked for Googlebot, Bingbot and the AI crawlers, against the live site rather than the files.
Indexing
The Search Console Page indexing report, stray noindex tags, canonicals, duplicate URLs, soft 404s and an XML sitemap that lists only pages you want indexed.
Redirects and status codes
Redirect chains and loops, 404 pages that internal links still point at, and a clean move from HTTP to HTTPS with one hop.
Rendering
What Googlebot sees after JavaScript runs compared with what is in the raw HTML, because not every crawler runs JavaScript.
Structured data
One clean JSON-LD graph using schema.org types for the business, services, articles and breadcrumbs, matched to visible content. It is also the base of my structured data and entity work for AI search.
Core Web Vitals and mobile
LCP, INP and CLS from real-visitor field data, fixed at template level, and the same content and links on mobile as on desktop.
Site architecture and internal links
Orphan pages, click depth, breadcrumbs, and hreflang where you sell in more than one country. For online stores: filter URLs, duplicate products and crawl waste.
AI-readiness extras
Content signals in robots.txt, server log checks for AI crawler visits, and llms.txt, with the honest caveat that it is optional and unproven.
Ecommerce technical SEO gets the same process with more URLs: faceted navigation, product variants and out-of-stock pages create duplicates quickly. Crawl budget is rarely a problem on small sites, but on a large store it is worth checking in the Crawl Stats report.
How I work
You work with me directly, from the first crawl to the last re-check. I've been working in search since 2006, across corporate, local and global campaigns, so I know which technical issues move rankings and which are noise in an audit tool's red list. Work is remote, for businesses NZ-wide and overseas, in the UK, Canada and the US.
Free AI Visibility Audit
Where you rank today, whether AI assistants mention you, and a first look at whether any crawler is being turned away.
Technical crawl and ranked fix list
A full crawl, Search Console, the live robots.txt and your CDN settings, turned into a list ordered by impact.
Fixes, done or briefed
I make the changes myself where I have access, or write a precise brief for your developer, as a technical SEO specialist and strategist rather than a report generator.
Re-check
Search Console, the live robots.txt and, where available, crawler logs, to prove each fix worked for Google, Bing and the AI crawlers.
What technical SEO costs
Technical SEO is part of my SEO retainer and of my one-off SEO and AI search audit, both with published prices on how much SEO costs in New Zealand, and anything larger is quoted after the free audit rather than guessed.
Technical SEO: questions people ask
What is the difference between SEO and technical SEO?
Technical SEO is one part of SEO: access, indexing, rendering, speed and markup. The rest is content that answers real questions, and the authority from links and mentions that makes Google and AI models trust you. You need all three. Technical problems put a ceiling on what the other two can achieve, however good they are.
Can you give me an example of technical SEO?
A common one: a site goes live with the noindex tag from its staging copy still on every page, so Google drops it from the results while the site looks perfectly normal to visitors. The fix is removing the tag, checking the live pages with the URL Inspection tool and asking Google to recrawl. Finding it early is the technical SEO part.
What does a technical SEO consultant do?
A technical SEO consultant works out why pages aren't being crawled, indexed or understood, then ranks the fixes by impact and either makes them or briefs a developer. That means reading Search Console, crawling the site, checking robots.txt and CDN rules, testing rendering and structured data, and re-checking after each change until the problem is gone.
Can I do technical SEO myself?
Yes, for a small site on a mainstream platform. Google Search Console, a free PageSpeed Insights test and my audit checklist will catch most problems. Get help when the site relies on JavaScript to show content, when you are migrating to a new platform or domain, when you sell in several countries, or when a CDN or firewall sits in front of the site.
Does blocking GPTBot or Google-Extended hurt my Google rankings?
No. Google says Google-Extended does not affect a site's inclusion or ranking in Google Search, and GPTBot is OpenAI's training crawler, which has nothing to do with Google. Blocking OAI-SearchBot is a different matter: OpenAI says sites opted out of it won't be shown in ChatGPT search answers. Decide on each crawler separately.
Do I need an llms.txt file?
No. It's a 2024 proposal, not a standard, and it is unproven. Google says you don't need AI text files to appear in its AI features, and the crawler documentation from OpenAI, Anthropic and Perplexity doesn't mention it. I publish one on this site because it costs almost nothing to maintain, but I don't sell it as a ranking fix.
Find out what's blocking Google and AI crawlers from your site
The free AI Visibility Audit shows where you rank in Google today, whether ChatGPT, Google AI Overviews and Perplexity mention your business, and whether anything is stopping their crawlers reading your pages. Free, no obligation.
Book Your Free AuditSources
- Google Search Central: In-depth guide to how Google Search works
- Google Search Central: AI features and your website
- Google: Google's common crawlers (Google-Extended)
- Google Search Central: Introduction to robots.txt
- Google Search Central: Understanding Core Web Vitals and Google search results
- Google Search Central: JavaScript SEO basics
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Cloudflare: Managed robots.txt and AI bot policies
- llms.txt: the llms.txt proposal