AI crawlers, end to end
Every AI assistant that can talk about your business got there one of two ways. Either your pages were part of what the model learned from, or something fetched your page while answering a question. Those are different systems, run by different crawlers, controlled by different rules, and worth different amounts to you. This guide covers which crawlers exist, how to decide about each, how to write the rules so they take effect, and how to confirm any of it worked.
TL;DR
- AI crawlers split into two jobs: collecting text to train a model, and fetching pages to answer a question now.
- Blocking the training crawler and blocking the retrieval crawler have opposite consequences, and several vendors run one of each.
- A named group in robots.txt replaces the wildcard group rather than adding to it, so naming a bot can silently unblock everything you had denied.
- Google-Extended is not a user agent you can see in your logs. It is a control token only.
- Nothing here needs a tool. All of it is checkable against vendor documentation and your own server logs.
Short answer
Controlling AI crawlers means deciding separately about training and retrieval, then writing robots.txt so the rules reach the bots you named. Most sites get the second part wrong: adding a group for a specific crawler removes the wildcard rules from that crawler entirely, so directives you thought applied everywhere quietly stop applying to the one bot you were trying to control.
What this guide covers
This is written for the person who has to make the change: someone with access to robots.txt and to server logs, deciding what an AI company may do with a site they are responsible for. It assumes no tooling, because none is needed. Every fact in it comes from a crawler operator documenting its own behavior, and every check in it can be run against your own server.
It deliberately avoids the numbers that circulate in this subject. Citation-lift percentages and visibility-uplift figures are repeated widely and traced to nothing, and a guide about verifying claims should not open by repeating unverifiable ones.
-
01
Which AI crawlers exist, and what each one wants
The AI crawlers that matter divide into training crawlers, which collect text that may be used to build future models, and retrieval crawlers, which fetch pages to answer a live question and usually cite what they fetched. OpenAI runs GPTBot and OAI-SearchBot, Anthropic runs ClaudeBot and Claude-SearchBot, Perplexity runs PerplexityBot for retrieval only, and Google uses the Google-Extended token to cover Gemini.
-
02
Deciding what to allow and what to block
Decide the training question and the retrieval question separately, because their consequences are opposite. Blocking a retrieval crawler removes you from answers that would have named and linked you, which is a direct traffic cost. Blocking a training crawler costs no traffic and gains none, so it turns on whether the content could be sold to an AI company or must be withheld for another reason.
-
03
Writing robots.txt rules that actually apply
A crawler uses the robots.txt group that matches its own product token, and only falls back to the wildcard group when no matching group exists. That means adding a group for a named crawler removes every wildcard rule from that crawler. Anything you want applied to it has to be repeated inside its own group, which is safe because groups sharing a token are merged.
-
04
Checking it worked
Check crawler access in your own server logs, which record what was actually requested, and verify identity by IP range rather than trusting the user agent string. Referral traffic from assistants shows up in analytics as ordinary referrers. Neither method tells you whether you were cited in an answer, and Google-Extended cannot be verified at all because it has no user agent of its own.
Common questions
Will blocking AI crawlers hurt my search rankings?
Blocking the AI-specific tokens does not affect Googlebot, so ordinary Google rankings are unaffected. The exception worth knowing is that Google AI Overviews are drawn from the normal Search index rather than a separate crawl, so there is no AI-specific token that removes you from them while leaving your rankings intact. Opting out of Overviews means opting out of the snippet, which is a different and larger decision.
Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “our generative AI features on Google Search are rooted in our core Search ranking and quality systems”
If I block a crawler, does the model forget what it already learned?
No. A robots.txt rule affects future fetches only. Anything already collected has already been collected, and no crawler directive reaches back into a model that has been trained. This is the main reason the decision is worth making deliberately rather than later.
Do I need an llms.txt file for any of this?
No. Google states that no AI system currently uses llms.txt, and crawler access is governed by robots.txt, which every one of these bots does read. If you want the reasoning and the evidence in full, it is set out on the page about what I will not sell you.
Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them.”
Worth saying plainly
This guide describes how crawler control works, not results it has produced for clients. Whether allowing a retrieval crawler leads to being cited depends on things no site controls, and anyone quoting you a figure for it is guessing. What is claimed here is narrower and checkable: these are the crawlers, this is what their operators say each one does, and this is how to make a rule apply.