Which AI crawlers exist, and what each one wants
The single most useful distinction in this subject is between a crawler that is collecting text for training and a crawler that is fetching your page because someone just asked a question. They look identical in a log file until you know the names, and they are worth opposite things to most businesses.
Part of the guide to AI crawlers, end to end .
TL;DR
- Training crawlers collect text that may shape a future model. Retrieval crawlers fetch a page to answer a question being asked now.
- OpenAI and Anthropic each run one of each, under separate names you can control separately.
- PerplexityBot is retrieval only, and Perplexity says so directly.
- Google-Extended covers Gemini training and grounding, and is a control token rather than a visible user agent.
- AI Overviews are served from ordinary Google Search, so no AI-specific token governs them.
Short answer
The AI crawlers that matter divide into training crawlers, which collect text that may be used to build future models, and retrieval crawlers, which fetch pages to answer a live question and usually cite what they fetched. OpenAI runs GPTBot and OAI-SearchBot, Anthropic runs ClaudeBot and Claude-SearchBot, Perplexity runs PerplexityBot for retrieval only, and Google uses the Google-Extended token to cover Gemini.
01
Two jobs, not one
A training crawler collects text that may contribute to a future version of a model. The value to you is indirect and delayed. Nothing links back, nothing is attributed, and the effect, if there is one, appears whenever the next model is built. Whether that is worth anything to a business is a genuine judgment call, and it is reasonable to answer no.
A retrieval crawler fetches a specific page because a user asked something now, and the answer that follows usually names and links its sources. That is much closer to ordinary search traffic: a real person, with a live question, who can click through. For most businesses this is the one that matters, and blocking it by accident is the expensive mistake in this whole subject.
The reason this matters practically is that several operators run one of each under different names. Blocking a company is not a single decision, and a rule written against the wrong token gives you the opposite of what you wanted.
02
The crawlers, and what their operators say they do
The table below uses each operator’s own description. That matters in a category where most published crawler lists are compiled from other published crawler lists, and where names change without announcement. Where a description is quoted, the quote is from the vendor page linked beneath.
Two entries are worth reading twice. Perplexity states outright that its crawler is not used to collect training data, which makes it the clearest retrieval-only case in the list. Google-Extended is not a user agent at all, which has a consequence covered in the chapter on checking your work: you cannot confirm it in a log file, because there is nothing to see.
| Token | Operator | Job | What it affects |
|---|---|---|---|
| GPTBot | OpenAI | Training | Content that may be used to train future models |
| OAI-SearchBot | OpenAI | Retrieval | Whether pages can surface and be linked in ChatGPT search |
| ClaudeBot | Anthropic | Training | Web content that may contribute to training |
| Claude-SearchBot | Anthropic | Retrieval | Quality of search results shown inside Claude |
| PerplexityBot | Perplexity | Retrieval only | Whether pages surface and are linked in Perplexity |
| Google-Extended | Control token | Gemini training and grounding. Not Google Search |
Source OpenAI: Bots and crawlers (opens in a new tab) “GPTBot is used to make our generative AI foundation models more useful and safe.”
Source OpenAI: Bots and crawlers (opens in a new tab) “OAI-SearchBot is used to surface websites in search results in ChatGPT's search features.”
Source Anthropic: Does Anthropic crawl data from the web? (opens in a new tab) “ClaudeBot helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training.”
Source Anthropic: Does Anthropic crawl data from the web? (opens in a new tab) “Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses.”
Source Perplexity: PerplexityBot (opens in a new tab) “PerplexityBot is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models.”
03
Google is the odd one out, in two ways
The first difference is that Google-Extended has no user agent of its own. Google states that crawling is done with existing Google user agent strings and the robots.txt token is used in a control capacity. So the rule works, but it is invisible: there is no line in your access log that says Google-Extended, and no way to confirm the directive took effect by watching traffic. You have to trust the documented behavior, which is unusual in this subject and worth knowing before you spend an afternoon looking for something that was never going to appear.
The second difference is AI Overviews. Google describes its generative features as rooted in core Search ranking and quality systems, which means the Overview is assembled from the ordinary index rather than from a separate AI crawl. There is no AI-specific token that removes you from Overviews while leaving your rankings alone. The controls that exist for it are the ordinary snippet controls, and they have ordinary snippet consequences.
The practical conclusion is that the Google decision and the OpenAI or Anthropic decision are not the same shape. With OpenAI and Anthropic you are choosing between two named crawlers doing two different jobs. With Google you are choosing about Gemini, and separately deciding how much of your page Google Search may quote at all.
Source Google Search Central: Google crawlers and fetchers (opens in a new tab) “Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.”
Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “our generative AI features on Google Search are rooted in our core Search ranking and quality systems”
04
The list changes, so date your decision
Crawler names are added, renamed and split without notice, and a rule written against a token that no longer exists fails silently. The failure is quiet in the worst way: nothing errors, nothing appears in a report, and the crawler you meant to control simply falls through to whatever your wildcard group says.
The maintenance habit that costs almost nothing is to record, in a comment inside robots.txt itself, the date the list was last checked and the vendor pages it came from. That turns an invisible staleness problem into a visible one. Any file this guide would have you write should carry that comment.
- Write the date the crawler list was last checked into robots.txt as a comment
- Link the vendor documentation pages in the same comment, so the next person can recheck
- Recheck when you hear a new assistant is retrieving pages, not on a fixed schedule
- Treat a token you cannot find documentation for as a token you should not write a rule against
Questions this raises
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects content that may be used to train future OpenAI models. OAI-SearchBot fetches pages so they can surface and be linked in ChatGPT search results. They are separately controllable, and for most businesses the second is far more valuable than the first, because it is the one that can send a visitor.
Source OpenAI: Bots and crawlers (opens in a new tab) “GPTBot is used to make our generative AI foundation models more useful and safe.”
Source OpenAI: Bots and crawlers (opens in a new tab) “OAI-SearchBot is used to surface websites in search results in ChatGPT's search features.”
Does PerplexityBot train on my content?
Perplexity states that PerplexityBot is designed to surface and link websites in its search results and is not used to crawl content for AI foundation models. That makes it the clearest retrieval-only crawler among the major ones, and the easiest decision on this list for a business that wants referral traffic.
Source Perplexity: PerplexityBot (opens in a new tab) “PerplexityBot is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models.”
Is there a crawler for Google AI Overviews I can block?
No. Google describes its generative features as built on core Search ranking, so Overviews draw on the ordinary Search index rather than a separate AI crawl. Google-Extended covers Gemini, not Search. Limiting what Overviews can show means limiting what Google Search may quote from your page generally, which affects ordinary results too.
Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “our generative AI features on Google Search are rooted in our core Search ranking and quality systems”
Source Google Search Central: Google crawlers and fetchers (opens in a new tab) “Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.”