Skip to content
Visibility Bureau
Menu
Chapter 02 of 04

Deciding what to allow and what to block

Most of the argument about blocking AI crawlers is conducted as a single yes or no, usually about how someone feels concerning AI companies using web content. That framing produces bad decisions, because the two things you are actually deciding have almost nothing in common.

Part of the guide to AI crawlers, end to end .

TL;DR

  • Answer the training question and the retrieval question separately. They are not the same decision.
  • Blocking retrieval removes you from answers that would have linked you. That is the costly one.
  • Blocking training costs you no traffic today, and buys no traffic today either.
  • Publishers who sell access and businesses wanting referrals reach opposite conclusions, and both are right.
  • A block only affects future fetches. It does not reach into a model already trained.

Short answer

Decide the training question and the retrieval question separately, because their consequences are opposite. Blocking a retrieval crawler removes you from answers that would have named and linked you, which is a direct traffic cost. Blocking a training crawler costs no traffic and gains none, so it turns on whether the content could be sold to an AI company or must be withheld for another reason.

01

Split it into two questions

The first question is whether you want your content used to train future models. There is no traffic attached to the answer in either direction. Nothing links back, nothing is attributed, and no visitor arrives because a model was trained on your page. The case for allowing it is that being represented in a model may make an assistant more likely to describe your business accurately when asked. The case against is that you are supplying an input to a commercial product for nothing, and if your content has licensing value you may prefer to sell it.

The second question is whether you want to appear in answers that assistants generate right now. Here there is traffic attached. A retrieval crawler fetches your page because a user asked something, and the answer typically names and links the sources. Blocking it removes you from those answers, and the removal is total rather than gradual: you do not rank lower, you are not eligible.

Answered separately, most businesses reach a stable position quickly. Allow retrieval, because it behaves like search traffic and is the whole point of being findable. Decide training on its own merits, where a no costs nothing measurable.

02

Who correctly reaches the other answer

Blocking is not a mistake for everyone, and a guide that pretends otherwise is selling something. There are three situations where a broad block is the right call.

The first is a publisher whose archive is the product. If people pay for access to your writing, supplying it free as training data undercuts the thing you sell, and having it summarized in an answer can substitute for the visit rather than causing it. The second is anyone whose content is distinctive enough that an AI company might pay for it, where giving it away first removes the only negotiating position you had. The third is content under a confidentiality or contractual obligation that does not permit redistribution, where the decision is not commercial at all.

What these have in common is that the content itself has independent value. If your pages exist to bring in enquiries rather than to be the product, none of the three applies, and blocking retrieval mostly means being absent from a place your buyers are asking questions.

How the decision usually resolves by situation
Situation Training crawlers Retrieval crawlers
Business that wants enquiries Either. Costs nothing to allow Allow
Subscription publisher Block Consider blocking, weigh referrals
Content with licensing value Block until a deal exists Usually allow
Confidential or contractual content Block Block
Site not meant to be public at all Use authentication, not robots.txt Same

03

What robots.txt does not do

Robots.txt is a request, honored by well-behaved crawlers because their operators choose to honor it. It is not access control. It does not stop a crawler that ignores it, it does not stop a person, and it does not stop anyone who has already copied the page. If the real requirement is that something must not be read, the answer is authentication, not a directive in a public text file.

It is also worth being clear that the file is public and, on many sites, quietly informative. Anyone can read it, including competitors, and a long list of oddly specific disallowed paths tells a reader more about your site than you probably intended.

Finally, a block is not retroactive. It governs the next fetch, not the last one. Content already collected is collected, and no directive written today reaches into a model that has already been trained on it. That asymmetry is the main argument for deciding this deliberately now rather than getting to it eventually.

04

The part nobody can tell you

What allowing retrieval is worth in practice is not something anyone can quantify for you honestly. Whether an assistant names your business in an answer depends on your prominence, on what third-party sources say about you, on the phrasing of the question, and on retrieval behavior that changes without announcement. The figures circulating for this are vendor marketing, and the ones this site has tried to trace lead to aggregator articles citing each other rather than to any study.

What can be said with confidence is the direction and the asymmetry. Allowing retrieval makes you eligible for something and costs you almost nothing. Blocking it makes you ineligible with certainty. Between an uncertain upside and a certain exclusion, the decision does not need a percentage to be clear.

Questions

Questions this raises

Does blocking AI crawlers affect my Google rankings?

Not through the AI-specific tokens. Google-Extended covers Gemini and Google states it has no effect on Google Search, and the OpenAI, Anthropic and Perplexity tokens have nothing to do with Googlebot. The one thing that does reach ordinary results is restricting what Google Search may quote from your page, because AI Overviews are drawn from the Search index rather than a separate crawl.

Source Google Search Central: Google crawlers and fetchers (opens in a new tab) “Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.”

Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “our generative AI features on Google Search are rooted in our core Search ranking and quality systems”

If I allow retrieval crawlers, am I guaranteed to be cited?

No, and anyone promising that is selling something they cannot deliver. Allowing the crawler makes you eligible. Whether an assistant names you depends on prominence, on what independent sources say about you, and on retrieval behavior that changes without notice. The honest version of the claim is that blocking guarantees absence while allowing only creates a chance.

Can I block training but allow retrieval from the same company?

Yes, for the operators that run separate crawlers. OpenAI splits GPTBot from OAI-SearchBot and Anthropic splits ClaudeBot from Claude-SearchBot, so you can deny one and allow the other. This is the configuration most businesses end up wanting, and it is only possible because the tokens are separate.

Source OpenAI: Bots and crawlers (opens in a new tab) “GPTBot is used to make our generative AI foundation models more useful and safe.”

Source Anthropic: Does Anthropic crawl data from the web? (opens in a new tab) “Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses.”