Skip to content
Visibility Bureau
Menu
Original research

How many sites broke robots.txt while trying to control AI crawlers

There are already several published counts of how many sites block AI crawlers. This is not one of them. It asks a different question: of the sites that changed robots.txt to control an AI crawler, how many broke their existing rules in the process without noticing.

Run 2026-08-28 · Tranco W36Q9 · 5000 domains

TL;DR

  • Naming a crawler in robots.txt replaces your wildcard rules for it rather than adding to them.
  • Across 5,000 domains, 168 of the 591 sites at risk had done exactly that, which is 28.4%.
  • The most common trigger is GPTBot, present in 108 of the 168 cases.
  • It is not an AI problem. Googlebot and Bingbot groups cause it slightly more often.
  • Most of what leaks is crawl hygiene, but not all of it: about half also expose account-shaped paths.

Short answer

Of 5,000 domains sampled from the Tranco list, 2,841 served a robots.txt and 591 both named an AI crawler and had specific wildcard rules to lose. 168 of those, 28.4%, now let that crawler reach at least one path blocked for every other bot. The cause is RFC 9309 group precedence: a named group replaces the wildcard group rather than adding to it.

The question

What this asked

Under RFC 9309, a crawler obeys the group matching its own name and falls back to the wildcard group only when no named group matches. Read in reverse, that means adding a group for GPTBot removes every wildcard rule from GPTBot. The rules you thought applied to everyone stop applying to the one crawler you were trying to control.

The mechanism is not a secret. Several articles explain it, including one titled "Your robots.txt Is Doing the Opposite of What You Think". What did not exist anywhere, as far as could be established before running this, was a number. Nobody had measured how often it actually happens.

That gap has a cause worth naming. Measuring it requires resolving group precedence the way a crawler does, and the tooling behind the published crawler censuses matches tokens with text search instead. A grep cannot see this defect, because the file contains no wrong characters. It is correct robots.txt that means something other than what its author intended.

Method

How it was measured

Every domain in the sample was asked for its robots.txt and nothing else. The file was then parsed with a resolver written against RFC 9309: groups sharing a token merge, a named group replaces the wildcard group, the most specific path match wins by pattern length, and an allow beats a disallow of equal length.

For each AI crawler token with its own group, the test is simple. Take every specific path the wildcard group denies, and ask whether this crawler is now permitted to fetch it. If it is, the site has lost a protection it still believes it has.

One distinction does most of the work, and it was added after checking early results by hand rather than being designed in. A file that denies everything with a blanket rule and then names the crawlers it permits is running an allowlist on purpose. Facebook and Netflix both do this. Counting them as broken would have inflated the headline by roughly a third, so sites whose only wildcard rule is a blanket deny are excluded and reported separately.

  • Sample frame: the Tranco top list, which is generated for research use and issued with a permanent identifier so the exact sample can be pulled again
  • Fetch https://{domain}/robots.txt, follow redirects, record the status, take no other request
  • Parse with an RFC 9309 resolver, verified against both worked examples in section 2.2.1
  • Population at risk: sites with a named AI crawler group AND specific wildcard disallows to lose
  • Defect: that crawler is allowed at least one path the wildcard group denies
  • Control: the identical test for Googlebot and Bingbot, so the study can tell an AI phenomenon from a robots.txt phenomenon

Source IETF: RFC 9309: Robots Exclusion Protocol (opens in a new tab) “If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.”

How this was collected

  • Only robots.txt was requested. No path exposed by the defect was ever fetched. The point is to count the mistake, not to exercise it.
  • Requests were rate limited with a descriptive user agent naming this page, and each host was asked once.
  • Results are published as aggregates. No affected site is named, here or anywhere else. They are real businesses with a misconfiguration, and publishing a target list would be indefensible.
  • The parser has its own spec tests, and the collection script refuses to run if they fail.
Findings

What it found

01

Nearly 3 in 10 at-risk sites have the defect

Of 5,000 domains, 2,841 served a parseable robots.txt. 763 of those named at least one AI crawler. Narrowing to the population where the defect is possible at all, meaning sites that also had specific wildcard rules to lose, leaves 591. Of those, 168 had lost at least one protection: 28.4%.

The study was run four times. The defective count landed between 168 and 174 and the share between 28.4% and 28.9%, with the variation coming from hosts that time out on one run and answer on the next. A pilot on the top 500, by the same method, found 32.3%. Treat the figure as close to 28% rather than as a precise value, and see the run history below.

The share of all reachable sites is much lower, at 5.9%, because most sites have not named an AI crawler at all. Both numbers are true and they answer different questions. The one that matters to a site owner is the first: if you have edited robots.txt for AI crawlers, this is roughly the chance you got it wrong.

Population at each step, 5,000 domains sampled 28 August 2026
Step Count Share
Domains sampled 5,000
Served a parseable robots.txt 2,841 56.8% of sampled
Named at least one AI crawler 763 26.9% of reachable
At risk: named a crawler and had specific rules to lose 591 20.8% of reachable
Defective: crawler reaches a path denied to others 168 28.4% of at risk
Deliberate allowlist, excluded as intentional 49

02

This is not an AI problem, and saying otherwise would be dishonest

The same test applied to Googlebot and Bingbot groups found 184 affected sites, slightly more than the 168 from AI crawler groups. That result is inconvenient for a headline and it is the most important thing in this study.

The defect is a long-standing property of how robots.txt works, and it has been quietly costing sites their crawl rules for as long as anyone has been adding named groups for search engines. What AI has changed is the rate of editing. A file that had sat untouched for years is now being edited to add GPTBot, ClaudeBot and half a dozen others, and every one of those edits is a chance to trigger it.

So the honest framing is not that AI crawlers are dangerous. It is that AI crawler adoption has become the largest new source of a mistake that was already common, and that most site owners have never been told the rule that causes it.

03

GPTBot is the most common trigger, by volume of adoption

GPTBot appears in 108 of the 168 defective files, followed closely by the other two OpenAI tokens. That ordering tracks adoption rather than anything specific to OpenAI: GPTBot was the first widely publicized AI crawler token, so it is the one most sites added first and the one most likely to be present at all.

The pattern worth noticing is that the counts cluster. A site that broke its rules for GPTBot usually broke them for every AI crawler it named at the same time, because the tokens were typically added together in one edit, in the same wrong shape.

Defective sites by crawler token, of 168 total
Token Sites Operator
GPTBot 108 OpenAI
OAI-SearchBot 101 OpenAI
ChatGPT-User 98 OpenAI
ClaudeBot 85 Anthropic
PerplexityBot 84 Perplexity
Google-Extended 78 Google
CCBot 49 Common Crawl
Claude-SearchBot 42 Anthropic
Amazonbot 39 Amazon
anthropic-ai 38 Anthropic

04

What actually leaks, and what it is not

Most of what these sites lost is crawl hygiene. The two most common categories are query and faceted URLs, at 55% of defective sites, and internal search results, at 51%. Those rules exist to stop crawlers wasting budget on infinite parameter combinations and generating duplicate content, not to protect anything.

The uncomfortable part is that about half also expose paths matching account, admin, login or profile patterns. That sounds worse than it is, and the distinction matters. Robots.txt is not access control. Anything genuinely protected sits behind authentication, and a crawler being permitted to request a URL is not the same as being able to read what is behind it. What has actually been lost is a request to keep away, which well-behaved crawlers honor and nothing else does.

The classifier behind this table is crude and the numbers should be read as indicative. 88% of defective sites leaked at least one path it could not confidently categorize, which is why the categories overlap and why none of them should be quoted as a precise figure. It is included because the shape is informative and excluding it would be hiding the least flattering part of the data.

Categories of leaked path across 168 defective sites. Sites appear in more than one row, and 88% also leaked an unclassified path.
What the rule was protecting Sites Share
Query and faceted URLs 93 55%
Internal search results 85 51%
Account, admin and profile paths 80 48%
Endpoints and static assets 72 43%
Alternate renderings and feeds 54 32%
Cart and checkout 27 16%
Unclassified 147 88%

05

The most common shape is a site that tried

The failures are not careless. Reading the affected files, the recurring pattern is a site that understood the rule well enough to repeat some of its wildcard disallows inside the named group, and missed the rest.

One site in the sample denies two paths to everyone, repeated one of them into its named AI crawler group, and forgot the other. Another repeated eight of its twelve rules and missed four, among them its internal search and its tracking-parameter URLs. Neither file contains an obvious error. Both look careful, and both mean something other than what their author intended.

That is why this is worth measuring rather than explaining again. The people getting it wrong are the ones paying attention. A rule you have to remember to apply in six places is a rule that will eventually be applied in five.

Checking this

What would show this is wrong

Every claim here is reproducible from the sample frame and the method above. If you repeat it and get a different answer, one of us has made a mistake and it is worth knowing which.

  • Pull the same Tranco list, re-run the method, and get a materially different share. The list identifier and sample size are recorded for exactly this.
  • Show that the RFC 9309 resolver used here mis-implements group precedence. It is tested against both worked examples in section 2.2.1, and if those tests are wrong the numbers are wrong.
  • Show that a major crawler does not follow the specification here, and instead merges the wildcard group with the named group. That would make the defect theoretical for that crawler.
  • Demonstrate that the excluded allowlist pattern is more common than measured, which would mean this study still over-counts.

Limits of this study

  • The frame is a top-sites list, so the finding describes prominent sites. It says nothing about small business sites, which are a different population and may well be worse.
  • Reachability varies between runs by around 1%, because some hosts time out or block automated requests. The defect count was identical across repeated runs, but the denominator moves slightly.
  • The test uses a representative path synthesized from each rule pattern. It is sound for deciding whether a rule applies, and it is not a claim that a specific URL exists.
  • Crawlers self-report their behavior. This measures what the specification says should happen, not confirmation that every operator implements it that way.
  • A defect here means a protection was lost, not that anything sensitive was exposed. The two are different, and the next section separates them.
Run history

Every time this was run

This page keeps one address and records each run rather than publishing a new page per quarter. Old figures stay so the trend can be checked, including the runs where the method changed.

Results by run
Run Frame Sample At risk Defective Note
2026-08-28 Tranco W36Q9 5000 591 168 (28.4%) Published run. Leak categories recorded.
2026-08-28 Tranco W36Q9 5000 593 170 (28.7%) Repeat run, kept to show the spread from host reachability.
2026-08-28 Tranco W36Q9 5000 602 174 (28.9%) First run at full size.
2026-08-28 Tranco W36Q9 500 65 21 (32.3%) Pilot. Method corrected after hand-checking: deliberate allowlists excluded.
Questions

Questions this raises

Does this mean those sites have a security problem?

Almost never. What is usually lost is a crawl-hygiene rule: internal search results, tracking-parameter URLs, faceted filters, print versions. Those are blocked to protect crawl budget and avoid duplicate content, not because they are sensitive. Robots.txt is not access control in any case, so nothing that needed protecting was ever protected by it.

How do I check my own site?

Take each named user-agent block in your file and cover up everything else. What remains is the complete set of rules that crawler obeys. If a protection you rely on is not inside that block, it is not protecting you. The guide on writing robots.txt rules works through a real file and shows the fix.

Source IETF: RFC 9309: Robots Exclusion Protocol (opens in a new tab) “If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.”

Why not just name the affected sites?

Because they are real businesses with a misconfiguration they did not know about, and a published list would function as a target list while helping nobody fix anything. The aggregate supports every claim made here. Anyone who wants to verify a specific site can run the method themselves against that site.

Will you re-run this?

Quarterly, on this page rather than a new one, with previous runs kept in the table above. The interesting question is whether the share falls as the rule becomes better known, and that can only be answered if the old figures stay visible.

Worth saying plainly

This study measures other people’s robots.txt files. It is not a client result and implies no track record. The reason to publish it is that a method anyone can repeat is better evidence of how this studio works than a testimonial nobody can check, and the studio had no case studies to offer instead.