Did naming an AI crawler break your robots.txt?
Adding a group for GPTBot removes your wildcard rules from GPTBot. This resolves your file the way a crawler does and tells you what each named crawler can actually reach, then compares you against 5,000 domains I measured.
Check a domain
Why this happens
RFC 9309 says a crawler obeys the group matching its own name, and falls back to the wildcard group only when no named group matches it. Read that in reverse and the trap appears: once a crawler has its own group, the wildcard group does not apply to it at all. Not partially, not as a default underneath. It is not consulted.
So a site that has quietly denied an internal search path to everyone for years, then adds a group naming GPTBot to control training access, has in the same edit given GPTBot more access than any other crawler has. Nothing errors. The file is still valid. No tester flags it, because it is not a syntax mistake.
Source IETF: RFC 9309: Robots Exclusion Protocol (opens in a new tab) “If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.”
The same intent, written two ways
| Approach | What the file says | What GPTBot obeys |
|---|---|---|
| Rules once at the top | Wildcard group denies /internal-search, then a GPTBot group allows / | Only the allow. The /internal-search rule never reaches it |
| Rules repeated per group | Wildcard group denies /internal-search, GPTBot group denies it too | Both rules, as intended |
The duplication is the point rather than a smell. Groups that share a token are merged by the specification, so repeating a rule inside a named group is safe. If you maintain the file by hand, the durable fix is to generate it from one shared list, which is how this site does it.
How common is this?
I measured it, because nobody had. Of 5,000 domains sampled from Tranco W36Q9 on 2026-08-28, 591 both named an AI crawler and had specific rules to lose. 168 of them, 28.4%, had lost at least one.
The method, the limitations and what would prove it wrong are all on the study page. It also reports the least convenient finding: Googlebot and Bingbot groups cause this slightly more often than AI crawlers do, so it is a long-standing robots.txt problem that AI adoption has multiplied rather than an AI problem.
About this checker
What does this actually check?
Whether a crawler you named in robots.txt can reach a path you block for everyone else. Under RFC 9309 a named group replaces the wildcard group rather than adding to it, so adding a group for GPTBot silently removes your wildcard rules from GPTBot. This resolves the file the way a crawler does and reports what each named crawler is actually allowed.
Source IETF: RFC 9309: Robots Exclusion Protocol (opens in a new tab) “If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.”
Does this fetch anything other than robots.txt?
No. It requests one URL, https://yourdomain/robots.txt, and nothing else. Redirects are followed only to that same path on another host, and only over https, because the site being checked writes the redirect and could otherwise choose the destination. It reports the rule patterns involved but never requests those paths, because the point is to count the problem rather than exercise it. That constraint is enforced on the server, not in the browser.
Why do the results differ from other robots.txt testers?
Most testers answer whether a syntax error exists, and this file usually has none. The failure here is a valid file that means something other than what its author intended, which only shows up if you resolve group precedence rather than search the text for a token. That is also why the published crawler censuses miss it.
It says my site is clean. Am I safe?
It means your named groups repeat the protections your wildcard group sets, which is the specific mistake this checks for. It is not a general robots.txt audit and it is not a security check. Robots.txt is a request that well-behaved crawlers honor, never access control, so anything that genuinely needs protecting belongs behind authentication.
Worth saying plainly
This checks one specific mistake. It is not a robots.txt audit, not an SEO audit, and not a security check. Robots.txt is a request that well-behaved crawlers honor, so it never protected anything from a crawler that ignores it.
If you want the detail
- Writing robots.txt rules that actually apply
- The full guide to AI crawlers
- The study behind the benchmark
- Every free tool here
Last updated 2026-09-04