Writing robots.txt rules that actually apply
This chapter covers one rule, because getting it wrong is common, the consequence is the opposite of what was intended, and nothing anywhere reports it. If you take a single thing from this guide, take this one.
Part of the guide to AI crawlers, end to end .
TL;DR
- A crawler obeys the group matching its own name. If one exists, the wildcard group does not apply to it at all.
- So adding a group for a named bot silently removes every wildcard rule from that bot.
- The fix is to repeat your shared rules inside every named group, not to add rules once at the top.
- Multiple groups with the same token merge, which is why the repetition is safe.
- This fails without an error, so it has to be tested by reading the file, not by watching for breakage.
Short answer
A crawler uses the robots.txt group that matches its own product token, and only falls back to the wildcard group when no matching group exists. That means adding a group for a named crawler removes every wildcard rule from that crawler. Anything you want applied to it has to be repeated inside its own group, which is safe because groups sharing a token are merged.
01
The rule, from the specification
The Robots Exclusion Protocol is standardized as RFC 9309, and it is explicit about how a crawler picks its rules. A crawler matches its product token against the user-agent lines, obeys the group it matches, and falls back to the wildcard group only when no matching group exists.
Read that in reverse and the trap appears. If a matching group does exist, the wildcard group does not apply to that crawler. Not partially, not as a default underneath. It is not consulted.
The specification also states that where more than one group matches the same token, the matching groups are combined into one. That is the property which makes the fix below safe: repeating a token is not a conflict, it is a merge.
Source IETF: RFC 9309: Robots Exclusion Protocol (opens in a new tab) “If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.”
02
What that means in practice
Picture the ordinary case. A site has a wildcard group disallowing a few paths that should never be crawled: an internal search endpoint, a staging directory, some parameterized URLs. It has worked for years. Then someone adds a group naming GPTBot, to control training access.
The moment that group is added, GPTBot stops obeying the wildcard disallows. It now has its own group, so the wildcard group is no longer its group, and every path that was denied to everyone is available to it. The site owner made a change intended to restrict one crawler and, in the same edit, gave that crawler more access than any other bot has.
Nothing reports this. The file is still valid, no tester flags it as an error because it is not an error, and traffic looks normal because a crawler quietly fetching a previously denied path does not look like anything. This site hit exactly this, and it was found by reading the generated file against the specification rather than by anything going wrong.
- Adding a named group is a permission change for that crawler, even when every line you added is a disallow
- The more named groups a file has, the more places a shared rule has to exist
- A robots.txt tester confirms syntax, not that your intent survived
- The symptom is silence, so the check has to be deliberate
03
How to write it so it holds
The correct shape is to treat every named group as self-contained. Whatever you want a crawler to be denied has to appear inside that crawler’s own group, even though it also appears in the wildcard group. The wildcard group then covers only the crawlers you did not name.
That means real duplication in the file, and the duplication is the point rather than a smell. If you maintain robots.txt by hand, the risk is that someone adds a disallow to the wildcard group later and does not copy it into the six named groups below. The durable fix is to generate the file from one list of shared rules and one list of tokens, so the repetition is produced rather than remembered. That is how this site does it: the shared disallows are written once in code and emitted into every group, which is what turned a rule people forget into a rule that cannot be forgotten.
If generating it is not practical, the fallback is a comment at the top of the file stating that shared rules must be copied into every group, positioned where the next person editing it will read it before they make the edit.
| Approach | What the file says | What GPTBot actually obeys |
|---|---|---|
| Rules once at the top | Wildcard group disallows /internal-search, then a GPTBot group disallows / | Only disallow: /. The /internal-search rule does not reach it |
| Rules repeated per group | Wildcard group disallows /internal-search, GPTBot group disallows /internal-search and / | Both rules, as intended |
04
Checking a file you did not write
Auditing someone else’s robots.txt takes about two minutes once you know what you are looking at. List every user-agent token in the file. For each named token, read only the lines in that group and ask whether they are sufficient on their own, ignoring everything above and below. Any rule you assumed applied but which is not inside that group does not apply to that crawler.
The pattern that should stop you is a file with a well-developed wildcard group at the top and several short named groups underneath, each containing one or two lines. That shape almost always means the named crawlers are exempt from rules the author believed were universal.
It is also worth checking that the file is served as plain text with a 200 status at the root of the domain, and that it is not accidentally being served from a redirect or a 404 handler that returns 200. A file that does not load correctly denies nothing, and this is a common failure on sites where the root is handled by an application rather than a static server.
Questions this raises
If I add a group for GPTBot, do my other robots.txt rules still apply to it?
No. RFC 9309 states that a crawler falls back to the wildcard group only when no group matches its own token. Once GPTBot has its own group, that group is the complete set of rules it obeys, and every wildcard rule stops applying to it. Anything you want it denied has to be repeated inside its group.
Source IETF: RFC 9309: Robots Exclusion Protocol (opens in a new tab) “If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.”
Is it safe to list the same crawler in two groups?
Yes. RFC 9309 says groups matching the same product token are combined into one group. That is what makes repeating shared rules inside a named group safe rather than contradictory, and it is why generating the file from a shared list works cleanly.
Source IETF: RFC 9309: Robots Exclusion Protocol (opens in a new tab) “If no matching group exists, crawlers MUST obey the group with a user-agent line with the "*" value, if present.”
Will a robots.txt tester catch this mistake?
Not usually. Testers check syntax and tell you whether a given URL is allowed for a given agent, which is useful only if you already suspected the problem and thought to test that combination. The file in the failing case is completely valid, so nothing flags it. Reading each named group in isolation is the check that finds it.