Skip to content
Visibility Bureau
Menu
Original research

How many sites actually have an llms.txt, and why the usual count is wrong

Adoption figures for llms.txt circulate freely and rarely state a method. Where one is stated, it is usually that the file was requested and the 200 responses were counted. That method has a specific, measurable flaw, and this study is about the size of it rather than about the file itself.

Run 2026-09-04 · Tranco W36Q9 · 5000 domains

TL;DR

  • Asking 5,000 domains for /llms.txt, 843 answered 200. Only 411 of those responses were an llms.txt.
  • More than half of what a status-code count calls adoption is something else, most often a site returning its own HTML for any path.
  • Real adoption in this sample is 8.2%, against the 16.9% a status count reports.
  • Some of the misses are genuine attempts: 15 sites served a substantial text file with the format’s one required element missing.
  • This says nothing about whether the file helps. Google states it ignores llms.txt.

Short answer

Of 5,000 domains sampled from the Tranco list, 843 returned HTTP 200 for /llms.txt but only 411 returned a file in the llms.txt format, which is 8.2% of the sample. The gap is mostly sites that answer 200 to every path and serve their own HTML, which a count based on status codes reads as adoption. Measured that way the figure is 16.9%, roughly twice the real one.

The question

What this asked

A large share of sites answer 200 to any path. A single-page application, a catch-all rewrite or an over-broad redirect will return the site’s own HTML rather than a 404, and it will do so for a path nobody has ever created. Counting those responses as adoption counts sites that have never heard of the format.

So the question is not how many sites have an llms.txt. It is how far apart the easy answer and the correct one are, and whether the gap is small enough to ignore. If it is small, the circulating figures are usable. If it is not, every argument built on them inherits the error.

The correct answer needs the response body classified rather than its status code counted. That is cheap to do and almost nobody does it, which is the same shape as the robots.txt study on this site: the naive measurement is a text or status match, and the sound one resolves what the thing actually is.

Method

How it was measured

Each domain was asked once for /llms.txt and for nothing else. The response body was then classified against the published format rather than counted by status.

Conformance here is the specification’s own bar and nothing stricter. llmstxt.org makes exactly one element required, an H1 naming the project, and everything after it optional. A response counts as an llms.txt when it is text, is not an HTML document, and contains that heading. A file failing this is not a near miss on a demanding standard; it is missing the only thing the format insists on.

The specification also permits a byte-order mark before the H1. An earlier version of the classifier tested for a heading at the start of the line and would have reported a conforming file with a BOM as broken. That was fixed before the run reported here, and the pilot was repeated. Getting it wrong would have been the same class of error this study is about, which is the reason for saying so rather than quietly correcting it.

  • Sample frame: the Tranco top list, generated for research use and issued with a permanent identifier so the same sample can be pulled again
  • Request https://{domain}/llms.txt once, follow redirects, record status, content type and body size
  • Classify the body: absent, HTML document, empty, text without the required H1, or a conforming file
  • Report the status-only count alongside the classified count, so the size of the difference is visible rather than asserted
  • Record the shape of the conforming files, because a file that exists and a file that is useful are different claims

Source llmstxt.org: The /llms.txt file, v2 (opens in a new tab) “An H1 with the name of the project or site. This is the only required section”

Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them.”

How this was collected

  • One request per host, for one path, with a user agent naming this page.
  • Only aggregates are published. No site is named as having a missing or malformed file. The few named below are cited as examples of a pattern, and none of them is doing anything wrong.
  • The collector is in the repository next to the robots.txt census and takes the same arguments, so the run can be repeated exactly.
  • Nothing was fetched beyond the single file. The links inside the conforming files were not requested, which is why this study makes no claim about whether they resolve.
Findings

What it found

01

Half of the apparent adoption is not adoption

Of 5,000 domains, 843 returned HTTP 200 for /llms.txt. Classifying the bodies leaves 411 that are actually the format. The other 432, slightly more than half of the apparent adopters, returned something else with a 200 attached.

Stated as shares of the sample, a status-only count reports 16.9% and the classified count gives 8.2%. The overstatement is a factor of 2.1. That is not a rounding difference, and any argument that leans on the higher figure is leaning on sites that do not have the file.

The direction of the error is worth noting. It only ever overstates, because a site cannot serve a conforming file and be counted as absent. So published adoption figures without a stated classification method should be read as an upper bound rather than an estimate.

What 5,000 domains returned for /llms.txt, sampled 4 September 2026
Response Count Share of sample
Answered HTTP 200 843 16.9%
Of those, a conforming llms.txt 411 8.2%
Of those, the site’s own HTML 386 7.7%
Of those, text without the required H1 41 0.8%
Of those, an empty body 30 0.6%
Absent or unreachable 4,132 82.6%

Source llmstxt.org: The /llms.txt file, v2 (opens in a new tab) “An H1 with the name of the project or site. This is the only required section”

02

What the other 432 responses turned out to be

The largest group by far is the site returning its own HTML: 386 responses, 45.8% of every 200 received. These are ordinary large sites whose routing answers any unmatched path with the application shell. Requesting a path that has never existed on one of them returns a full page and a 200, and nothing about that response says the file is missing.

Thirty responses were a 200 with an empty body, which is the same phenomenon with less to show for it. A further three were not text at all: two returned a tracking pixel and one returned JSON, because those hosts answer every path with the same thing.

The remaining 38 are the interesting ones, because they are real attempts. Each served plain text or markdown at that address, and each lacked the one element the format requires. Fifteen of them were over a kilobyte, so this is not a stub left behind by accident. One well-known site serves a megabyte of correctly formatted markdown links with no heading above them. Another opens with a line stating that it was produced by an llms.txt generator, and then goes straight to a second-level heading.

That last case is the one worth dwelling on. A tool built to produce this format is producing files that miss its only requirement, which means the error is being manufactured rather than typed.

The 843 responses that returned HTTP 200
What it actually was Count Share of the 200s
A conforming llms.txt 411 48.8%
The site’s own HTML page 386 45.8%
Text or markdown, no H1 38 4.5%
Empty body 30 3.6%
Not text at all (image or JSON) 3 0.4%

03

The files that are real are substantial

Among the 411 conforming files the median is 8,429 bytes, with 8 second-level sections and 41 links. These are not tokens dropped at the root to tick a box. Whoever published them assembled a structured index of the site.

The optional summary blockquote, which the specification suggests immediately after the H1 to give a reader the context for everything below, appears in 317 of them, or 77.1%. That is a high rate for an optional element and suggests most publishers are following the specification rather than guessing at a format.

So the picture is not one of widespread token effort. Adoption is narrow and the sites that have adopted have generally done the work properly. The measurement problem is entirely on the other side of the line, among sites that never adopted at all and get counted anyway.

04

What this study does not say

It does not say that publishing an llms.txt improves visibility. Google states plainly that it ignores the file, and nothing measured here bears on whether any other engine reads it. An adoption count is a count of a behavior, not evidence that the behavior works.

It also does not say the 411 files are useful. Conformance was tested, not accuracy: no link inside any file was requested, so a file listing pages that no longer exist counts here exactly the same as one that is current. That is a real gap and a candidate for the next run.

What it does support is narrower and, for anyone quoting a figure, more useful. If you read that some share of the web has adopted llms.txt, ask how the count was made. If it counted status codes, the true figure is roughly half of what you were told.

Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them.”

Checking this

What would show this is wrong

Every claim here is reproducible from the sample frame and the method above. If you repeat it and get a different answer, one of us has made a mistake and it is worth knowing which.

  • A repeat run at the same size returning a materially different conforming share. The two pilot runs on 500 domains both found exactly 52 conforming files while the raw 200 count moved between 115 and 120, so the classified figure should be the stable one.
  • A site classified as serving HTML that is in fact serving a genuine llms.txt with an HTML content type. That would mean the classifier is too strict and the real figure is higher.
  • A published adoption figure produced by classifying bodies rather than counting statuses that disagrees with this one. Two sound methods disagreeing would mean at least one of them is wrong.
  • The specification changing what it requires, which would make the conformance test measure the wrong thing.

Limits of this study

  • Only the root path was requested. The specification allows an llms.txt at any subpath and says the most specific one applies, so a site with a file at /docs/llms.txt and nothing at the root is counted here as absent. The true figure for "publishes one anywhere" is therefore higher than 8.2%.
  • One request per host at one moment. Reachability varies between runs: the two pilots differed by five hosts on the raw 200 count alone.
  • Tranco top domains are not a random sample of the web. They are large sites, which are both likelier to have heard of the format and likelier to run the routing that produces the false 200, so both figures here would move on a different frame.
  • Existence is not usefulness. Nothing here tests whether the links inside a file resolve, whether the content is current, or whether any engine reads it.
  • A site could serve a conforming file to a browser and something else to this collector. Nothing was done to detect that, and no attempt was made to disguise the request.
Run history

Every time this was run

This page keeps one address and records each run rather than publishing a new page per quarter. Old figures stay so the trend can be checked, including the runs where the method changed.

Results by run
Run Frame Sample HTTP 200 Conforming Note
2026-09-04 Tranco W36Q9 5000 843 411 (8.2%) Published run.
2026-09-04 Tranco W36Q9 500 120 52 (10.4%) Pilot, repeated after the byte-order mark fix. Same conforming count as the first pilot.
2026-09-04 Tranco W36Q9 500 115 52 (10.4%) First pilot. The raw 200 count moved between runs; the conforming count did not.
Questions

Questions this raises

Should I publish an llms.txt?

This study cannot tell you, and anyone citing an adoption figure to answer it is arguing from popularity. Google states it ignores the file. It costs very little to generate one and there is no evidence it does harm, so treat it as cheap and unproven rather than as a ranking factor. This site publishes one and refuses to sell llms.txt implementation as a service, for exactly that reason.

Source Google Search Central: Optimizing your website for generative AI features on Google Search (opens in a new tab) “Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them.”

Why is your figure lower than the ones I have seen?

Because the counts you have seen most likely treated an HTTP 200 as adoption. Slightly more than half of the 200 responses in this sample were not an llms.txt, usually because the site answers any unmatched path with its own HTML page. Classifying the body rather than the status halves the figure.

How do I check whether my own file conforms?

Open it and confirm the first heading is a single H1 naming the site, because that is the only element the specification requires and it is the one the failures in this study were missing. Then confirm your server returns a 404 for a path that does not exist, since a site that answers 200 to everything will appear to have the file whether or not it does.

Source llmstxt.org: The /llms.txt file, v2 (opens in a new tab) “An H1 with the name of the project or site. This is the only required section”

Does a missing H1 mean the file is ignored?

Unknown, and worth being clear about. The specification defines the format, it does not describe what any particular consumer does with a file that breaks it. A tolerant parser may well read the links anyway. The finding here is that the files exist and do not conform, not that they have been rejected by anything.

Worth saying plainly

This measures other people’s sites, not client work, and implies no track record. It was published because the method can be repeated by anyone who doubts it, which is a stronger form of evidence than a case study nobody outside the studio can verify.