Skip to content
Visibility Bureau
Menu
Original research

Research

Studies run on this site, published with the method, the sample and the limitations attached. With no case studies yet, a dataset is the honest substitute for proof, and a better one: a testimonial has to be taken on trust, a method can be repeated.

  • Run 2026-08-28 · 5000 domains

    How many sites broke robots.txt while trying to control AI crawlers

    Of 5,000 domains sampled from the Tranco list, 2,841 served a robots.txt and 591 both named an AI crawler and had specific wildcard rules to lose. 168 of those, 28.4%, now let that crawler reach at least one path blocked for every other bot. The cause is RFC 9309 group precedence: a named group replaces the wildcard group rather than adding to it.

    Frame: Tranco W36Q9. Method and limitations are on the page, and the result is reproducible.

  • Run 2026-09-04 · 5000 domains

    How many sites actually have an llms.txt, and why the usual count is wrong

    Of 5,000 domains sampled from the Tranco list, 843 returned HTTP 200 for /llms.txt but only 411 returned a file in the llms.txt format, which is 8.2% of the sample. The gap is mostly sites that answer 200 to every path and serve their own HTML, which a count based on status codes reads as adoption. Measured that way the figure is 16.9%, roughly twice the real one.

    Frame: Tranco W36Q9. Method and limitations are on the page, and the result is reproducible.

Every study here can be re-run by anyone who wants to check it, and the figures from previous runs stay on the page rather than being quietly replaced. The guides explain the mechanisms these studies measure, the free tools let you run the same test against your own site, and what I will not sell you sets out the evidence standard all of it is written to.