Skip to content
Visibility Bureau
Menu
Chapter 04 of 04

Checking it worked

Having made a decision and written the rules, the natural next question is whether any of it took effect. Most of that is answerable from data you already have, and the parts that are not answerable are worth knowing about, because they are the parts vendors sell tools for.

Part of the guide to AI crawlers, end to end .

Last updated 2026-08-28

TL;DR

  • Your server log is the primary evidence. It records what was actually fetched, not what should have been.
  • A user agent string is self-reported and trivially faked, so verify by IP range before trusting it.
  • Google-Extended cannot be confirmed in logs at all, because it is not a user agent.
  • Referrals from assistants appear in analytics as ordinary referrers and can be segmented without any tool.
  • No check tells you whether you were cited. That needs prompts run and recorded, and it is a weaker signal than it looks.

Short answer

Check crawler access in your own server logs, which record what was actually requested, and verify identity by IP range rather than trusting the user agent string. Referral traffic from assistants shows up in analytics as ordinary referrers. Neither method tells you whether you were cited in an answer, and Google-Extended cannot be verified at all because it has no user agent of its own.

01

Start with the server log

The access log is the only record of what actually happened. It shows which agent requested which path, when, and what status it received. Analytics will not show you this, because analytics runs JavaScript and crawlers generally do not, so a crawler visit is invisible to it by design.

The useful first pass is simply to count requests by user agent over a month and look for the tokens from the first chapter. Three outcomes are informative. Seeing a crawler you intended to block means the rule is not reaching it, and the group-matching trap from the previous chapter is the first thing to check. Seeing nothing at all from a crawler you allowed usually means it has not found reason to fetch you rather than that something is broken. Seeing a crawler hit paths you thought were denied is the specific signature of the named-group problem.

If your host does not give you raw logs, the equivalent is usually available in a CDN dashboard, which will typically break traffic down by user agent. What matters is that it is request-level data from the server side rather than anything measured in a browser.

  • Count requests by user agent over a month, not a day. Crawl frequency is lumpy
  • Check status codes too. A crawler receiving 403 or 404 is a different problem from one being denied
  • Compare the paths fetched against what you intended to allow
  • A crawler absent entirely is normal, and is not evidence of a misconfiguration

02

Do not trust the user agent string

A user agent is self-reported text. Anything can claim to be GPTBot, and scrapers routinely do, because it is a single header and there is no barrier to setting it. Treating that string as identity will give you a picture of AI crawler activity that is partly fiction, and it inflates in the direction people want to believe.

The check that works is to verify the requesting IP address against the ranges the operator publishes. The major operators publish these deliberately so that site owners can do exactly this. The workflow is unglamorous and reliable: take the IPs behind the requests claiming a given token, check them against the published ranges, and discard the ones that do not match. If a large share fails, you have learned something more useful than the original count.

This matters beyond tidiness. Decisions get made on these numbers, and a report claiming heavy AI crawler interest that turns out to be scrapers wearing a borrowed name is worse than no report, because it is confident.

03

What cannot be verified, and what to do instead

Google-Extended is the clean example. Google states it has no separate user agent string and that the robots.txt token is used in a control capacity, so there is nothing to find in a log. The directive either works as documented or it does not, and no amount of log analysis will tell you which. This is worth knowing before someone spends a day looking.

The larger thing you cannot check from logs is whether you were actually named in an answer. A retrieval crawler fetching your page does not mean the answer used it, and an answer naming you does not always involve a fresh fetch. The two are related but neither implies the other, so a log cannot settle it.

The honest method for that question is to run a fixed set of prompts on a schedule and record what comes back: whether you were named, how you were described, and which sources the answer drew on. It is manual and it is a weaker instrument than it appears, because answers to identical questions vary between runs. What makes it worth doing is the description rather than the count. Being named inaccurately is a fixable problem, and you can only fix it if you know about it.

What each check can and cannot tell you
Question Where the answer is What it cannot settle
Did a crawler fetch my pages? Server or CDN access log Whether the fetch was used in an answer
Was it really that crawler? IP checked against the published ranges Nothing. This one is definitive
Is Google-Extended being obeyed? Nowhere. It has no user agent Everything. The rule is documented, not observable
Did anyone arrive from an assistant? Analytics referrer, segmented by hostname Visits that carry no referrer, so it reads as a floor
Am I named in answers? A fixed prompt set, run and recorded over time A stable figure. Identical prompts vary between runs

Source Google Search Central: Google crawlers and fetchers (opens in a new tab) “Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.”

04

Finding the traffic that did arrive

When an assistant links you and someone clicks, that visit lands in your analytics as an ordinary referral, with the assistant’s domain as the referrer. No special integration is required to see it, and no tool is required to segment it. Building a segment or filter for the assistant hostnames gives you the same view a paid dashboard would sell you.

Two cautions make the numbers honest. The volumes are usually small compared with search, and a low number is not evidence that something is broken. And some assistant traffic arrives without a referrer at all, for instance from a desktop application rather than a browser, which means what you can see is a floor rather than a total. Report it as a floor.

The pattern worth watching is not the daily figure but the direction over months, alongside what the prompt log says about how you are being described. Traffic tells you something arrived. The prompt log tells you what was said. Neither is sufficient alone, and together they are about as good as this gets without inventing precision.

Questions

Questions this raises

How do I know if GPTBot is really GPTBot?

Verify the requesting IP address against the ranges OpenAI publishes, rather than trusting the user agent header. The header is self-reported text that anything can set, and scrapers commonly claim well-known crawler names. Any count built on the header alone will overstate real crawler activity.

Source OpenAI: Bots and crawlers (opens in a new tab) “GPTBot is used to make our generative AI foundation models more useful and safe.”

Why do AI crawlers not show up in my analytics?

Because analytics runs in the browser and crawlers generally do not execute JavaScript, so their visits are never recorded. Crawler activity lives in server logs or your CDN dashboard. Analytics will show you the human who clicked through from an assistant, which is a different and also useful measurement.

Can I measure whether I am being cited in AI answers?

Only by running prompts and recording the results, and the measurement is weaker than it looks because answers to identical questions vary between runs. It is worth doing for what it says about how you are described rather than for a citation count. Any product offering a precise citation figure is reporting its own sampling, not a number that exists.