Try it on your own data
Every report in this article is in the free plan. One snippet, cookieless, up to 10 sites.
Start freeA year ago AI crawlers were a curiosity in the server log. Today they are a measurable share of requests on most content sites, and the questions have changed from “what is this user agent?” to “should we let it in, and what do we get for it?” This article explains how TapCub identifies the main assistants, how the crawler ledger lines up your robots.txt and llms.txt against what actually happened, and how to read the one column that settles the debate: visitors referred.
What you will learn
- How GPTBot, ClaudeBot, PerplexityBot and Googlebot are told apart, and why user agent alone is not enough
- How to read the ledger: fetches vs. robots.txt vs. llms.txt vs. visitors referred
- A simple allow / limit / block decision rule and how to handle false positives
Who is fetching your pages
Three kinds of automated visitors matter here. Search crawlers (Googlebot, Bingbot) index pages so people can find them. AI training crawlers (GPTBot, ClaudeBot, Google-Extended, Bytespider) collect text to build models. AI answer crawlers (PerplexityBot, ChatGPT-User, Claude-User) fetch a page at the moment a person asks a question, so the assistant can cite it. The third group is the one that sends visitors back; the second rarely does.
The distinction matters because a single vendor often runs two or three of these with different user agents and different purposes. Blocking “OpenAI” as one thing blocks both the training crawler and the live-answer fetcher. The ledger keeps them on separate lines so you can treat them differently.
Five layers, because user agents lie
A declared user agent is a claim, not proof. Plenty of scrapers borrow a well-known bot name hoping to inherit its permissions. TapCub confirms the claim with four more layers: request headers (real crawlers send a consistent, sparse header set), client signals (a browser executes the snippet, a crawler does not), the ASN behind the IP address (a GPTBot request should come from the network range the vendor publishes), and request rate (crawlers fetch in bursts people cannot produce). A hit passes all five or it gets a different label.
Across the 197+ rules, each well-known assistant gets its own row, unknown automated traffic is grouped as “other bots”, and anything that passes as a browser with a human rhythm is counted as a person. The same classification drives the humans-only switch on the website analytics overview.
robots.txt says allowed; llms.txt says preferred
Two files talk to crawlers, and they say different things. robots.txt is a permission list: which user agents may fetch which paths. llms.txt is an invitation: a plain-text index of the pages you most want assistants to read and cite, with one-line descriptions. One is enforced by the crawler’s good behavior; the other is a suggestion that good crawlers follow.
The ledger reads both from your site and puts them next to the hits. That makes three kinds of row visible at a glance: a bot you allowed that fetched what you offered (fine), a bot you disallowed that fetched anyway (the row to open first), and a bot you allowed that never came (nothing to decide yet). A minimal robots.txt that blocks training but allows answers looks like this:
# Training crawlers: not allowed
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Bytespider
Disallow: /
# Answer / citation crawlers: allowed
User-agent: PerplexityBot
User-agent: ChatGPT-User
User-agent: Claude-User
Allow: /
# Search engines: allowed as before
User-agent: Googlebot
User-agent: Bingbot
Allow: /
Sitemap: https://example.com/sitemap.xmlReading a month of the ledger
Below is a sample month for a mid-sized content site. Read it left to right: how many fetches, what robots.txt said about that agent, whether the fetched pages were in llms.txt, and how many human visitors arrived with that assistant as the referrer in the same period.
Bot fetches · 30 days
6,254
▲ 31%
Blocked attempts
998
Visitors referred by AI
150
▲ 64%
AI crawler ledger · March 2026
robots.txt read 2 min ago| Agent | Purpose | Fetches | robots.txt | In llms.txt | Visitors referred |
|---|---|---|---|---|---|
| Googlebot | Search | 3,812 | Allowed | n/a | 9,640 |
| GPTBot | Training | 1,204 | Allowed | 62% | 38 |
| PerplexityBot | Answers | 240 | Allowed | 91% | 112 |
| ClaudeBot | Training | 0 (86 tried) | Disallowed | — | 0 |
| Bytespider | Training | 912 | Disallowed | — | 0 |
| Other bots | Mixed | 86 | Mixed | — | 0 |
Bytespider fetched 912 pages despite Disallow: block it at the edge, not in the file.
The column that settles arguments
Visitors referred is the only column that turns an opinion into a number. In the sample, PerplexityBot made a fifth of GPTBot’s fetches and sent three times the visitors, because it fetches to answer rather than to train. Googlebot still dominates both, which is a useful reminder that search is not going anywhere. ClaudeBot was disallowed and the ledger shows 86 attempts and zero fetches, which is exactly what a well-behaved crawler does.
Referrals from assistants show up in the sources report like any other channel, and you can build a segment of “visitors referred by AI assistants” to see which pages they land on and whether they convert. The marketing analytics view groups them under an AI assistants channel so they do not get lost in “direct”.
Allow, limit or block: a rule you can defend
With the ledger in front of you, the decision becomes a simple rule. If a crawler sends visitors, allow it and make sure llms.txt points it at your strongest pages. If a crawler takes a lot and sends nothing, decide whether being in the model is worth the bandwidth; many publishers say no and block training agents while allowing answer agents. If a crawler ignores robots.txt, block it at the edge rather than in the file, because the file clearly is not working.
Review the ledger monthly. New agents appear every quarter, and vendors change which agent does what. The AI crawler check gives you a quick read on a single user agent string, and the llms.txt guide covers writing the invitation file.
When the rules get it wrong
Five layers are good, not perfect. A corporate proxy that strips headers, a privacy browser that blocks the snippet, or a QA team running automated tests from an office network can all look like a bot. When you see one, mark it as human in the ledger. The reclassification applies to past data and to future hits from the same signature, so the correction sticks.
The reverse case, a scraper that passes as a person, is caught by request rate and ASN over time. If you suspect one, open the live visitors view, filter to that network, and watch the rhythm. People pause; scrapers do not.
TapCub Team
The people who design, build and support TapCub. We write about what we measure on our own site and what customers ask us most.