TapCub is live — analytics, insights and live chat, free on one platform
AI crawlers

GPTBot, ClaudeBot and the AI content ledger: who takes what

A ledger puts every AI crawler on its own line: fetches, what your robots.txt said, what llms.txt offered, and how many visitors came back. Then the allow-or-block decision is arithmetic.

TCTapCub Team Published Mar 26, 2026 Last updated Sep 4, 2026 11 min read AI crawlers

A year ago AI crawlers were a curiosity in the server log. Today they are a measurable share of requests on most content sites, and the questions have changed from “what is this user agent?” to “should we let it in, and what do we get for it?” This article explains how TapCub identifies the main assistants, how the crawler ledger lines up your robots.txt and llms.txt against what actually happened, and how to read the one column that settles the debate: visitors referred.

What you will learn

  • How GPTBot, ClaudeBot, PerplexityBot and Googlebot are told apart, and why user agent alone is not enough
  • How to read the ledger: fetches vs. robots.txt vs. llms.txt vs. visitors referred
  • A simple allow / limit / block decision rule and how to handle false positives

Who is fetching your pages

Three kinds of automated visitors matter here. Search crawlers (Googlebot, Bingbot) index pages so people can find them. AI training crawlers (GPTBot, ClaudeBot, Google-Extended, Bytespider) collect text to build models. AI answer crawlers (PerplexityBot, ChatGPT-User, Claude-User) fetch a page at the moment a person asks a question, so the assistant can cite it. The third group is the one that sends visitors back; the second rarely does.

The distinction matters because a single vendor often runs two or three of these with different user agents and different purposes. Blocking “OpenAI” as one thing blocks both the training crawler and the live-answer fetcher. The ledger keeps them on separate lines so you can treat them differently.

Five layers, because user agents lie

A declared user agent is a claim, not proof. Plenty of scrapers borrow a well-known bot name hoping to inherit its permissions. TapCub confirms the claim with four more layers: request headers (real crawlers send a consistent, sparse header set), client signals (a browser executes the snippet, a crawler does not), the ASN behind the IP address (a GPTBot request should come from the network range the vendor publishes), and request rate (crawlers fetch in bursts people cannot produce). A hit passes all five or it gets a different label.

Across the 197+ rules, each well-known assistant gets its own row, unknown automated traffic is grouped as “other bots”, and anything that passes as a browser with a human rhythm is counted as a person. The same classification drives the humans-only switch on the website analytics overview.

robots.txt says allowed; llms.txt says preferred

Two files talk to crawlers, and they say different things. robots.txt is a permission list: which user agents may fetch which paths. llms.txt is an invitation: a plain-text index of the pages you most want assistants to read and cite, with one-line descriptions. One is enforced by the crawler’s good behavior; the other is a suggestion that good crawlers follow.

The ledger reads both from your site and puts them next to the hits. That makes three kinds of row visible at a glance: a bot you allowed that fetched what you offered (fine), a bot you disallowed that fetched anyway (the row to open first), and a bot you allowed that never came (nothing to decide yet). A minimal robots.txt that blocks training but allows answers looks like this:

robots.txttraining blocked, answers allowed
# Training crawlers: not allowed
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Bytespider
Disallow: /

# Answer / citation crawlers: allowed
User-agent: PerplexityBot
User-agent: ChatGPT-User
User-agent: Claude-User
Allow: /

# Search engines: allowed as before
User-agent: Googlebot
User-agent: Bingbot
Allow: /

Sitemap: https://example.com/sitemap.xml

Reading a month of the ledger

Below is a sample month for a mid-sized content site. Read it left to right: how many fetches, what robots.txt said about that agent, whether the fetched pages were in llms.txt, and how many human visitors arrived with that assistant as the referrer in the same period.

app.tapcub.com/sites/demo/bots/ledger

Bot fetches · 30 days

6,254

▲ 31%

Blocked attempts

998

Visitors referred by AI

150

▲ 64%

AI crawler ledger · March 2026

robots.txt read 2 min ago
AgentPurposeFetchesrobots.txtIn llms.txtVisitors referred
GooglebotSearch3,812Allowedn/a9,640
GPTBotTraining1,204Allowed62%38
PerplexityBotAnswers240Allowed91%112
ClaudeBotTraining0 (86 tried)Disallowed—0
BytespiderTraining912Disallowed—0
Other botsMixed86Mixed—0

Bytespider fetched 912 pages despite Disallow: block it at the edge, not in the file.

Sample data
Sample ledger. Highlighted row: fewest fetches per visitor referred.

The column that settles arguments

Visitors referred is the only column that turns an opinion into a number. In the sample, PerplexityBot made a fifth of GPTBot’s fetches and sent three times the visitors, because it fetches to answer rather than to train. Googlebot still dominates both, which is a useful reminder that search is not going anywhere. ClaudeBot was disallowed and the ledger shows 86 attempts and zero fetches, which is exactly what a well-behaved crawler does.

Referrals from assistants show up in the sources report like any other channel, and you can build a segment of “visitors referred by AI assistants” to see which pages they land on and whether they convert. The marketing analytics view groups them under an AI assistants channel so they do not get lost in “direct”.

Allow, limit or block: a rule you can defend

With the ledger in front of you, the decision becomes a simple rule. If a crawler sends visitors, allow it and make sure llms.txt points it at your strongest pages. If a crawler takes a lot and sends nothing, decide whether being in the model is worth the bandwidth; many publishers say no and block training agents while allowing answer agents. If a crawler ignores robots.txt, block it at the edge rather than in the file, because the file clearly is not working.

Review the ledger monthly. New agents appear every quarter, and vendors change which agent does what. The AI crawler check gives you a quick read on a single user agent string, and the llms.txt guide covers writing the invitation file.

When the rules get it wrong

Five layers are good, not perfect. A corporate proxy that strips headers, a privacy browser that blocks the snippet, or a QA team running automated tests from an office network can all look like a bot. When you see one, mark it as human in the ledger. The reclassification applies to past data and to future hits from the same signature, so the correction sticks.

The reverse case, a scraper that passes as a person, is caught by request rate and ASN over time. If you suspect one, open the live visitors view, filter to that network, and watch the rhythm. People pause; scrapers do not.

TC

TapCub Team

The people who design, build and support TapCub. We write about what we measure on our own site and what customers ask us most.

See your own crawler ledger

The bot view separates search, AI and other crawlers from people, and reads your robots.txt and llms.txt automatically.

See your data clearly. Find your growth.

Every click, backed by data. Install one line of code and see your first numbers in a minute.

No credit card · Free plans for analytics and chat · Cookieless analytics