TapCub is live — analytics, insights and live chat, free on one platform
Help center

Bots and AI crawlers

Recognized bots are counted separately by default. Here is how detection works, how to switch between humans and all traffic, and how to read the AI ledger.

Website analytics6 min readLast updated Sep 25, 2026TapCub team

Still have a question?

Could not find it? A person reads every message, usually within a few hours on business days.

Chat with us

What happens by default

New sites start in recognize and count separately mode. A request recognized as a bot goes to the Bots report and is left out of visitors, sessions, pageviews and every human report. It also does not consume your monthly event quota.

You can change the mode under Web analyticsConfigurationBots: drop discards bot hits entirely, and do not recognize mixes them into human numbers. The last one is kept for comparisons with older tools and is not recommended.

Five detection layers, 197+ rules

Detection runs on every hit before it is stored. A hit is marked as a bot when enough layers agree:

  1. User agent. Known crawler strings for search engines, AI companies, monitoring services and SEO tools — the largest share of the rule set.
  2. Request headers. Missing or inconsistent headers that real browsers always send.
  3. Client signals. Whether the script could run, screen and timing signals that headless browsers get wrong.
  4. Network (ASN). Ranges that belong to cloud providers and known crawler operators.
  5. Frequency. Request rates no person produces, such as hundreds of pages per minute from one source.

The rule set is maintained centrally and currently holds 197+ rules; you do not update anything. Each bot row shows which layers matched and a trust level — verified, suspected or unknown. Unknown means we could not confirm the operator, not that the traffic is fake.

Read the Bots report

Open Web analyticsBots. There are three tabs:

  • AI ledger — per AI operator: fetches for training and for answer engines, how many human visitors that operator sent back, and the ratio between the two. A high ratio means many reads and few people.
  • AI visibility — the same data from the other direction: visits sent back divided by fetches, so you can see which assistants actually cite you.
  • All bots — name, operator, purpose (search, AI training, AI answers, monitoring), requests, pages, most common status code, first and last seen. Search engines also get a small health table with success rate, 404s and 5xx.
The toggle at the top of every analytics report switches between Humans and All traffic. Humans is the default and the number to report; All traffic is for debugging crawl volume.

Compare with robots.txt and llms.txt

The Crawl health view puts your robots.txt and llms.txt rules next to what crawlers actually fetched. A crawler that is disallowed but still appears in the ledger is worth knowing about; so is a search engine that is allowed but has stopped fetching.

If a row is wrong — a real person marked as a bot, or a bot you want treated as human — use Report false positive on that row. A reviewer checks it, and the fix applies to every site.

Limits

  • With only the page script, you see bots that execute JavaScript. Crawlers that fetch raw HTML without running scripts never appear. For those, use the server-side SDK or a log import; both are described in the developer docs.
  • Status codes and response times come from server collection, not the page script.
  • AI source in the Sources report is a person who clicked through from an AI product. The ledger counts programs reading pages. The two are related but never equal.

For the full story on reading the ledger, see the AI crawler ledger post.

Was this guide helpful?

Your answer helps us decide which guides to expand next.

Thanks for the feedback. If something is still unclear, tell us what you were looking for.

See your data clearly. Find your growth.

Every click, backed by data. Install one line of code and see your first numbers in a minute.

No credit card · Free plans for analytics and chat · Cookieless analytics