More in Website analytics
Still have a question?
Could not find it? A person reads every message, usually within a few hours on business days.
Chat with usWhat happens by default
New sites start in recognize and count separately mode. A request recognized as a bot goes to the Bots report and is left out of visitors, sessions, pageviews and every human report. It also does not consume your monthly event quota.
You can change the mode under Web analyticsConfigurationBots: drop discards bot hits entirely, and do not recognize mixes them into human numbers. The last one is kept for comparisons with older tools and is not recommended.
Five detection layers, 197+ rules
Detection runs on every hit before it is stored. A hit is marked as a bot when enough layers agree:
- User agent. Known crawler strings for search engines, AI companies, monitoring services and SEO tools — the largest share of the rule set.
- Request headers. Missing or inconsistent headers that real browsers always send.
- Client signals. Whether the script could run, screen and timing signals that headless browsers get wrong.
- Network (ASN). Ranges that belong to cloud providers and known crawler operators.
- Frequency. Request rates no person produces, such as hundreds of pages per minute from one source.
The rule set is maintained centrally and currently holds 197+ rules; you do not update anything. Each bot row shows which layers matched and a trust level — verified, suspected or unknown. Unknown means we could not confirm the operator, not that the traffic is fake.
Read the Bots report
Open Web analyticsBots. There are three tabs:
- AI ledger — per AI operator: fetches for training and for answer engines, how many human visitors that operator sent back, and the ratio between the two. A high ratio means many reads and few people.
- AI visibility — the same data from the other direction: visits sent back divided by fetches, so you can see which assistants actually cite you.
- All bots — name, operator, purpose (search, AI training, AI answers, monitoring), requests, pages, most common status code, first and last seen. Search engines also get a small health table with success rate, 404s and 5xx.
Compare with robots.txt and llms.txt
The Crawl health view puts your robots.txt and llms.txt rules next to what crawlers actually fetched. A crawler that is disallowed but still appears in the ledger is worth knowing about; so is a search engine that is allowed but has stopped fetching.
If a row is wrong — a real person marked as a bot, or a bot you want treated as human — use Report false positive on that row. A reviewer checks it, and the fix applies to every site.
Limits
- With only the page script, you see bots that execute JavaScript. Crawlers that fetch raw HTML without running scripts never appear. For those, use the server-side SDK or a log import; both are described in the developer docs.
- Status codes and response times come from server collection, not the page script.
- AI source in the Sources report is a person who clicked through from an AI product. The ledger counts programs reading pages. The two are related but never equal.
For the full story on reading the ledger, see the AI crawler ledger post.