Free tool AI crawler robots.txt checker
Paste the robots.txt from your site and see, crawler by crawler, what it allows. Then tick the crawlers you want to block and copy the rules. Parsed in your browser — no request is made to your site.
- 11 AI crawler user agents
- Shows the rule that decides
- Runs locally
Paste robots.txt, read the verdict
Open https://your-site.com/robots.txt, copy everything, paste it below. The table updates as you type. The default text is an example file you can replace.
Example file preloaded — replace it with your own. Comments and Sitemap lines are ignored.
Allowed
0
Partly restricted
8
Blocked
3
Not declared
0
Crawler by crawler
A crawler follows the group written for it; if there is none, it follows User-agent: *; if there is neither, it may crawl everything.
| Crawler | Purpose | Deciding rule | Status |
|---|---|---|---|
| GPTBotOpenAI | Model training | Disallow: /own User-agent group | Blocked |
| OAI-SearchBotOpenAI | AI search index | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
| ChatGPT-UserOpenAI | Fetches on a user’s request | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
| ClaudeBotAnthropic | Model training | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
| PerplexityBotPerplexity | AI search index | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
| Google-ExtendedGoogle | Model training | Disallow: /own User-agent group | Blocked |
| Applebot-ExtendedApple | Model training | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
| BytespiderByteDance | Model training | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
| CCBotCommon Crawl | Model training | Disallow: /own User-agent group | Blocked |
| meta-externalagentMeta | Model training | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
| AmazonbotAmazon | AI search index | Disallow: /admin/ Disallow: /cartinherits User-agent: * | Partly restricted |
Generate rules
Tick the crawlers you want to keep out. Paste the result at the end of your robots.txt — one group per crawler, nothing else changes.
# AI crawlers — generated with tapcub.com/tools/ai-crawler
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Bytespider
Disallow: /
User-agent: CCBot
Disallow: /robots.txt is a request, not a lock. Well-known crawlers honor it; the ones that do not are what the bot report is for.
The four rules the checker applies
The same logic well-behaved crawlers use, simplified to the cases that matter for AI bots.
-
1
Find the group for the crawler
A group starts with one or more User-agent lines. The crawler matches a group whose token matches its name, case-insensitively. Only that group applies to it.
-
2
Fall back to the wildcard
No group of its own? The crawler follows User-agent: *. The table shows "inherits *" so you know the verdict is indirect.
-
3
Read Disallow and Allow
Disallow: / blocks everything; an empty Disallow allows everything; other paths mean partly restricted. Allow: / overrides a root block.
-
4
No group at all means allowed
A robots.txt that never mentions a crawler and has no * group places no restriction on it. "Not declared" is effectively "allowed".
Not every AI crawler does the same thing
Blocking all of them is a choice, not a default. Training crawlers copy content for model training; search crawlers index it for AI answers that can link back; assistant fetchers load a page because a person asked.
- Training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Bytespider, CCBot, meta-externalagent
- AI search: OAI-SearchBot, PerplexityBot, Amazonbot — blocking these removes you from AI answers
- Assistant fetch: ChatGPT-User — triggered by a user action, closer to a browser than a crawler
- Policies change; check the vendor documentation for the current token names
What a sensible policy looks like
Block training crawlers if you do not want your content in the next model; keep search and assistant agents if you want to be cited. The snippet generator defaults to that split.
- GPTBotBlocked
- ClaudeBotBlocked
- OAI-SearchBotAllowed
- PerplexityBotAllowed
- ChatGPT-UserAllowed
What the bot report adds
robots.txt tells crawlers what you want. The bot report in TapCub shows what they did: requests per crawler per day, which pages, and whether a blocked one kept coming.
AI crawler requests · 7 days
What the default file says
The example robots.txt blocks three training crawlers by name and gives everyone else the wildcard rules, which only restrict /admin/ and /cart.
User-agent: *
Disallow: /admin/
Disallow: /cart
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /GPTBot, Google-Extended, CCBot → Blocked
Each has its own group with Disallow: /. Nothing in the wildcard group applies to them.
ClaudeBot, PerplexityBot, Bytespider… → Partly restricted
No group of their own, so they inherit User-agent: *, which only keeps them out of /admin/ and /cart. Everything else is open.
If you removed the * group → Not declared
With no wildcard group, crawlers that are not named face no restriction at all. That is the state most sites are in.
Change the text in the checker above and watch the statuses follow.
Why a robots.txt does not do what you meant
The file is simple; the matching rules are not obvious.
Expecting * rules to add to a named group
They do not. A crawler with its own group ignores the wildcard entirely — including your Disallow: /admin/.
Wrong token
"ChatGPT" or "OpenAI" are not tokens. The crawler reads GPTBot, OAI-SearchBot or ChatGPT-User. Typos silently match nothing.
Treating robots.txt as enforcement
It is a request. Compliant crawlers honor it; others ignore it. Watch the bot report to see who actually complies.
Blocking search crawlers by accident
Disallowing OAI-SearchBot or PerplexityBot removes you from AI answers that would have linked to you. Decide per crawler.
AI crawler questions
Still have a question?
Chat with the team behind TapCub — we usually reply within a few hours on business days.
Does the checker fetch my site’s robots.txt?
No. It only parses the text you paste, in your browser. Nothing is requested from your domain and nothing is sent to TapCub.
Will blocking AI crawlers hurt my search ranking?
Blocking training crawlers such as GPTBot or Google-Extended does not affect regular search indexing. Blocking AI search crawlers removes you from those AI answers specifically.
How current is the crawler list?
It covers the user agents most often seen in bot reports as of this writing. Vendors add and rename tokens; check their documentation before publishing a policy.
What if a crawler ignores robots.txt?
Then robots.txt cannot help, and you need to see the traffic. TapCub’s bot report lists requests per crawler so you can decide on a server-side block. See bot detection.
Know who is crawling, not just who you asked not to
Install the ≈2 KB script and the bot report starts separating humans, search engines and AI crawlers from day one — free plan included.