TapCub is live — analytics, insights and live chat, free on one platform

Free tool AI crawler robots.txt checker

Paste the robots.txt from your site and see, crawler by crawler, what it allows. Then tick the crawlers you want to block and copy the rules. Parsed in your browser — no request is made to your site.

  • 11 AI crawler user agents
  • Shows the rule that decides
  • Runs locally
Checker

Paste robots.txt, read the verdict

Open https://your-site.com/robots.txt, copy everything, paste it below. The table updates as you type. The default text is an example file you can replace.

Example file preloaded — replace it with your own. Comments and Sitemap lines are ignored.

4 User-agent group(s) parsed.
Runs entirely in your browser. Nothing you type here is uploaded or stored by TapCub.

Allowed

0

Partly restricted

8

Blocked

3

Not declared

0

Crawler by crawler

A crawler follows the group written for it; if there is none, it follows User-agent: *; if there is neither, it may crawl everything.

AI crawler status according to the pasted robots.txt
CrawlerPurposeDeciding ruleStatus
GPTBotOpenAIModel trainingDisallow: /own User-agent groupBlocked
OAI-SearchBotOpenAIAI search indexDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted
ChatGPT-UserOpenAIFetches on a user’s requestDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted
ClaudeBotAnthropicModel trainingDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted
PerplexityBotPerplexityAI search indexDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted
Google-ExtendedGoogleModel trainingDisallow: /own User-agent groupBlocked
Applebot-ExtendedAppleModel trainingDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted
BytespiderByteDanceModel trainingDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted
CCBotCommon CrawlModel trainingDisallow: /own User-agent groupBlocked
meta-externalagentMetaModel trainingDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted
AmazonbotAmazonAI search indexDisallow: /admin/ Disallow: /cartinherits User-agent: *Partly restricted

Generate rules

Tick the crawlers you want to keep out. Paste the result at the end of your robots.txt — one group per crawler, nothing else changes.

robots.txt
# AI crawlers — generated with tapcub.com/tools/ai-crawler

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: CCBot
Disallow: /

robots.txt is a request, not a lock. Well-known crawlers honor it; the ones that do not are what the bot report is for.

How it reads the file

The four rules the checker applies

The same logic well-behaved crawlers use, simplified to the cases that matter for AI bots.

  1. 1

    Find the group for the crawler

    A group starts with one or more User-agent lines. The crawler matches a group whose token matches its name, case-insensitively. Only that group applies to it.

  2. 2

    Fall back to the wildcard

    No group of its own? The crawler follows User-agent: *. The table shows "inherits *" so you know the verdict is indirect.

  3. 3

    Read Disallow and Allow

    Disallow: / blocks everything; an empty Disallow allows everything; other paths mean partly restricted. Allow: / overrides a root block.

  4. 4

    No group at all means allowed

    A robots.txt that never mentions a crawler and has no * group places no restriction on it. "Not declared" is effectively "allowed".

Who is who

Not every AI crawler does the same thing

Blocking all of them is a choice, not a default. Training crawlers copy content for model training; search crawlers index it for AI answers that can link back; assistant fetchers load a page because a person asked.

  • Training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Bytespider, CCBot, meta-externalagent
  • AI search: OAI-SearchBot, PerplexityBot, Amazonbot — blocking these removes you from AI answers
  • Assistant fetch: ChatGPT-User — triggered by a user action, closer to a browser than a crawler
  • Policies change; check the vendor documentation for the current token names

What a sensible policy looks like

Block training crawlers if you do not want your content in the next model; keep search and assistant agents if you want to be cited. The snippet generator defaults to that split.

  • GPTBotBlocked
  • ClaudeBotBlocked
  • OAI-SearchBotAllowed
  • PerplexityBotAllowed
  • ChatGPT-UserAllowed
Sample data

What the bot report adds

robots.txt tells crawlers what you want. The bot report in TapCub shows what they did: requests per crawler per day, which pages, and whether a blocked one kept coming.

AI crawler requests · 7 days

GPTBot1,284
ClaudeBot812
PerplexityBot506
Bytespider377
Sample data
Worked example

What the default file says

The example robots.txt blocks three training crawlers by name and gives everyone else the wildcard rules, which only restrict /admin/ and /cart.

robots.txt
User-agent: *
Disallow: /admin/
Disallow: /cart

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

GPTBot, Google-Extended, CCBot → Blocked

Each has its own group with Disallow: /. Nothing in the wildcard group applies to them.

ClaudeBot, PerplexityBot, Bytespider… → Partly restricted

No group of their own, so they inherit User-agent: *, which only keeps them out of /admin/ and /cart. Everything else is open.

If you removed the * group → Not declared

With no wildcard group, crawlers that are not named face no restriction at all. That is the state most sites are in.

Change the text in the checker above and watch the statuses follow.

Common mistakes

Why a robots.txt does not do what you meant

The file is simple; the matching rules are not obvious.

Expecting * rules to add to a named group

They do not. A crawler with its own group ignores the wildcard entirely — including your Disallow: /admin/.

Wrong token

"ChatGPT" or "OpenAI" are not tokens. The crawler reads GPTBot, OAI-SearchBot or ChatGPT-User. Typos silently match nothing.

Treating robots.txt as enforcement

It is a request. Compliant crawlers honor it; others ignore it. Watch the bot report to see who actually complies.

Blocking search crawlers by accident

Disallowing OAI-SearchBot or PerplexityBot removes you from AI answers that would have linked to you. Decide per crawler.

FAQ

AI crawler questions

Still have a question?

Chat with the team behind TapCub — we usually reply within a few hours on business days.

Does the checker fetch my site’s robots.txt?

No. It only parses the text you paste, in your browser. Nothing is requested from your domain and nothing is sent to TapCub.

Will blocking AI crawlers hurt my search ranking?

Blocking training crawlers such as GPTBot or Google-Extended does not affect regular search indexing. Blocking AI search crawlers removes you from those AI answers specifically.

How current is the crawler list?

It covers the user agents most often seen in bot reports as of this writing. Vendors add and rename tokens; check their documentation before publishing a policy.

What if a crawler ignores robots.txt?

Then robots.txt cannot help, and you need to see the traffic. TapCub’s bot report lists requests per crawler so you can decide on a server-side block. See bot detection.

Next step

Know who is crawling, not just who you asked not to

Install the ≈2 KB script and the bot report starts separating humans, search engines and AI crawlers from day one — free plan included.

See your data clearly. Find your growth.

Every click, backed by data. Install one line of code and see your first numbers in a minute.

No credit card · Free plans for analytics and chat · Cookieless analytics