AI Crawler Checker: Can GPTBot and ClaudeBot Crawl Your Site?

By Linus Li · Updated · Free, runs in your browser

Rate this tool1 rating so far

Enter a domain or paste your site’s robots.txt and the tool works out, following RFC 9309, whether each AI and search crawler may fetch the path you enter, and which line decided it.

Why AI search and AI training are separate

Many sites block AI crawlers in robots.txt to keep their content out of model training. AI companies, however, usually run different crawlers for different jobs:

  • AI search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot and others) decide whether your pages can appear, with a citation, in ChatGPT, Claude and Perplexity answers.
  • AI training crawlers (GPTBot, ClaudeBot, CCBot and others) collect content to train models and do not send citations or clicks directly.
  • Google-Extended and Applebot-Extended are control tokens, not separate crawlers. Blocking Google-Extended does not affect Google Search indexing or AI Overviews, which are governed by Googlebot.

If you want AI search to cite you without providing training data, allow the search crawlers and block only the training crawlers. “Load example” shows such a configuration.

Matching rules

The tool follows RFC 9309 and Google’s implementation:

  1. A crawler uses the group whose User-agent matches its name (case-insensitive), or the User-agent: * group if none does. When several groups match, their rules are combined.
  2. Within the group, the rule with the longest matching path wins; when Allow and Disallow are equally long, Allow wins.
  3. * matches any characters and $ marks the end of the path. An empty Disallow: blocks nothing.
  4. /robots.txt itself is always allowed.
For practitioners: details that trip people up
  • Once a crawler has its own group, it ignores the User-agent: * group. If you add Disallow: /private/ for GPTBot alone, the Disallow: /cart in the * group no longer applies to GPTBot.
  • Consecutive User-agent lines form one group and share the rules below them.
  • Google ignores Crawl-delay; Bing and Yandex support it.
  • User-initiated fetchers such as ChatGPT-User and Perplexity-User are documented by their operators as possibly not following robots.txt.
  • robots.txt controls crawling, not indexing. A blocked URL with external links can still appear in results without a snippet.

FAQ

The tool says allowed, so why can’t AI reach my site?

robots.txt is only a declaration. A CDN “block AI bots” switch (Cloudflare has one), WAF rules or user-agent blocking on the server reject requests regardless of robots.txt. Most AI crawlers also do not run JavaScript, so client-side-only content looks empty to them.

How does fetching a domain work?

The browser’s same-origin policy stops scripts on one site from reading another, so our server requests https://domain/robots.txt for you and returns only that file. Fetching is rate limited (30 per IP per hour), and the same site checked again within 10 minutes is served from cache. If a site serves different content to servers and browsers, you can still open its robots.txt yourself and paste it.

Where do the crawler names come from?

From each company’s published crawler documentation. AI companies add and rename crawlers from time to time; the list is updated accordingly, and the update date is shown at the top of the page.

Further reading

Scan with WeChat to follow my official account (in Chinese)

Scan with WeChat to follow my official account (in Chinese)