AI Crawlability Checker: What GPTBot and ClaudeBot See
Enter a URL and the tool requests it once each as a normal browser, Googlebot and five AI crawlers, then reports on four layers: can they get in, can they read it, can it be used in answers, and do the dates agree. Each finding comes with fix code.
Key takeaways
- The tool requests the page as a normal browser, Googlebot and five AI crawlers, and reports on four layers: can they get in, can they read it, can it be used in answers, and do the dates agree.
- Requests come from a Vercel data centre and only imitate the user agent, so a block is reported as suspected and needs confirming in your server or CDN logs.
- The tool parses the raw HTML without running JavaScript. The main AI crawlers do not run JavaScript either, so content built by JavaScript may be invisible to them.
- Hard checks are pass or fail, everything else is a suggestion with an evidence level, and there is no overall score.
What this tool checks
Plenty of sites allow AI crawlers in robots.txt and still never get cited by ChatGPT, Claude or Perplexity. There are three common reasons: a CDN or firewall blocks the crawler by user agent; the content is built by JavaScript, so the HTML the crawler receives is empty; or the page carries directives such as noindex or nosnippet. A robots.txt checker only sees a small part of the first problem. This tool requests the page itself and checks all of them.
The seven user agents:
| User agent | Operator | Purpose |
|---|---|---|
| Chrome (baseline) | — | Control, to tell whether a block depends on the user agent |
| Googlebot | Google Search, AI Overviews and AI Mode | |
| OAI-SearchBot | OpenAI | ChatGPT search results |
| Claude-SearchBot | Anthropic | Claude search results |
| PerplexityBot | Perplexity | Perplexity search index |
| GPTBot | OpenAI | Model training |
| ClaudeBot | Anthropic | Model training |
The full user-agent strings and their sources are listed under “User-agent strings sent, with sources” in the results.
How to use it
- Enter the URL. A long article or product page is more telling than a home page, which is mostly navigation.
- Keep “Also read robots.txt” ticked. The tool applies RFC 9309 to decide whether each crawler may fetch the path, with the same rules as the AI Crawler robots.txt Checker.
- Click “Check”. Our server sends the seven requests; this usually takes 3–15 seconds.
- Start with the hard checks at the top and fix any failures first. Each finding has a “How to fix” section with code you can copy.
- If you suspect JavaScript rendering, open the page in your browser, select all, copy, and paste it into “Compare with the rendered text”. The tool estimates how much of it is already in the raw HTML.
You can copy the result as a Markdown report for your developers, or copy the raw JSON. There is no overall score: hard checks are pass or fail, and everything else is a suggestion with an evidence level.
Layer 1: can they get in
The tool compares the status code, response size and a hash of the body for the browser and each crawler. If the browser gets a 200 while GPTBot gets a 403, or gets Cloudflare’s “Just a moment” challenge page, the block depends on the user agent and the tool marks it as suspected blocking.
It can only say “suspected” because the tool requests from a Vercel data centre and only imitates the user agent. Its IP is not in the ranges OpenAI, Anthropic, Perplexity or Google publish for their crawlers. Many firewalls check both, which can skew the result either way:
- The spoofed request is blocked while the real crawler, coming from an official range, is allowed.
- The spoofed request gets through while the real crawler is blocked by an IP or region rule.
To confirm, filter your server or CDN logs by the official IP ranges and look at the real crawlers’ status codes. The signals the tool looks for include Cloudflare’s cf-mitigated: challenge header (the Cloudflare docs say every challenge page carries it), the names of bot-management cookies such as __cf_bm, and the marker text of block pages from DataDome, PerimeterX, Imperva, AWS WAF, Akamai and others.
Since July 2025 Cloudflare blocks AI crawlers by default on newly added domains. If your site joined Cloudflare after that and you have not changed the setting, GPTBot, ClaudeBot and the rest are probably blocked.
For experts: why training and search crawlers are judged differently
OpenAI and Anthropic both separate training from search. Blocking GPTBot or ClaudeBot only affects model training; blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot directly affects citations in AI search. So a blocked search crawler is a failed hard check, while a blocked training crawler is only a suggestion, since many sites opt out of training on purpose.
Google works differently: AI Overviews and AI Mode are controlled by Googlebot. Google-Extended only controls training and grounding in other systems such as Gemini, so blocking it does not affect AI Overviews.
Layer 2: can they read it
The server parses the raw HTML without running JavaScript. It counts visible words, the heading structure and JSON-LD types, and shows up to the first 300 characters of the text. It also looks for client-side rendering markers: an empty <div id="root"></div>, __next or __NUXT__ mount points, and <noscript> notices such as “You need to enable JavaScript”.
This layer matters because the major AI crawlers do not run JavaScript. In 2024 Vercel and MERJ analysed hundreds of millions of crawler requests and found that GPTBot, ClaudeBot, PerplexityBot and others download JavaScript files but do not execute them (Vercel’s write-up, evidence level B). Googlebot does render JavaScript (Google documentation, level A), so a page can look fine in Google Search and still be unreadable to ChatGPT. One or two sources claim OAI-SearchBot has started rendering JavaScript. That is not officially confirmed, so the tool assumes it does not.
The fix is to put the key content in the HTML the server returns: Server Components or getStaticProps / getServerSideProps in Next.js, SSR or pre-generated pages in Nuxt, and for a plain React or Vue single-page app, an SSR / SSG framework or prerendering for crawlers.
Layer 3: can it be used in answers
Once a crawler can read the page, directives on the page still decide whether the content can appear in AI answers. Google’s AI features documentation says that to appear in AI Overviews or AI Mode a page must be indexed and eligible to be shown with a snippet. The tool therefore applies these rules (level A):
noindex(meta robots, meta googlebot or theX-Robots-Tagheader): fail.nosnippetormax-snippet:0: fail, because Google will not use the content in AI features.- A short limit such as
max-snippet:50: suggest relaxing it, because Google applies it to AI features too. data-nosnippet: the count is shown so you can confirm the excluded parts are not the main content.noaiandnoimageai: non-standard. Google’s robots meta documentation does not list them and OpenAI and Anthropic do not document support, so they have no effect.
Bing said in 2023 that pages with noarchive are not included in Bing Chat (now Copilot) answers, and pages with nocache show only the URL, title and snippet. The tool flags both.
Layer 4: freshness and consistency
The tool reads JSON-LD datePublished / dateModified, meta tags such as article:modified_time, <time datetime>, dates in the visible text and the HTTP Last-Modified header, and checks whether they agree. Typical problems are templates that swap the two date fields, a structured-data modified date that differs from the one on the page, and dates in the future. Google’s byline date guidance recommends that the visible date and structured data match.
The idea that AI search prefers fresh content comes mainly from industry observation, with no official statement, so the tool labels it level D. Consistent dates matter because they stop systems from misreading when a page was written.
FAQ
The tool says a crawler is blocked, but my logs show GPTBot getting through. Which is right?
Your logs. The tool’s requests come from a data-centre IP, and a firewall may treat them as impostors. If requests from OpenAI’s official IP ranges get a 200 in your logs, the real crawler is not blocked and the tool’s result is a false positive. If the tool shows a normal page while the real crawler gets a 403 in your logs, the rule blocks by IP or bot score.
The browser row shows a 403 too. What does that mean?
The block is not based on the user agent alone. It may be based on data-centre IPs, region or bot score, or the site may be failing. The per-crawler comparison is of limited use in that case; check your server logs.
Why not show the full page the AI crawler received?
So that this site cannot be used as a general-purpose proxy crawler. The endpoint returns only status codes, selected headers and counts, plus at most the first 300 characters of text. Each IP can run 10 checks an hour, each target site has an hourly request cap, and repeat checks of a URL within 10 minutes come from the cache.
My raw HTML has little text but Google indexes the page fine. Do I need to change anything?
Yes, if you want ChatGPT, Claude or Perplexity to cite it. Googlebot renders JavaScript and AI crawlers usually do not, so normal Google indexing says nothing about what AI crawlers can read.
I want to opt out of AI training but stay in AI search. How?
Split the rules in robots.txt: allow OAI-SearchBot, Claude-SearchBot and PerplexityBot, and disallow GPTBot, ClaudeBot and Google-Extended. CDN “block AI bots” switches usually do not distinguish training from search, so allow the search crawlers there separately. “How to fix” in the tool has examples for robots.txt and Cloudflare rules.
How is the rendered-text percentage calculated?
The number of non-space characters in the raw HTML divided by the number in the pasted text, capped at 100%. Navigation, collapsed content and cookie banners all affect it, so treat it as an estimate. Below 50% means most of the content only appears with JavaScript.
Are results stored?
The server does not store page content. The statistics for a URL stay in the cache for 10 minutes, for rate limiting and to avoid requesting the target site again. Rendered text you paste is processed in your browser and never uploaded.
Related tools
- AI Crawler robots.txt Checker: see which robots.txt rule applies to each of 20+ crawlers.
- Schema Visualizer: find JSON-LD syntax errors and see how entities connect.
- SERP Snippet Preview: check in pixels whether titles and descriptions get truncated.
- All free SEO and GEO tools