AI Crawler Log Analyzer: See What GPTBot and ClaudeBot Fetch
Drop in your server’s access log and the tool finds the requests from GPTBot, ClaudeBot, PerplexityBot and other AI crawlers as well as Googlebot, separates real crawlers from impostors using the operators’ published IP lists, and shows which pages they fetched and which errors they hit.
Key takeaways
- The tool finds AI crawler and Googlebot requests in your access log and shows which pages they fetched and which status codes they received.
- It separates real crawlers from impostors using the operators’ published IP lists.
- It detects nginx, Apache, Cloudflare Logpush, Vercel and other log formats, and
.gzfiles do not need unpacking. Shopify does not give merchants access logs, so Shopify stores cannot be analysed. - The log is processed line by line in your browser and never uploaded to this site or any third party.
What this tool answers
The AI crawler robots.txt checker tells you who your rules allow. Whether AI crawlers actually come, which pages they fetch and whether your server answers them properly is only visible in your access logs. This tool uses your own log to answer:
- Have the ChatGPT, Claude and Perplexity crawlers visited recently, how often per day, and how does that compare with Googlebot?
- Which URLs do they fetch: product and article pages, or old links that no longer exist?
- Are the status codes they receive healthy? Many 403s usually mean a CDN or firewall is blocking crawlers you meant to allow.
- Of the requests claiming to be GPTBot or Googlebot, how many come from official IPs and how many are scrapers in disguise?
- Do they request robots.txt and llms.txt, and do they request JS and CSS files?
How to use it
- Download an access log from your server, hosting panel, CDN or platform (see the next section). Compressed
.gzfiles do not need to be unpacked. - Drop the file on the dashed box above or click “Choose file”. The format is detected automatically and progress is shown; you can cancel a large file at any time.
- When reading finishes, start with “Findings and fixes”, then expand a row in “Crawler details” to search its URLs.
- Click the export buttons to continue in a spreadsheet.
No log at hand? Click “Load sample log”. The sample is synthetic data built into the tool, with a made-up site and made-up requests, and only demonstrates the report.
Everything runs in your browser. The file is read and counted line by line in a background thread (Web Worker) and is never uploaded to this site or any third party.
How to export an access log
nginx
nginx logs in the combined format by default, usually at /var/log/nginx/access.log on Debian and Ubuntu. Rotated files such as access.log.2.gz can be dropped in as they are. If your access_log directive uses a custom format, make sure it includes $remote_addr, $time_local, $request, $status and $http_user_agent, or crawlers cannot be identified:
1 | access_log /var/log/nginx/access.log combined; |
Apache
Apache must use the combined format; the common format does not record the User-Agent. The log is usually /var/log/apache2/access.log on Debian and /var/log/httpd/access_log on CentOS and RHEL:
1 | CustomLog ${APACHE_LOG_DIR}/access.log combined |
aaPanel / BT Panel
BT Panel (宝塔面板, known internationally as aaPanel) stores site access logs in /www/wwwlogs/, usually named yourdomain.log. Open that folder in the panel’s file manager and download the file. The panel uses the nginx or Apache combined format by default.
Cloudflare
If the site sits behind Cloudflare, your origin logs record Cloudflare’s edge IPs, and the tool will warn that “the log may record your CDN or proxy IP”. You have two options:
- Restore the visitor IP on the origin with nginx’s realip module (
real_ip_header CF-Connecting-IP, with every range from Cloudflare’s IP list inset_real_ip_from). New log lines can then be verified. - Export HTTP request logs with Cloudflare Logpush. Logpush writes one JSON record per line, and the tool reads the
ClientIP,ClientRequestUserAgent,ClientRequestURI,EdgeResponseStatusandEdgeStartTimestampfields. Check Cloudflare’s current documentation for which plans include Logpush.
Vercel
Vercel Log Drains push request logs as JSON to an endpoint you choose, with the request data in a proxy object (clientIp, userAgent, path, statusCode, timestamp); the tool reads it directly. The dashboard’s Logs view can also export results as CSV or JSON, and the tool matches the IP, user agent, path, status and time columns by name.
Shopify
Shopify does not give merchants access to server access logs, so this tool cannot analyze crawler traffic for a Shopify store. Shopify merchants can still check their robots.txt rules with the AI crawler robots.txt checker. Hosted builders such as Wix and Squarespace are in the same position: without raw logs there is nothing to analyze.
Other sources
cPanel hosts offer Apache combined logs under “Raw Access”. IIS and Amazon CloudFront use the W3C extended format (with a #Fields: header line), which the tool reads by field name. CSV exports from other platforms usually work as long as the first row holds column names and one of them is the user agent.
How to read the report
Three kinds of AI crawler
| Kind | Crawlers | What it means for you |
|---|---|---|
| AI search | OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot | Decide whether pages can appear, with a citation, in ChatGPT, Claude and Perplexity answers |
| User-triggered | ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User | Live visits when a user asks the assistant to open a page; someone mentioned or looked at your page in an AI chat |
| AI training | GPTBot, ClaudeBot, CCBot, meta-externalagent, Bytespider, Amazonbot | Collect training data; no direct citations or clicks |
Classic search crawlers (Googlebot, Bingbot, Applebot, Baiduspider, YandexBot, DuckDuckBot) are counted alongside as a baseline. Google’s AI Overviews and AI Mode use content crawled by Googlebot, so Googlebot data matters for AI search too.
IP verification: real, likely fake, cannot verify
A user agent is whatever the client writes, and anyone can write GPTBot. OpenAI, Anthropic, Google, Microsoft, Apple and Perplexity publish the IP ranges of their crawlers, and the tool checks each request’s source IP against them:
- Real: the IP is on the official list, so the request is confirmed to come from that company (evidence grade A).
- Likely fake: the user agent claims to be a crawler but the IP is not on the official list. The report lists these IPs and generates an nginx config that identifies spoofers by the official ranges. Because the official lists change and blocking straight away can lock out real crawlers, the config only writes suspected spoofers to a separate log by default; enable the 403 once a few days of that log show no false matches, and run
nginx -tafter editing. - Cannot verify: the operator publishes no IP list (CCBot, Bytespider and others), or the tool has no data for that list right now.
The lists are fetched from the official URLs when the site is built; if a fetch fails, a snapshot stored in the repository is used. Expand “IP verification uses the official published IP lists” on the tool to see each list’s source and date.
For practitioners: list URLs and verification details
- OpenAI:
openai.com/gptbot.json,openai.com/searchbot.json,openai.com/chatgpt-user.json, each verifying its own crawler. - Anthropic:
claude.com/crawling/bots.json; the tool uses this single list for ClaudeBot, Claude-SearchBot and Claude-User. - Perplexity:
www.perplexity.com/perplexitybot.jsonandwww.perplexity.com/perplexity-user.json. - Google:
developers.google.com/static/crawling/ipranges/common-crawlers.json, covering Googlebot and the other common crawlers. The oldgooglebot.jsonURL now redirects there. - Microsoft:
www.bing.com/toolbox/bingbot.json; Apple:search.developer.apple.com/applebot.json. - All lists share one format:
ipv4Prefix/ipv6Prefixentries in aprefixesarray. IPv6 and IPv4-mapped addresses are supported. - Google and Bing also support reverse DNS verification. Browsers cannot do reverse DNS lookups, so the tool uses the IP lists only.
- When the logged IP is a private address, the tool uses the first public IP in
X-Forwarded-Forinstead. It ignores that header when the source IP is public, because clients can forge it.
Status codes
The report counts 2xx, 3xx, 4xx and 5xx responses per crawler and lists the URLs with the most errors. Google’s HTTP status code documentation says 5xx and 429 make its crawlers slow down, while 4xx codes other than 429 do not affect crawl rate (A). AI companies have published nothing comparable, so for AI crawlers the tool only suggests what to check (D).
Typical fixes: 403s usually come from WAF or CDN bot blocking, so check whether the rules block crawlers you want to allow; 404s usually come from old links, so 301-redirect old URLs that still have links; for 5xx and 429, check server load and rate limits, and return 429 with Retry-After when rate limiting.
JS and CSS requests
A 2024 study by Vercel of crawler traffic on its network found that the major AI crawlers do not execute JavaScript, and that some download JS files without running them (D). Google documents that Googlebot renders JavaScript (A). The report therefore puts the JS / CSS requests of AI crawlers next to Googlebot’s. If your body text, prices or structured data are inserted by client-side scripts, AI crawlers probably do not see them, and you should switch to server-side rendering or static generation.
robots.txt and llms.txt
Google generally caches robots.txt for up to 24 hours, so a short log without robots.txt requests is normal. If the log spans several days and a crawler keeps fetching pages without ever requesting robots.txt, it is worth a closer look. User-triggered fetchers (ChatGPT-User, Perplexity-User and others) are documented by their operators as possibly not following robots.txt.
llms.txt is a community proposal, and the major AI companies have not publicly confirmed that they use it in answers (D). The report only counts which crawlers requested /llms.txt and /llms-full.txt; fetching the file does not mean its content is used.
Other bot-like user agents
The bottom of the report lists user agents that are not on the tool’s list but contain bot, spider, crawler and similar words, plus scripts such as python-requests and curl. Most are SEO tools and monitoring services; some may be new AI crawlers that are not on the list yet. When an unfamiliar name sends many requests, look up its documentation before deciding whether to handle it in robots.txt or your firewall.
Evidence grades
Every finding carries an evidence grade, the same scale used across this site’s tools:
- A Confirmed: documentation or public statements from Google, OpenAI or another operator.
- B Documented: court filings, leaked documents and similar; they exist, but how the information is used is unknown.
- C Patented: patents or papers; technically possible, not proof of deployment.
- D Speculative: third-party studies, correlation analyses and industry experience.
The counts in your log are facts about your own site; the grade describes how reliable the tool’s reading of those counts is. There is no overall score.
FAQ
Is my log uploaded?
No. The browser reads the file and a background thread in the browser does the counting; the results only appear on this page. This site’s server only serves the page, the scripts and the IP lists, and never receives log content.
Will a multi-GB log freeze the browser?
No. The file is streamed and read line by line, and only the counts are kept in memory, not the file. Each crawler keeps up to 20,000 distinct URLs; beyond that, new URLs are only added to the total. You can cancel at any time while it reads.
Why do Perplexity or Bingbot show “cannot verify”?
Their IP lists have to be fetched from the official URLs when the site is built. If a build could not fetch a list and the repository snapshot has no data for it either, the tool can only count by user agent. Expand the IP list details on the tool to see the status of each list.
Is every “likely fake” request a fake crawler?
Almost always. There are two exceptions: the operator added IPs that the list does not include yet, or the log records a CDN or load balancer IP. The tool warns about the second case separately. Before blocking anything, make sure the IPs in your log are the visitors’ real IPs.
Why don’t I see Google-Extended in my logs?
Google-Extended and Applebot-Extended are robots.txt control tokens without their own user agent. Google documents that Google-Extended has no separate HTTP request user agent; crawling is done by Google’s other crawlers. Whether they take effect can only be checked in robots.txt.
Which log formats are supported?
nginx and Apache combined logs (including vhost_combined and the X-Forwarded-For field at the end of nginx’s main format), JSON with one record per line (Cloudflare Logpush, Vercel Log Drains, Google Cloud and others), whole JSON arrays up to 150 MB, CSV / TSV with a header row, and the W3C extended format (IIS, CloudFront). The common format has no user agent and cannot be used.
Are more AI crawler visits always better?
No. Heavy traffic from AI training crawlers only means your content is being collected for training; it brings no citations or clicks directly. For GEO, the more useful signals are the requests from AI search and user-triggered crawlers, and whether the pages they fetch are the ones you want cited.
Further reading
- AI crawler robots.txt checker: see which AI crawlers your robots.txt allows.
- robots.txt and Meta Robots: Controlling Search and AI Crawlers (Chinese)
- More free SEO and GEO tools