Most AI Crawler Blocks Were Never a Decision
Why robots.txt alone doesn't tell you whether GPTBot or PerplexityBot can read your site - and how to check what your CDN is doing underneath it, for free.
We built the free audit to answer "can AI crawlers read this page." Running it for a while surfaced a narrower, sharper question underneath that one: when a crawler can't read a page, was that actually decided by anyone? Often the answer is no. So we split that question out into its own tool - the AI Crawler Access Checker - and building it made the "nobody decided this" pattern a lot more visible than we expected.
robots.txt is not the whole story
Ask most site owners whether they block AI crawlers and they'll point you at their robots.txt. That's a reasonable place to look, and it's the first thing our checker reads - properly, with wildcard groups, Allow overrides, and longest-match precedence, not a regex skimming for "Disallow." But robots.txt is a request, not an enforcement mechanism. It tells a well-behaved crawler what you'd prefer. It says nothing about what your CDN, WAF, or hosting platform does to that same crawler's actual HTTP request.
That gap is where most of the confusion lives. A site can have a completely open robots.txt and still return a 403 to GPTBot, because something in front of the origin server made that call independently - most commonly a managed bot-protection rule that shipped enabled by default.
The Cloudflare default
Cloudflare's managed rules for AI bots are the single biggest source of this we see. They're a reasonable product to offer - a lot of site owners genuinely want to block AI scrapers - but "reasonable to offer" and "consciously chosen by the person whose site it protects" are different claims. A setting flipped on somewhere in a dashboard, possibly by whoever set up the CDN years ago, possibly as a default nobody touched, isn't the same thing as a decision made about this site's AI visibility strategy.
[TK: stat - % of checked sites where a bot UA got 403 while the browser control request got 200, once we clear 500 runs]
The way to tell the two apart from the outside is the part robots.txt-only tools skip: run a normal browser request against the same URL and compare it to the bot request. If the browser gets through and the bot doesn't, and robots.txt didn't say to block it, that's not a robots.txt decision - it's something downstream making the call. Our checker always runs that control request, and when the pattern matches Cloudflare's fingerprint specifically, it says so and links straight to the toggle.
Not every block is a problem
Once you can see what's actually blocked, the more important question is which crawler it was, because we don't think all fourteen entries in our registry deserve equal alarm.
Split them by purpose. Training crawlers - GPTBot, ClaudeBot, CCBot, Bytespider, and the two robots.txt-only directives, Google-Extended and Applebot-Extended - only affect whether your content trains a future model. Blocking those is a legitimate, defensible call about your content and your terms. We don't moralize about it in the tool, and we're not going to here either.
Search and live-browsing crawlers are a different category entirely: OAI-SearchBot, PerplexityBot, Claude-SearchBot on the search side; ChatGPT-User, Perplexity-User, Claude-User on the live-browsing side. These are the crawlers standing between your page and someone asking ChatGPT or Perplexity a question right now. Blocking one of those isn't a stance on AI training - it's opting your page out of getting cited in an answer someone is looking at today. If that block turns out to be a CDN default nobody chose, it's costing you citations for no reason at all.
Our tool's headline verdict leads with this distinction on purpose: it tells you first whether a search or live-browsing crawler is blocked, because that's the one worth acting on quickly, and only mentions training-crawler blocks as a secondary, neutral note.
What the check actually does
For a URL, it:
- Fetches and properly parses robots.txt.
- Requests the page under all fourteen registered bot user-agents (skipping the two directive-only tokens, which never send a live request - we report those from robots.txt alone).
- Runs one plain-browser control request against the same URL.
- Classifies each bot as able to read the page, blocked by robots.txt, or blocked by the server - and flags when a server block looks like a Cloudflare managed rule specifically.
Same no-signup shape as the free audit: a URL and an email, no credit card, no GetBotRank credits. It's free, it doesn't call a model, and there's nothing to meter -- the email just gets you a copy in your inbox and a shareable link alongside the on-page result.
Try it
getbotrank.com/tools/crawler-access - enter a URL, see exactly which AI crawlers can reach it and why, in a few seconds.