Free AI Crawler Access Checker

Is any AI crawler being blocked outright?

We request your page as GPTBot, ClaudeBot, PerplexityBot, and 11 others -- for real, from our own servers -- then check that against a normal browser request on the same URL.

No signup required

Why this needs more than reading robots.txt

robots.txt tells you what you told crawlers to do. It doesn't tell you what your CDN, WAF, or hosting platform is doing on top of that -- and a large share of AI-crawler blocks turn out to be a managed default nobody at the company actually chose. The only way to find that gap is to actually make the requests and compare them.

How the check works

We read robots.txt properly

Full parsing -- wildcard groups, Allow overrides, longest-match precedence -- not a regex guess. We check every registered AI crawler token, including the two directive-only ones (Google-Extended, Applebot-Extended) that never send a live request.

We request your page as each crawler

One request per bot, using that vendor’s real user-agent string, sent from our own infrastructure -- identified honestly in every request (see /bot). We record the actual HTTP status each one gets back.

We run one control request as a normal browser

If a bot gets a 403 but a browser UA gets 200 on the same URL, that is not a robots.txt decision -- it is usually a CDN or WAF rule, most often Cloudflare’s managed AI-bot rules, which a lot of site owners never consciously turned on.

Frequently asked questions

What counts as "blocked"?

Either robots.txt disallows that crawler’s token for the page, or the live request came back 401, 403, 429, or a CDN/WAF challenge page (most commonly Cloudflare’s). We report which one it was -- they mean different things.

Is blocking an AI crawler always bad?

No. Blocking a training crawler (GPTBot, ClaudeBot, CCBot, Bytespider) is a defensible business decision about whether your content trains future models. Blocking a search or live-browsing crawler (OAI-SearchBot, PerplexityBot, ChatGPT-User) is different -- that is what actually keeps you out of AI answers people are asking right now.

Why do Google-Extended and Applebot-Extended show "directive only"?

Those two are robots.txt tokens, not crawlers that ever send a live HTTP request under that name. There is nothing to probe -- the only signal is whether your robots.txt disallows the token, so that is all we check for them.

Why do you need my email?

It doesn't require a signup, a credit card, or any GetBotRank credits -- just a URL and an email. The result renders on this page immediately after you submit, and we send you a copy plus a shareable link at the same time.

How is this different from just reading my robots.txt myself?

robots.txt only tells you what you told crawlers to do. It doesn’t tell you what your CDN or WAF is doing underneath it -- which is often the actual block. Running a live control request alongside the bot requests is what surfaces that gap.