What an AI Crawler Actually Sees On Your Page
We built a free tool that shows the raw text GPTBot recovers from your page, side by side with what your browser renders - and the gap is usually bigger than people expect.
The free audit caps a page's score hard the moment it detects a client-rendered shell, regardless of how clean everything else is. That single check turned out to explain more bad scores than the other four combined. So we pulled it into its own tool - Crawler View - that does one thing: shows you exactly what an AI crawler recovers from a page, next to what your own browser renders, with nothing in between.
The gap is invisible until you look for it
Open your site in Chrome and it looks fine. Chrome runs JavaScript, fills in the content, and you never see the in-between state. GPTBot, ClaudeBot, and PerplexityBot largely skip that step entirely - they read the HTML response and move on. If your content lives in a <div id="root"> that only fills in after a script runs, what they receive is close to nothing, and there is no way to notice that from a normal browser tab.
Crawler View exists because "trust me, it's server-rendered" and "actually check" produce different answers more often than you'd think.
Why we didn't reach for a headless browser
The obvious way to build a "what does a crawler see" tool is to spin up Puppeteer or Playwright, load the page, and diff the rendered DOM against the raw response. We didn't do that, on purpose. A headless browser answers "what would this page look like if a crawler executed JavaScript" - and the entire point of this tool is that the crawlers we care about don't. Adding one would have made the left panel and the right panel converge on cases where they're supposed to diverge, which is the one thing this comparison can't afford to get wrong.
So the right panel is a single HTTP request, under a real crawler user-agent, with nothing rendering it - the same request GPTBot itself would make. The left panel is your own browser, in a sandboxed iframe, actually running the page. If a site blocks iframe embedding (plenty do, deliberately, for clickjacking protection), that panel can come back blank even though the site is fine - that's a browser security policy, not a crawler-visibility finding, and we say so directly in the UI rather than let it read as a false negative.
What it actually does
- Fetches the page once, under a real crawler user-agent (GPTBot), with JavaScript never executed - no headless browser, because that would defeat the point.
- Runs the raw HTML through the same Readability-based extractor the paid audit pipeline uses, to recover the text a crawler would actually work with.
- Puts that side by side with your own browser rendering the live page in an iframe, so the comparison isn't an abstraction.
- Computes word count, character count, heading structure, and the ratio of script markup to recovered text.
When the recovered text drops under roughly 200 words and the page is dominated by a client-side mount node, that's the headline finding, stated in one sentence at the top - not buried in a metrics table.
[TK: stat - % of checked pages that come back as an empty shell, once we clear 500 runs]
What "fine" and "not fine" actually look like
A server-rendered marketing page, a static blog post, most e-commerce product pages built on conventional stacks - these come back with the full text intact. The right-hand panel reads almost identically to what a human sees, just without the styling.
A single-page app that renders its entire body client-side comes back nearly blank on the right, no matter how polished the left panel looks. That's not a subtle scoring nuance - it's the difference between a crawler having something to cite and having nothing at all.
To make that concrete: picture a typical client-rendered product page. A human visitor sees a hero image, a paragraph or two of copy, a spec table, and a review section - easily a few hundred words. The raw HTML response, before any script runs, is often little more than <div id="root"></div> plus a handful of <script src="/static/js/main.a1b2c3.js"> tags. Run that through the extractor and the recovered text is close to zero words against a script-tag count in the double digits - which is exactly the script-to-text ratio metric flagging what the word count alone might undersell. That's not a hypothetical edge case; it's the default output of several popular frontend frameworks unless someone has deliberately configured server-side rendering.
What this doesn't tell you
Crawler View is a rendering check, not a quality check. A page can pass this - full text, clean structure - and still be poorly written, badly organized, or missing the answer someone actually asked for. That's a separate, harder problem, and it's the one the full multi-model audit (GPT-4o, Claude, Gemini scoring actual content quality) is built to catch. Code can tell you a page is readable. Only a model can tell you if it's any good.
Try it
getbotrank.com/tools/crawler-view - paste a URL and an email, see both sides in a few seconds. No signup, no credit card, no credits spent.