← Blog

20 July 2026

AI tools that don't crawl are guessing

There's a fast way to check how AI models talk about your brand, and it's tempting: open ChatGPT, paste in your URL, and ask "how's my SEO?" You'll get an answer in seconds, formatted nicely, with a handful of specific-sounding recommendations. It feels like an audit.

It isn't one. Here's why.

What actually happens when you ask an AI to "look at your website"

When you paste a URL into a chat interface and ask for feedback, the model doesn't crawl your site the way a search engine does. Depending on the tool, it either fetches a single page, skims whatever a general web-search plugin happens to surface, or, worse, draws on stale training data that has no idea what your site looks like today. There's no sitemap discovery, no systematic pass through every page, no consistent starting point. Ask the same question tomorrow and you'll likely get a different answer, because the underlying "audit" was never a fixed, repeatable process to begin with: it was a best guess, generated fresh each time.

That's fine for a quick gut check. It's not a foundation for deciding what to fix on your site this week.

What a real crawl actually involves

A proper site audit starts the same way a search engine does: by finding out what pages exist in the first place. That means discovering your sitemap (sitemap.xml, sitemap_index.xml, or wp-sitemap.xml, whichever your site publishes) and then fetching and indexing every page listed in it. Not a sample. Not "the homepage and a couple of links that looked important." Every page, so that the picture of your site is complete before any judgment gets made about it.

This matters more than it sounds like it should. A page with a broken canonical tag, a duplicate meta description, or missing schema markup doesn't show up if nobody looked at it. An AI model asked to "review your website" from a skim of three pages will simply never mention the problem on page forty, because it never saw page forty. A systematic crawl doesn't have that blind spot.

Then, and only then, the AI testing

Once there's a real, complete inventory of a site, testing how AI models talk about the brand becomes a much more precise exercise. Instead of asking a model to freeform-review a URL, the right approach is to run the actual queries a buyer would type, such as "best project management tool for small agencies", against Claude, ChatGPT, Gemini and Perplexity, and then cross-reference what those models say against what's genuinely on the crawled pages.

That cross-reference is the part that turns a guess into a fact. If a competitor gets recommended instead of you, the question isn't "why does the AI like them more": it's "what's on their page that answers this query better than anything on ours, specifically." That's an answerable, fixable question, because it's grounded in real content on both sides, not in an AI's vague impression.

Why this is the harder way to build it, and the only honest one

It would have been much faster to skip the crawl step entirely and just prompt an AI model to summarise a site on demand. It's also close to worthless as an audit: the same shallow skim every tool that takes the shortcut ends up doing, dressed up with a confident tone.

The alternative is slower to build and slower to run, because a full crawl of a real site takes real time. But it's the difference between a recommendation that's a hit-or-miss guess and one that's grounded in the actual crux of what's on your pages: consistent and repeatable, so you can check it again next week and trust that a change in the number reflects a real change on your site, not noise from a different random sample.

If a tool can't tell you how it arrived at a recommendation, it's worth asking whether it actually looked.

See what AI is actually saying about your site.