Skip to main content

What's the difference between the AI crawlers hitting my site?

AI Search Published September 11, 2026
Short Answer

They do three different jobs. Training crawlers collect content that may be used to train future models. Indexing crawlers build the search index an assistant retrieves from. On-demand fetchers pull a single page in real time because a person just asked a question. They use different user agent tokens and are controlled separately in robots.txt, so blocking one does not block the others — and blocking the wrong one removes you from live answers.

Three jobs with three signatures

You can usually tell them apart from the request pattern alone, without knowing the vendor.

  • Training crawlers. Broad, periodic, indifferent to which pages matter. They walk large portions of the site over days and rarely return to the same URL quickly.
  • Indexing crawlers. Sitemap-driven and more selective. They revisit pages that change and behave a lot like a conventional search crawler.
  • On-demand fetchers. One or two URLs, no crawl of neighbors, arriving in bursts that correlate with the moment a user asked something. These are the ones whose failure you feel immediately.

The failure nobody looks for

The single most common reason a site is absent from AI answers has nothing to do with content. It is that a bot-mitigation rule, a CDN security setting, or a rate limiter is returning a challenge page or an error to these agents. The page looks perfect in a browser. The fetcher gets a block page and moves on.

Check your logs by user agent and look at the status codes, not just the hit counts. A steady stream of denials or challenges from an agent you meant to allow is a configuration problem, and it is the cheapest thing on this list to fix.

Robots.txt controls them independently

Each agent has its own token, and a rule for one has no effect on another. That means a blanket decision is rarely what people actually want: most businesses are comfortable being retrievable live while being more cautious about training collection, and those are separate lines in the file.

Vendor tokens and behavior change. Confirm the current names in each vendor's own documentation rather than copying a robots.txt off a forum post, and re-check it periodically. Also remember robots.txt is a request, not an enforcement mechanism.

What to actually check, in order

Confirm the fetcher reaches you at all. Confirm it gets a 200. Confirm the served HTML contains the substance. Confirm robots.txt does not disallow the path. Only then start editing content.

This is ordinary infrastructure work, and it belongs in the same review as the rest of your site's technical health. If you have a security layer in front of the site, the person who administers it needs to be in the conversation — this is usually an access and configuration question, not a marketing one.

Topics: crawlers · server logs · robots.txt · diagnostics

Have a version of this question about your own business?

The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.

Related Answers

People who read this also asked

Browse the Answer Hub →

AI is easy to access. Making it useful is hard.

Bluefrog makes AI useful by integrating it with the way your business actually works — your software, your calls, your customers, your marketing and your revenue.

Technology development since 1997 · AI integration platforms since 2001