An AI crawler is a program run by an AI company that fetches pages from your website automatically, the same way Googlebot has fetched pages for Google Search since the 1990s. What changed in the last few years is why they fetch: alongside search-index crawlers, AI companies now also run crawlers that collect training data for language models, and crawlers that fetch one specific page only when a user of their assistant asks it to. Those are three different jobs with three different consequences, and conflating them leads to confused decisions about what to block.

Three kinds of AI crawler, not one

Every major AI company that publishes a documented crawler falls into one or more of these categories:

A single company can run crawlers in more than one category. Anthropic, for example, operates ClaudeBot for training, Claude-SearchBot for its search index, and Claude-User for on-demand fetches, each with its own robots.txt user-agent, so you can allow one and block another.

The well-known AI crawler user-agents, as of late 2026

User-agentCompanyKindWhat it is for
GPTBotOpenAITrainingCollects content to train OpenAI's models.
OAI-SearchBotOpenAISearch indexIndexes pages for ChatGPT's search and citation features, not training.
ChatGPT-UserOpenAIOn-demand fetchFetches a page only when a ChatGPT user asks it to browse that URL.
ClaudeBotAnthropicTrainingCollects content to train Claude models.
Claude-SearchBotAnthropicSearch indexIndexes pages for Claude's search and citation features.
Claude-UserAnthropicOn-demand fetchFetches a page only when a Claude user asks it to browse that URL.
anthropic-aiAnthropicTraining (legacy)An older Anthropic training crawler name, largely superseded by ClaudeBot.
PerplexityBotPerplexitySearch indexIndexes pages for Perplexity's search and answer features.
Perplexity-UserPerplexityOn-demand fetchFetches a page only when a Perplexity user asks it to browse that URL.
Google-ExtendedGoogleTraining opt-outNot a separate crawler; a token controlling whether already-crawled content trains Gemini and powers certain AI features.
CCBotCommon CrawlTraining (general)A general web archive used as training data by many AI labs, not one company.
BytespiderByteDanceTrainingCollects content to train ByteDance's AI models.

New crawlers appear and old ones get renamed. Before treating a name as current, check the company's own documentation. Common Crawl, for one, warns directly that it is aware of crawlers falsely identifying themselves as CCBot, which is a reminder that a user-agent string in a request header is a claim, not a guarantee; a site that genuinely needs to verify a crawler's identity should check the IP ranges or reverse DNS the crawler operator publishes, not just the string.

What robots.txt actually does

robots.txt is a plain text file at the root of your domain (for example, https://example.com/robots.txt) that lists rules per user-agent: which paths a named crawler may or may not fetch. It is part of the Robots Exclusion Protocol, a convention that predates AI crawlers by decades and was formalized as an actual internet standard, RFC 9309, in 2022.

The important limitation: robots.txt is voluntary. Nothing about the file technically prevents a request from reaching your server. It works because major crawler operators choose to check the file before fetching and choose to respect what it says. Anthropic, for instance, states that all three of its Claude bots honor robots.txt, including Claude-User, and notes that this is notable because other vendors' user-initiated fetchers may bypass or generally ignore robots.txt for that same category of request. In other words, even among companies that generally respect the file, the on-demand fetch category is treated inconsistently. robots.txt is not a security control and should never be relied on to keep something actually private; use authentication or simply do not publish sensitive content if that is the goal.

A basic example

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

Sitemap: https://www.example.com/sitemap.xml

This example opts a site out of OpenAI's training crawler while still allowing its search-index crawler and its on-demand fetch crawler, which is a common pattern for a site that wants to be discoverable and citable by an AI assistant without having its content used for training. The reverse pattern, allowing training but blocking search indexing, is less common but equally valid; there is no single correct default, only a choice that matches what you actually want.

Deciding what to allow

Our free AI crawler robots.txt generator lists the crawlers above with their kind and purpose, lets you toggle allow or disallow per crawler, and builds a ready-to-paste robots.txt, including room for any custom rules a crawler not on this list might need.

The mistake we see most
Treating "block AI crawlers" as one decision. Blocking GPTBot does nothing to stop ChatGPT-User from fetching a page a user directly asks about, and blocking a training crawler does not remove your site from an AI assistant's live search results. Decide separately for training, search indexing and on-demand fetching, based on what outcome you actually want.

How this fits with llms.txt and general AEO work

robots.txt controls access. It says nothing about whether a page is well-written, clearly structured, or likely to be the passage an AI answer engine chooses to quote. That is a separate, larger topic covered in our guide to answer-first content for AI search. A companion but unrelated file, llms.txt, gives AI assistants a curated summary and link list rather than access rules; the two files solve different problems and neither is a substitute for the other.

Sources: OpenAI's publishers and developers FAQ, Anthropic's Claude Help Center on web crawling, Perplexity's crawler documentation, Google's crawler overview.