An AI crawler is a program run by an AI company that fetches pages from your website automatically, the same way Googlebot has fetched pages for Google Search since the 1990s. What changed in the last few years is why they fetch: alongside search-index crawlers, AI companies now also run crawlers that collect training data for language models, and crawlers that fetch one specific page only when a user of their assistant asks it to. Those are three different jobs with three different consequences, and conflating them leads to confused decisions about what to block.
Three kinds of AI crawler, not one
Every major AI company that publishes a documented crawler falls into one or more of these categories:
- Training crawlers fetch content in bulk to build or refresh a dataset used to train a model. OpenAI's GPTBot, Anthropic's ClaudeBot and ByteDance's Bytespider are training crawlers. Blocking one of these stops that company from adding your future content to its training data; it does not remove content already trained on, and it has no effect on whether the resulting assistant can still browse your site live.
- Search-index crawlers fetch and index content so an AI assistant's own answer engine can search it and cite it, similar to how a traditional search engine indexes pages. OpenAI's OAI-SearchBot and Anthropic's Claude-SearchBot fall here. These are explicitly not used for model training; OpenAI has said OAI-SearchBot exists purely to power ChatGPT's search and citation features.
- On-demand fetch crawlers only make a request when a specific user asks the assistant to read or summarize a specific URL right now. OpenAI's ChatGPT-User, Anthropic's Claude-User and Perplexity's Perplexity-User are this type. Because the request is triggered by one person asking about one page, some AI companies treat these fetches differently from bulk crawling; OpenAI has noted that because these fetches are user-initiated, robots.txt rules may not apply to them the same way they do to automated bulk crawlers.
A single company can run crawlers in more than one category. Anthropic, for example, operates ClaudeBot for training, Claude-SearchBot for its search index, and Claude-User for on-demand fetches, each with its own robots.txt user-agent, so you can allow one and block another.
The well-known AI crawler user-agents, as of late 2026
| User-agent | Company | Kind | What it is for |
|---|---|---|---|
GPTBot | OpenAI | Training | Collects content to train OpenAI's models. |
OAI-SearchBot | OpenAI | Search index | Indexes pages for ChatGPT's search and citation features, not training. |
ChatGPT-User | OpenAI | On-demand fetch | Fetches a page only when a ChatGPT user asks it to browse that URL. |
ClaudeBot | Anthropic | Training | Collects content to train Claude models. |
Claude-SearchBot | Anthropic | Search index | Indexes pages for Claude's search and citation features. |
Claude-User | Anthropic | On-demand fetch | Fetches a page only when a Claude user asks it to browse that URL. |
anthropic-ai | Anthropic | Training (legacy) | An older Anthropic training crawler name, largely superseded by ClaudeBot. |
PerplexityBot | Perplexity | Search index | Indexes pages for Perplexity's search and answer features. |
Perplexity-User | Perplexity | On-demand fetch | Fetches a page only when a Perplexity user asks it to browse that URL. |
Google-Extended | Training opt-out | Not a separate crawler; a token controlling whether already-crawled content trains Gemini and powers certain AI features. | |
CCBot | Common Crawl | Training (general) | A general web archive used as training data by many AI labs, not one company. |
Bytespider | ByteDance | Training | Collects content to train ByteDance's AI models. |
New crawlers appear and old ones get renamed. Before treating a name as current, check the company's own documentation. Common Crawl, for one, warns directly that it is aware of crawlers falsely identifying themselves as CCBot, which is a reminder that a user-agent string in a request header is a claim, not a guarantee; a site that genuinely needs to verify a crawler's identity should check the IP ranges or reverse DNS the crawler operator publishes, not just the string.
What robots.txt actually does
robots.txt is a plain text file at the root of your domain (for example, https://example.com/robots.txt) that lists rules per user-agent: which paths a named crawler may or may not fetch. It is part of the Robots Exclusion Protocol, a convention that predates AI crawlers by decades and was formalized as an actual internet standard, RFC 9309, in 2022.
The important limitation: robots.txt is voluntary. Nothing about the file technically prevents a request from reaching your server. It works because major crawler operators choose to check the file before fetching and choose to respect what it says. Anthropic, for instance, states that all three of its Claude bots honor robots.txt, including Claude-User, and notes that this is notable because other vendors' user-initiated fetchers may bypass or generally ignore robots.txt for that same category of request. In other words, even among companies that generally respect the file, the on-demand fetch category is treated inconsistently. robots.txt is not a security control and should never be relied on to keep something actually private; use authentication or simply do not publish sensitive content if that is the goal.
A basic example
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
Sitemap: https://www.example.com/sitemap.xml
This example opts a site out of OpenAI's training crawler while still allowing its search-index crawler and its on-demand fetch crawler, which is a common pattern for a site that wants to be discoverable and citable by an AI assistant without having its content used for training. The reverse pattern, allowing training but blocking search indexing, is less common but equally valid; there is no single correct default, only a choice that matches what you actually want.
Deciding what to allow
- Want to be found and cited by AI assistants? Allow the search-index and on-demand fetch crawlers for the companies whose assistants you care about, so pages can be indexed and quoted.
- Don't want your content used to train future models? Disallow the training crawlers (GPTBot, ClaudeBot, Bytespider, CCBot and similar) while still allowing the search and on-demand crawlers if citation matters to you. These are independent decisions.
- Running a private or paywalled section? robots.txt rules there are advisory at best; use real access control (authentication, a login wall) for anything that actually needs to stay restricted, and remember search engines and AI crawlers alike should generally be disallowed from any path that should not be publicly discoverable at all.
Our free AI crawler robots.txt generator lists the crawlers above with their kind and purpose, lets you toggle allow or disallow per crawler, and builds a ready-to-paste robots.txt, including room for any custom rules a crawler not on this list might need.
How this fits with llms.txt and general AEO work
robots.txt controls access. It says nothing about whether a page is well-written, clearly structured, or likely to be the passage an AI answer engine chooses to quote. That is a separate, larger topic covered in our guide to answer-first content for AI search. A companion but unrelated file, llms.txt, gives AI assistants a curated summary and link list rather than access rules; the two files solve different problems and neither is a substitute for the other.
Sources: OpenAI's publishers and developers FAQ, Anthropic's Claude Help Center on web crawling, Perplexity's crawler documentation, Google's crawler overview.