Pick which well-known AI crawlers may read your site, add any custom rules, and get a ready-to-paste robots.txt. Training crawlers, on-demand fetch crawlers and search-index crawlers are all labeled, since allowing or blocking each one means something different.
A training crawler like GPTBot or ClaudeBot fetches pages in bulk to build a dataset used to train a model. An on-demand crawler like ChatGPT-User or Claude-User only fetches a specific page when a user of that assistant asks it to read or browse that exact URL. Blocking one does not block the other.
Only if the crawler chooses to follow it. robots.txt is a voluntary convention: well-behaved crawlers from major AI companies generally respect it, but robots.txt cannot technically prevent a request from reaching your server. It is not a security control.
Google-Extended is not a separate crawler; Googlebot still fetches your pages for Search. Google-Extended is a token that controls whether that already-crawled content may also be used to train Gemini models and power certain AI features, separate from normal Search indexing and ranking.
It covers the well-known AI crawlers that publish a respected, documented user-agent string. New crawlers appear and old ones change; use the custom rules box to add any other user-agent your logs show, and check the company's own documentation for the current name before you rely on it.