There are more than thirty AI crawler tokens in circulation and one question worth asking about each. Does blocking it protect your work, or does it delete you from an answer? Those are opposite outcomes, and a pasted list of tokens cannot tell them apart.
Below is every token I could verify against the company that runs it, with the operator's own wording. Where a company publishes nothing, the row says so rather than guessing.
The short answer
- Tokens do one of three jobs: feed model training, feed a search or answer product, or fetch a page a person asked for.
- Blocking a training token costs you nothing in visibility. Blocking a search token removes you from that product's answers.
- Most major operators now split the two. Amazon states it directly: "Each user agent setting is independent of the others."
- Meta is the exception that bundles training and indexing into one token, and Bytespider is the one with no documentation at all.
The three jobs a token can govern
Training. The crawler collects content that may end up in a model's training data. Blocking it changes nothing about whether you get cited today. This is the safe block.
Search or answers. The crawler collects content so the product can retrieve and link your page when a user asks something. Blocking it removes you from that product's answers. This is the expensive block, and it is the one people make by accident.
A user's fetch. Someone pasted your URL into a chat. Almost every operator says the same two things about this category: it is not automated crawling, and robots.txt may not apply to it. Google's own wording is "Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules."
Sort every token into one of those three and the file writes itself.
The table
| Token | Operator | Job | What the operator says |
|---|---|---|---|
GPTBot |
OpenAI | Training | "crawl content that may be used in training our generative AI foundation models" |
OAI-SearchBot |
OpenAI | Search | "surface websites in search results in ChatGPT's search features" |
ChatGPT-User |
OpenAI | User fetch | "certain user actions in ChatGPT and Custom GPTs" |
OAI-AdsBot |
OpenAI | Ads | Validates the safety of pages submitted as ads on ChatGPT |
Googlebot |
Search | Fills the Search index, which is what AI Overviews are served from | |
Google-Extended |
Training and grounding | Manages "whether content Google crawls from their sites may be used for training future generations of Gemini models" and for grounding in Gemini Apps | |
GoogleOther |
Neither | "the generic crawler that may be used by various product teams", one-off internal crawls | |
Google-CloudVertexBot |
Owner-requested | Crawls a site owner requests for building Vertex AI Agents. "It has no effect on Google Search or other products" | |
ClaudeBot |
Anthropic | Training | "collecting web content that could potentially contribute to their training" |
Claude-SearchBot |
Anthropic | Search | "navigates the web to improve search result quality for users" |
Claude-User |
Anthropic | User fetch | "When individuals ask questions to Claude, it may access websites using a Claude-User agent" |
PerplexityBot |
Perplexity | Search | "designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models" |
Perplexity-User |
Perplexity | User fetch | "supports user actions within Perplexity" |
Applebot |
Apple | Search and training | Powers "Spotlight, Siri, and Safari", and crawled data "may also be used to help train Apple foundation models" |
Applebot-Extended |
Apple | Training opt-out | Publishers "can opt-out from having their content used to train generative foundation models by disallowing Applebot-Extended" |
Meta-ExternalAgent |
Meta | Training and indexing | "crawls the web for use cases such as training foundation AI models or improving products by indexing content directly" |
Meta-WebIndexer |
Meta | Search | "navigates the web to improve Meta AI search result quality for users" |
Meta-ExternalFetcher |
Meta | User fetch | "fetches individual links at a user's request" |
FacebookExternalHit |
Meta | Link previews | Crawls content shared on Facebook, Instagram, or Messenger |
Amazonbot |
Amazon | Training and products | "used to improve our products and services... and may be used to train Amazon AI models" |
Amzn-SearchBot |
Amazon | Search | "used to improve search experiences in Amazon products and services... does not crawl content for generative AI model training" |
Amzn-User |
Amazon | User fetch | "supports user actions, such as responding to Alexa queries that require up-to-date information" |
MistralAI-Training |
Mistral | Training | "crawls web content to help build datasets for training Mistral generative AI models" |
MistralAI-Index |
Mistral | Search | "automated crawling of the web for indexing purposes only. It indexes content for Mistral search" |
MistralAI-User |
Mistral | User fetch | "for user actions in Vibe" |
DuckAssistBot |
DuckDuckGo | Search | "crawls pages in real-time for our AI-assisted answers". "This data is not used in any way to train AI models" |
CCBot |
Common Crawl | Archive | Maintains "an open repository of web crawl data that is universally accessible and analyzable by anyone" |
Bytespider |
ByteDance | Unknown | No operator documentation exists |
Sources, in the order they appear: OpenAI, Google, Anthropic, Perplexity, Apple, Meta, Amazon, Mistral, DuckDuckGo, and Common Crawl. All retrieved 2026-08-19.
Where the table gets uncomfortable
Meta-ExternalAgent does two jobs in one token. Meta's own description bundles "training foundation AI models" and "improving products by indexing content directly" into one sentence. There is no way to accept the indexing and refuse the training. Blocking it is a real tradeoff, and it is the only row here where I cannot tell you the safe answer. Meta does run a separate search token, Meta-WebIndexer, so blocking the agent does not necessarily take you out of Meta AI search results.
Applebot is not an AI token, and blocking it hurts. Apple explicitly says content stays "discoverable through Spotlight, Siri, and Safari" when you disallow Applebot-Extended, because Applebot itself keeps crawling. Disallow the wrong one of those two and you leave Apple's search surfaces to protect against training you could have opted out of for free.
Bytespider has nothing behind it. ByteDance publishes no crawler documentation, no purpose statement, no IP range, and no reverse DNS pattern for verification. So a request claiming to be Bytespider cannot be authenticated, and any description of what it is "for" comes from third-party log analysis rather than the operator. I am not going to invent a purpose for it, and neither should the blog post that told you to block it.
Tokens that no longer do anything
Three of these still appear in copied snippets and are worth removing from yours.
anthropic-ai is the notable one. Anthropic's current support article documents exactly three robots, and this is not among them. Say "no longer documented" rather than "deprecated", because Anthropic has not announced a deprecation, and the difference matters: one is an observation about a page, the other is a claim about intent. Either way, if anthropic-ai is your only Anthropic rule, ClaudeBot is not blocked.
Claude-Web is in the same position and absent from the same page.
Google-Extended is not obsolete, but it is misfiled. It is not a crawler. Google states it "doesn't have a separate HTTP request user agent string" and that the token "is used in a control capacity". You will never find it in an access log, however long you grep.
Microsoft does not use a token at all
Bing is the outlier worth knowing about, because there is no BingBot-Extended to add to your list. Microsoft handles this with page-level meta tags instead.
NOCACHE means "only URLs, Titles and Snippets may be used in training Microsoft's generative AI foundation models". NOARCHIVE means content "will not be included in Bing Chat answers, not be linked to in the answers" and "we will not use the content for training". Microsoft adds that content with either tag "will still appear in our search results" (Bing Webmaster Blog).
If your AI crawler policy lives entirely in robots.txt, Microsoft's surfaces are not covered by it.
The newer line in the file: Content-Signal
Cloudflare added a second grammar to robots.txt that is not a user agent rule at all. A Content-Signal line declares intent for three uses: search (building an index and returning links and excerpts), ai-input (real-time use in a generative answer), and ai-train (training or fine-tuning).
It used to appear at the top of our own file as Content-Signal: search=yes,ai-train=no,use=reference. No AI operator has committed to honouring it. It is a stated reservation of rights, written to be readable by a lawyer as much as by a crawler, and treating it as enforcement would be a mistake. If you use Cloudflare, it is probably in your file already and worth reading before someone quotes it back to you.
What our own file said, and what we did about it
On 2026-08-15, https://seobuilder.tech/robots.txt disallowed nine tokens: Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, CloudflareBrowserRenderingCrawler, Google-Extended, GPTBot, and meta-externalagent. Eight matched the list Cloudflare documents for its Managed robots.txt feature.
None of them were in this codebase. src/app/robots.ts produces eight Disallow lines for authenticated paths and a sitemap reference, and nothing else. The AI crawler block was injected at the edge, so the file on localhost and the file customers saw were different documents, and no code review would ever have shown the difference.
We found it by pointing our own agent readiness checker at ourselves, and we turned the setting off the same day. The live file is now byte for byte what the repository generates. The lesson is not that Cloudflare's default is wrong, it is that a setting which changes production and appears in no diff is one you have not really made.
Read against the table above, the shape is defensible: every blocked token is a training or archive token, and every search token (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amzn-SearchBot, Meta-WebIndexer) is allowed. Meta-ExternalAgent is the one row we are on the wrong side of without having decided to be, since it carries indexing as well as training.
Check yours rather than reading about it
Fetch the live file, not the repo copy: curl -s https://yourdomain.com/robots.txt. Then match it against the table. Any Disallow on a row marked Search is costing you answers.
Our AI crawler checker does that match automatically and prints the operator's wording beside each verdict. The robots.txt generator builds a file with one group per token, so a future edit cannot catch the wrong one. Both are free and need no account. Both are our own fetch and parse, not a paid API behind an open endpoint.
One disclosure, since we sell in this category: SEOBuilder asks seven answer engines the same buyer questions and reports which of them cited you, with the citation URLs from each. Claude answers through Anthropic's API rather than claude.ai.
Getting the tokens right makes you eligible to be cited. It does not make you cited, and the second half is the harder measurement problem. How to measure AI visibility covers what can actually be counted. If you are also being told to publish machine-readable files for agents, well-known files for AI agents covers which of those anyone reads.