AI crawlers: every bot, and how to allow or block it
The bots that AI companies send to websites, what each one actually does, and the exact robots.txt lines to let it in or keep it out. Every entry is checked against the company’s own documentation.
Last checked 2026-10-08
Not every AI bot does the same job
Blocking “AI bots” as one group is the most common mistake. Most AI companies run separate bots for training, for search and for user requests, and blocking the wrong one can take you out of AI answers without stopping the thing you meant to stop.
Builds an AI search index
It reads pages so the AI assistant can find and link to them when answering questions. Blocking it is the usual way to drop out of that assistant’s search results.
Fetches a page when a user asks
It only visits when a person asks the assistant something that needs your page. Because a person triggered the visit, its operator says robots.txt rules may not apply.
Collects pages to train AI models
Pages it collects may be used to train future AI models. Blocking it is an opt-out from training, and does not by itself remove you from any AI search results.
Search engine crawler that also feeds AI
A classic search engine crawler. The same index it builds is used by that company’s AI answers, so blocking it removes you from normal search results too.
Control setting, not a crawler
It never visits your site. It is a name you can use in robots.txt to tell the company how pages its normal crawler already collected may be used for AI.
The full list
The name in the first column is exactly what goes after User-agent: in robots.txt. Open any bot for its full details.
| Bot | Company | What it does | Follows robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | AI training | Yes |
| OAI-SearchBot | OpenAI | AI search | Yes |
| ChatGPT-User | OpenAI | User request | May not, for user requests |
| ClaudeBot | Anthropic | AI training | Yes |
| Claude-SearchBot | Anthropic | AI search | Yes |
| Claude-User | Anthropic | User request | Yes |
| PerplexityBot | Perplexity | AI search | Yes |
| Perplexity-User | Perplexity | User request | May not, for user requests |
| Google-Extended | Control token | Not a crawler | |
| Googlebot | Search engine | Yes | |
| Bingbot | Microsoft | Search engine | Yes |
| Applebot-Extended | Apple | Control token | Not a crawler |
| meta-externalagent | Meta | AI training | Yes |
| meta-webindexer | Meta | AI search | Yes |
| CCBot | Common Crawl | AI training | Yes |
| Amazonbot | Amazon | AI training | Yes |
| Bytespider | ByteDance | AI training | Documented, but reports vary |
Deciding what to allow
For most businesses that want to be found through AI assistants, the search bots and the user-request bots should be let in. The training bots are a separate decision about how your content may be used, and blocking them does not remove you from AI search results. Our guide on blocking AI crawlers walks through that choice, and the robots.txt guide shows how the rules combine in one file.
robots.txt is a request, not a lock. To see which bots really visit, and whether they are who they say they are, check your server logs against the IP lists linked on each bot’s page.
See which AI bots your site lets in
The free AI readiness check reads your robots.txt and reports which of the main AI crawlers are allowed or blocked, along with everything else that decides whether AI can understand your site.
Check your site free