The short answer: allow the bots that search and the bots that fetch on a person’s behalf; decide separately about the bots that train. Blocking a training bot does not remove you from an answer. Blocking a search bot does.
Every major AI company now publishes more than one crawler, and the names look alike. Sites end up blocking the wrong one — usually through a “block AI scrapers” switch nobody remembers turning on. This guide sorts nine names by what they actually do, using each vendor’s own documentation.
01Three jobs behind nine names
Every AI bot does one of three things, and the cost of blocking is completely different for each.
What blocking each kind costs you
PER VENDOR DOCUMENTATION · LINKED BELOW
Search
BUILDS THE INDEX
Block it: you can’t be shown.
OpenAI states sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Claude-SearchBot and PerplexityBot play the same role.
User fetch
OPENS A PAGE ON REQUEST
Block it: a customer asked, and it can’t open the door.
ChatGPT-User, Claude-User, Perplexity-User. Triggered by a person’s question about you, not by a crawl schedule.
Training
FEEDS FUTURE MODELS
Block it: your text stays out of training data.
GPTBot, ClaudeBot, Google-Extended. A real choice with arguments both ways — and not what decides whether you are named today.
Perplexity states neither of its bots is used for foundation-model training. Anthropic and OpenAI keep training on separate, separately blockable bots.
02The map, vendor by vendor
OpenAI
SOURCE: OPENAI BOT DOCUMENTATION
OAI-SearchBot- Search. Surfaces sites in ChatGPT’s search features. Allowing it is recommended for search visibility; OpenAI says robots.txt changes take about 24 hours to apply.
ChatGPT-User- User fetch. Visits a page when someone asks ChatGPT or a custom GPT. Not used to decide whether content appears in Search. Because a user triggers it, robots.txt rules may not apply.
GPTBot- Training. Crawls content that may be used to train generative AI models. Disallowing it has no direct effect on search appearance.
Anthropic
SOURCE: ANTHROPIC SUPPORT
Claude-SearchBot- Search. Indexes content to improve the relevance of Claude’s search results. Block it and you are not in the index Claude retrieves from.
Claude-User- User fetch. Fetches a page when a person asks Claude something that needs it. Block it and Claude cannot open your page for the customer who asked.
ClaudeBot- Training. Collects public content that may be used to train models. Blocking opts you out of training data only.
Anthropic states its bots honour robots.txt and publishes verifiable IP ranges, so you can tell a real Claude request from an imitation. Our walkthrough of how Claude reads a page covers what happens after the door opens.
Perplexity and Google
SOURCES: PERPLEXITY · GOOGLE SEARCH CENTRAL
PerplexityBot- Search. Built to surface and link websites in Perplexity results. Respects robots.txt. Perplexity states it is not used for model training.
Perplexity-User- User fetch. Visits pages to answer a user’s query. Perplexity notes it generally ignores robots.txt because the user initiated the request.
Google-Extended- Training and grounding token. Not a separate crawler: a robots.txt token that controls whether Google-crawled content may be used to train Gemini models and for grounding. Google states it does not affect inclusion in Google Search and is not a ranking signal.
03How good sites lock themselves out
The blanket switch. A CDN setting or WordPress plugin labelled “block AI bots” usually covers all of them at once, search and user fetch included. The intent was to stop training scrapers; the effect is that the assistant a customer just asked about you is turned away.
The half-written group. A crawler obeys only the most specific group that names it and ignores User-agent: * once such a group exists. Add one named block and every rule you wanted must be repeated inside it.
The firewall you forgot. robots.txt is a request. A bot-protection rule at your CDN or host is a wall. You can pass the first and fail the second, so test them separately.
04A robots.txt that keeps you findable
One reasonable default: let search and user-fetch bots in, and make an explicit call on training. This example allows everything that decides visibility and opts out of training — flip the last block if you would rather contribute.
# search + user fetch: these decide whether you can be named User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User Allow: / # training: your call. Blocking these does not remove you from answers User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended Disallow: / User-agent: * Allow: / Sitemap: https://yoursite.com/sitemap.xml
Then check it from outside. Ask for the page as each bot would and look for anything other than a normal response:
# 200 means the door is open. 403 means a firewall answered for you
curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot" https://yoursite.com/
curl -s -o /dev/null -w "%{http_code}\n" -A "Claude-User" https://yoursite.com/
curl -s -o /dev/null -w "%{http_code}\n" -A "PerplexityBot" https://yoursite.com/
A user-agent test only shows how your server treats the name, not whether a real bot from a real IP is admitted. Where a vendor publishes IP ranges, use them to allow-list the genuine bot at the firewall.
05What robots.txt cannot fix
Letting the bots in only opens the door. A page that is empty until JavaScript runs, a price locked inside an image, or a phone number nobody wrote as text is still unreadable to a fetcher. The same is true of a page that never states its own answer. Access is the first gate; being quotable is the second, and the basics of SEO and AEO in 2026 cover it.
Where should I start?
- Open
yoursite.com/robots.txtand search it for the nine names above. - Check your CDN and plugins for an “AI bots” switch and see exactly which bots it covers.
- Decide on training separately from search. Write the decision down.
- Run the three
curlchecks and look for a 200. - Run the free scan: it checks crawler access as one of its 17 signals.
06Four short answers
Should I block GPTBot?
It is your call about training, not about visibility. OpenAI documents that blocking it has no direct effect on search appearance. OAI-SearchBot is the one that matters for that.
Which crawlers must I allow to be recommended?
OAI-SearchBot, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User. Training bots are a separate choice.
Does blocking Google-Extended hurt my Google ranking?
No. Google says it does not affect inclusion in Search and is not a ranking signal.
Does robots.txt stop assistants from reading my page?
Not always. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it. Only a firewall rule actually stops them.