AI crawler user agents

GPTBot, ClaudeBot, PerplexityBot,
Google-Extended and more.

The AI crawlers you can control in robots.txt, what each is used for, and where its operator documents it. Verified from vendor documentation on 2026-10-06.

AI crawlers are not one thing. Some collect pages that may be used to train models, some fetch pages to power AI search results, and some visit a page only because a person asked an assistant a question. The bots below are listed with the purpose their own operator states.

The list

User-agent tokenOperatorUsed forWhat the operator saysrobots.txtSource
GPTBot OpenAI Training Crawls content that may be used in training OpenAI's generative AI foundation models. Respects robots.txt Vendor docs
Verified 2026-10-06
OAI-SearchBot OpenAI Search Surfaces websites in search results in ChatGPT's search features. Respects robots.txt Vendor docs
Verified 2026-10-06
ChatGPT-User OpenAI User-triggered Visits a page when a user asks ChatGPT or a custom GPT a question. OpenAI says robots.txt rules may not apply, because the user starts the action Vendor docs
Verified 2026-10-06
ClaudeBot Anthropic Training Collects web content that could contribute to training Anthropic's generative AI models. Respects robots.txt (Anthropic also supports Crawl-delay) Vendor docs
Verified 2026-10-06
Claude-SearchBot Anthropic Search Navigates the web to improve the quality of search results for Claude users. Respects robots.txt Vendor docs
Verified 2026-10-06
Claude-User Anthropic User-triggered Accesses a website when a person asks Claude a question. Anthropic documents blocking it with robots.txt Vendor docs
Verified 2026-10-06
PerplexityBot Perplexity Search Surfaces and links websites in Perplexity search results. Perplexity says it is not used to crawl content for AI foundation models. Respects robots.txt Vendor docs
Verified 2026-10-06
Perplexity-User Perplexity User-triggered Visits a page to help answer a question a person asks Perplexity. Perplexity says it generally ignores robots.txt, because the user starts the request Vendor docs
Verified 2026-10-06
Google-Extended Google Training A robots.txt token that lets publishers manage whether content Google crawls may be used to train Gemini models and for grounding in Gemini Apps and Vertex AI. It has no separate user agent string. Control token in robots.txt. Google says it does not affect inclusion in Google Search or ranking Vendor docs
Verified 2026-10-06
Applebot-Extended Apple Training Lets publishers opt out of having content crawled by Applebot used to train Apple's generative AI foundation models. It does not crawl pages itself. Control token in robots.txt. Apple says pages that disallow it can still appear in search results Vendor docs
Verified 2026-10-06
Applebot Apple Search Crawled data powers features such as search in Spotlight, Siri and Safari. Respects robots.txt Vendor docs
Verified 2026-10-06
CCBot Common Crawl Data collection Builds Common Crawl's open repository of web crawl data, which anyone can access. Common Crawl does not say what third parties use it for. Respects robots.txt Vendor docs
Verified 2026-10-06

Each row was read from the vendor's own documentation on 2026-10-06. Crawlers we could not verify from the operator's documentation (for example from Meta, Amazon and ByteDance) are left out rather than guessed. Vendors change their bots, so check the source link before relying on a row.

What the three uses mean

  • Training: content may be used to train generative AI models. Blocking these opts your pages out of training; the operators say these tokens are separate from their search bots.
  • Search: content is used to show and link pages in an AI product's search results. Blocking these can stop your pages from appearing there.
  • User-triggered: a fetch that happens because a person asked an assistant about a page. Whether robots.txt applies is up to the vendor, as the table shows.

robots.txt snippets

Copy, then adjust. Each uses only tokens from the table above.

Opt out of training, stay in AI search

# Opt out of AI model training, stay visible in AI search
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

# Allow the bots that power AI search results
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

Optional: also stop Common Crawl

# Also stop Common Crawl's open dataset from collecting your pages
User-agent: CCBot
Disallow: /

Block everything on this page

# Block every crawler listed on this page, including search and user-triggered ones
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Applebot
User-agent: CCBot
Disallow: /

Conflicting rules resolve by the matching standard: the most specific user-agent group applies, so keep the groups separate as shown. robots.txt is a request, not access control. Blocking Google-Extended does not affect Google Search, and no robots.txt rule guarantees or prevents a citation. Test the file before you ship it with the AI crawler checker, or build one with the robots.txt generator.

Questions

Which AI crawler should I block to stay out of model training but still be cited?

The training tokens are GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Google) and Applebot-Extended (Apple). The search bots are OAI-SearchBot (ChatGPT search), Claude-SearchBot and PerplexityBot. Blocking the first group and allowing the second is the usual way to opt out of training while remaining eligible to appear in AI search answers. Each vendor documents its own bots, and none of them can promise a citation.

Does blocking Google-Extended remove me from Google Search?

No. Google states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal. It only controls use of content for Gemini training and grounding in Gemini Apps and Vertex AI.

Do user-triggered bots obey robots.txt?

It depends on the vendor. OpenAI says robots.txt rules may not apply to ChatGPT-User because a person starts the action. Perplexity says Perplexity-User generally ignores robots.txt for the same reason. Anthropic documents blocking Claude-User with robots.txt. Check each vendor's page, linked in the table, because these policies can change.

Why is the list short?

Every row was read from the vendor's own documentation on 2026-10-06, and anything we could not confirm that way is left out. Other crawlers exist, such as those run by Meta, Amazon and ByteDance. We will add them when we can verify them from the operator's own documentation.

Is robots.txt enough to keep AI bots out?

robots.txt is a request that well-behaved crawlers follow. It is not access control. Content that must stay private should sit behind a login. Some vendors also publish IP ranges so you can verify a crawler, for example Anthropic links a JSON list of its crawler IPs in its documentation.

How do I check that my robots.txt does what I intend?

Use the free AI crawler checker, which reads a site's robots.txt and shows which AI bots it allows and blocks. The robots.txt generator can build a file with presets for blocking AI training bots while keeping AI search.

Sources

Related: Is your robots.txt blocking ChatGPT? · What 99 popular sites block · All free tools

Let Balzac write the articles.

Type your website. Balzac finds the searches you can win, writes articles built to rank on Google and get cited by ChatGPT, and publishes them for you.

3 free articles · No card needed