AI assistants and your site

What the AI vendors say
about accessing your site.

What OpenAI, Anthropic, Perplexity and Microsoft say in their own documentation about how their AI search and assistants access your site. Verified on 2026-10-06.

This page restates each vendor's own documentation and nothing else: no statistics, no claims about what gets cited or ranked. Quotes are verbatim from the pages linked in each row. For Google, see what Google says about AI Overviews and AI Mode.

Vendor by vendor

VendorQuestionWhat the vendor saysSource
OpenAI Which robots does OpenAI use, and how are they managed? “OpenAI uses OAI-SearchBot and GPTBot robots.txt tags to enable webmasters to manage how their sites and content work with AI. Each setting is independent of the others – for example, a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI’s generative AI foundation models.” The page also documents ChatGPT-User and OAI-AdsBot. OpenAI: Overview of OpenAI Crawlers
Verified 2026-10-06
OpenAI How does a site appear in ChatGPT search? “OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features.” OpenAI says: “To help ensure your site appears in search results, we recommend allowing OAI-SearchBot in your site’s robots.txt file and allowing requests from our published IP ranges.” It does not promise inclusion. OpenAI: Overview of OpenAI Crawlers
Verified 2026-10-06
OpenAI How do I opt out of ChatGPT search? “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.” OpenAI: Overview of OpenAI Crawlers
Verified 2026-10-06
OpenAI How do I opt out of training? GPTBot “is used to crawl content that may be used in training our generative AI foundation models. Disallowing GPTBot indicates a site’s content should not be used in training generative AI foundation models.” OpenAI: Overview of OpenAI Crawlers
Verified 2026-10-06
OpenAI Does robots.txt apply to user-triggered fetches? For ChatGPT-User: “Because these actions are initiated by a user, robots.txt rules may not apply. ChatGPT-User is not used to determine whether content may appear in Search. Please use OAI-SearchBot in robots.txt for managing Search opt outs and automatic crawl.” OpenAI: Overview of OpenAI Crawlers
Verified 2026-10-06
OpenAI How long do robots.txt changes take? “For search results, please note it can take ~24 hours from a site’s robots.txt update for our systems to adjust.” OpenAI: Overview of OpenAI Crawlers
Verified 2026-10-06
OpenAI What is OAI-AdsBot? “OAI-AdsBot is used to validate the safety of web pages submitted as ads on ChatGPT.” It “only visits pages submitted as ads, and the data collected by OAI-AdsBot is not used to train generative AI foundation models.” OpenAI: Overview of OpenAI Crawlers
Verified 2026-10-06
Anthropic Which robots does Anthropic use? Three: ClaudeBot (“collecting web content that could potentially contribute to their training”), Claude-User (“When individuals ask questions to Claude, it may access websites using a Claude-User agent”) and Claude-SearchBot (“navigates the web to improve search result quality for users”). Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated April 7, 2026)
Verified 2026-10-06
Anthropic What happens if I disable each one? ClaudeBot: “When a site restricts ClaudeBot access, it signals that the site's future materials should be excluded from our AI model training datasets.” Claude-User: disabling it “prevents our system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search.” Claude-SearchBot: disabling it “prevents our system from indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results.” Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated April 7, 2026)
Verified 2026-10-06
Anthropic How do I block a bot? Add “User-agent: ClaudeBot” and “Disallow: /” (or the other bot name) to the robots.txt in your top-level directory, “for every subdomain that you wish to opt out from.” Anthropic also supports the non-standard Crawl-delay extension. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated April 7, 2026)
Verified 2026-10-06
Anthropic Should I block by IP address instead? “Alternate methods like blocking IP address(es) from which Anthropic Bots operates may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file.” Anthropic publishes a list of its crawler source IPs. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated April 7, 2026)
Verified 2026-10-06
Anthropic Does Anthropic respect robots.txt? “Anthropic’s Bots respect ‘do not crawl’ signals by honoring industry standard directives in robots.txt” and “respect anti-circumvention technologies (e.g., we will not attempt to bypass CAPTCHAs for the sites we crawl.)” Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler? (updated April 7, 2026)
Verified 2026-10-06
Perplexity How does a site appear in Perplexity search results? PerplexityBot “is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models. To ensure your site appears in search results, we recommend allowing PerplexityBot in your site’s robots.txt file and permitting requests from our published IP ranges.” Perplexity: Perplexity Crawlers
Verified 2026-10-06
Perplexity Does robots.txt apply to user-triggered fetches? Perplexity-User “supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a link to the page in its response … Since a user requested the fetch, this fetcher generally ignores robots.txt rules.” Perplexity says it is not used for web crawling or to collect content for training AI foundation models. Perplexity: Perplexity Crawlers
Verified 2026-10-06
Perplexity How long do changes take, and what about firewalls? “Each setting works independently, and it may take up to 24 hours for our systems to reflect changes.” For a web application firewall, Perplexity recommends combining User-Agent matching with IP address verification, using its published IP ranges. Perplexity: Perplexity Crawlers
Verified 2026-10-06
Microsoft (Bing and Copilot) Can I see how my content is cited in Copilot and Bing AI answers? Bing announced AI Performance in Bing Webmaster Tools, “a new set of insights that shows how publisher content appears across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations.” It says average cited pages “does not indicate ranking, authority, or the role of any page within an individual answer,” and: “Bing respects all content owner preferences expressed through robots.txt and other supported control mechanisms.” Bing blog: Introducing AI Performance in Bing Webmaster Tools Public Preview (February 10, 2026)
Verified 2026-10-06
Microsoft (Bing and Copilot) How do I keep part of a page out of Bing snippets and AI answers? The data-nosnippet attribute lets you “mark specific sections of a webpage’s HTML so they don’t appear in Bing Search snippets or AI-generated answers.” Marked content “is still indexed normally, but it will be excluded from snippets and AI summaries,” and “is available for ranking.” Bing also lists noindex, nosnippet and max-snippet, max-image-preview and max-video-preview among its directives. Bing blog: Bing Introduces Support for the data-nosnippet HTML Attribute (October 15, 2025)
Verified 2026-10-06
Microsoft (Bing and Copilot) How do sitemaps and freshness fit in? Bing says the lastmod field in a sitemap “remains a key signal, helping Bing prioritize URLs for recrawling and reindexing,” and that the optional changefreq and priority tags “are ignored by Bing.” It says sitemaps can be referenced in robots.txt or submitted in Bing Webmaster Tools. Bing blog: Keeping Content Discoverable with Sitemaps in AI Powered Search (July 31, 2025)
Verified 2026-10-06
Microsoft (Bing and Copilot) What does IndexNow do, and does it guarantee anything? Bing's AI Performance post says IndexNow “helps keep information fresh across search and AI experiences by notifying participating search engines whenever content is added, updated, or removed.” Bing's IndexNow page says: “Using IndexNow does not guarantee that web pages will be crawled or indexed by search engines.” Bing: How to add IndexNow to your website
Verified 2026-10-06

What to do, using only documented controls

  • Decide per purpose, because the vendors say their bots work independently: search bots (OAI-SearchBot, Claude-SearchBot, PerplexityBot), training bots (GPTBot, ClaudeBot) and user-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User). The AI crawler user agents reference has each token with its vendor's documentation and robots.txt snippets.
  • If you want a vendor's search bot to reach your site, OpenAI and Perplexity recommend allowing its robots.txt token and its published IP ranges, including in any firewall. Anthropic warns that blocking by IP can stop it reading robots.txt.
  • Check what your robots.txt actually says for 16 AI crawlers with the free AI crawler checker. Allow about a day for OpenAI and Perplexity to pick up changes.
  • To keep sections of a page out of Bing snippets and AI answers, Microsoft documents data-nosnippet; the robots meta tag generator writes the meta directives Google documents, which Bing also lists.
  • To tell Bing and other participating engines about changed URLs, use IndexNow; the free IndexNow generator builds the key file and request. Bing says it does not guarantee crawling or indexing.

Questions

How do I allow ChatGPT search, Claude search and Perplexity but opt out of training?

Each vendor documents separate robots.txt tokens. OpenAI: allow OAI-SearchBot for ChatGPT search and disallow GPTBot for training. Anthropic: Claude-SearchBot for search and ClaudeBot for training. Perplexity: PerplexityBot, which it says is not used to crawl content for AI foundation models. Google uses Google-Extended for Gemini training and grounding, covered on our separate Google page. Copy-paste snippets using verified tokens are on the AI crawler user agents reference. Each vendor also asks you to allow its published IP ranges if you want its search bot to reach you.

Do user-triggered fetchers obey robots.txt?

It differs, per the vendors' own pages. OpenAI says robots.txt rules may not apply to ChatGPT-User. Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says disabling Claude-User prevents retrieval in response to a user query. Check each vendor's page, linked in the table, because policies change.

Does allowing these bots get my site cited?

No vendor page we read promises that. OpenAI and Perplexity recommend allowing their search bots to help ensure a site can appear, and Bing says its AI citation data does not indicate ranking or authority. This page makes no claim about what gets cited.

How long do changes take?

OpenAI says about 24 hours for its search systems to adjust after a robots.txt update, and Perplexity says up to 24 hours. Anthropic and Microsoft do not give a figure on the pages we read.

Is IndexNow worth setting up?

IndexNow notifies participating search engines, including Bing, about changed URLs, and Bing says it does not guarantee crawling or indexing. Google is not listed as a participant on indexnow.org. Our free IndexNow generator builds the key file and request; it sends nothing.

How current is this page?

Every row was read from the vendor's own pages on 2026-10-06. We could not read OpenAI's help center articles for publishers (they returned an access error), so nothing from them is used, and Microsoft's Bing Webmaster Help pages need a browser to render, so only Microsoft's own Bing blog and IndexNow pages are cited. Vendors change these pages; the links are the final word.

Sources

Related: What Google says about AI Overviews and AI Mode · AI visibility checker · All free tools

Let Balzac write the articles.

Type your website. Balzac finds the searches you can win, writes articles built to rank on Google and get cited by ChatGPT, and publishes them for you.

3 free articles · No card needed