Is Your robots.txt Blocking ChatGPT? What Each AI Crawler Controls (2026)

You blocked GPTBot, refreshed your robots.txt, and still can’t figure out why your pages aren’t showing up in ChatGPT-style answers. Or you did the opposite: you meant to opt out of training, and you accidentally shut the door on being cited.

The problem is that “AI bots” aren’t one thing. Some crawlers control whether an assistant can fetch your pages for search/answer features. Others only control whether your content can be collected for model training. If you mix those up, a single robots.txt line can quietly change your visibility.

OpenAI makes the split explicit in its Overview of OpenAI Crawlers: OAI-SearchBot is the crawler used to surface websites in ChatGPT’s search features, and OpenAI recommends allowing it if you want to appear in ChatGPT search results. OpenAI also says that opting out of OAI-SearchBot means your site won’t be shown in ChatGPT search answers.

GPTBot is a separate switch. OpenAI says GPTBot is for crawling content that may be used in training its generative AI foundation models, and disallowing GPTBot indicates you don’t want your content used for training. OpenAI also states the robots.txt settings for OAI-SearchBot and GPTBot are independent, so “block GPTBot” is a training choice, not a ChatGPT search opt-out.

This guide shows you how to tell which bots matter for citations vs training, how robots.txt matching works when multiple groups and rules collide, and how to confirm what your site is doing right now (you can also sanity-check it with the free checker at https://hirebalzac.ai/free-seo-tools/ai-crawler-checker/). You’ll also see the gotcha that trips up even careful teams: your CDN or firewall can block AI crawlers before robots.txt even gets a vote. For a token-by-token list with each vendor’s own documentation, see our AI crawler user agents reference.

Which AI Bots Matter for Citations (Search) vs Training?

ChatGPT-User is a distraction for most robots.txt decisions. The practical question is simpler: which user-agents control whether your pages can be fetched for AI answers (citations), and which ones only control model training. If you are debugging ai crawlers robots.txt rules, start by separating “search/answer” bots from “training” bots.

The table below lists the names you will most often see in robots.txt discussions. Use it to decide what to allow, what to block, and what changes will not affect citations.

Bot (User-Agent) Operator Purpose Robots.txt Impact (Allow vs Block)
OAI-SearchBot OpenAI Search and retrieval for ChatGPT search features Controls whether content can appear in ChatGPT search results. OpenAI recommends allowing it, and says opting out means your site will not be shown in ChatGPT search answers. Source: Overview of OpenAI Crawlers.
GPTBot OpenAI Training data collection for generative AI foundation models Controls training use, not ChatGPT search eligibility. OpenAI says GPTBot and OAI-SearchBot settings are independent. Source: Overview of OpenAI Crawlers.
ChatGPT-User OpenAI User-initiated fetching OpenAI says it is not used for automatic web crawling and is not used to determine whether content may appear in Search. Source: Overview of OpenAI Crawlers.
Claude-SearchBot Anthropic Search and retrieval for Claude answers Typically affects whether Claude can fetch and cite pages during search-style experiences. Verify the exact user-agent your logs show.
ClaudeBot Anthropic Training crawler Blocking ClaudeBot targets training access. Some CDNs (example: Cloudflare) include rules that disallow ClaudeBot by default in managed robots.txt. Source: Cloudflare managed robots.txt.
Claude-User Anthropic User-initiated fetching Often behaves more like a browser request. Robots.txt handling can differ from automated crawlers.
PerplexityBot Perplexity Search and retrieval for Perplexity answers Blocking can reduce Perplexity’s ability to fetch and cite. Confirm the user-agent name in your server logs.
Perplexity-User Perplexity User-initiated fetching May not behave like a classic crawler. Treat it separately from PerplexityBot in your analysis.
Googlebot Google Google Search crawling (also feeds Google’s AI experiences) Blocking affects indexing and any downstream use in Google surfaces that rely on Search crawling.
Google-Extended Google Training and model improvement controls (separate from Googlebot) Blocking targets training use. Some CDNs (example: Cloudflare) include rules that disallow Google-Extended in managed robots.txt. Source: Cloudflare managed robots.txt.
Bingbot Microsoft Bing Search crawling (also influences Microsoft Copilot via Bing) Blocking affects Bing indexing and any AI features that depend on Bing’s index.
Applebot-Extended Apple Training controls (separate from Applebot) Blocking targets training use. Some CDNs (example: Cloudflare) include rules that disallow Applebot-Extended in managed robots.txt. Source: Cloudflare managed robots.txt.
Meta-ExternalAgent Meta Training crawler Blocking targets training access. It does not directly control ChatGPT citations.
CCBot Common Crawl Web crawl dataset used by many research and commercial projects Blocking limits inclusion in Common Crawl snapshots. This is training-adjacent, not a direct “AI answers” switch.
Bytespider ByteDance Training and data collection crawler Blocking targets training access. It does not answer “is my site blocking chatgpt”.

Quick Rules for Robots.txt Decisions

If your goal is citations, prioritize the search crawlers: OAI-SearchBot for ChatGPT search, plus the search bots you see in your own logs (PerplexityBot, Claude-SearchBot, Googlebot, Bingbot). If your goal is training opt-out, block training bots like GPTBot and Google-Extended. Remember the common misconception behind “robots.txt gptbot”: block GPTBot to limit training, but do not expect it to remove you from ChatGPT search.

How Robots.txt Matching Works (So Your Rules Do What You Think)

Most “is my site blocking ChatGPT” mistakes come from robots.txt matching, not from the bot name you picked. A single User-agent: * block can quietly override your intent if you do not understand how crawlers choose rules.

Robots.txt is a set of groups. Each group starts with one or more User-agent lines, then lists Allow and Disallow paths. A crawler reads your robots.txt, finds the best matching group for its user-agent string, then applies the best matching rule within that group.

User-Agent Group Selection: Specific Beats User-agent: *

The most specific user-agent group wins over User-agent: *. If you define a group for OAI-SearchBot and another group for *, OAI-SearchBot should follow the OAI-SearchBot group. If you forget the OAI-SearchBot group, then the * group becomes the default and can block search crawlers you meant to allow.

Two practical implications:

  • If you want to block GPTBot but allow ChatGPT search crawling, put GPTBot in its own group and keep your * group permissive.
  • If you want to opt out of ChatGPT search, you must target OAI-SearchBot specifically (OpenAI treats it separately from GPTBot per its Overview of OpenAI Crawlers).

Rule Matching: Longest Path Wins, Allow Wins Ties

Inside the chosen group, crawlers compare rules by how specifically they match the URL path. The longest matching path wins. If an Allow rule and a Disallow rule match the same URL with the same length, Allow wins the tie. This is why you can block a directory but open a specific file inside it.

Example logic (hypothetical):

  • Disallow: /private/ blocks everything under /private/.
  • Allow: /private/press-kit.pdf re-allows that one file because it matches more specifically.

Wildcards: * And $ Change What “Matches”

Robots.txt supports pattern matching that trips people up:

  • * matches any sequence of characters. Use it to match URL patterns like parameters or file types (for example, “block any URL containing ?”).
  • $ anchors the match to the end of the URL. Use it when you mean “only this exact ending” (for example, “block URLs ending in .pdf”).

If your robots.txt uses broad wildcards in User-agent: *, treat it as risky. One overly generic pattern can block OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, or Bingbot and you will only notice when citations and previews disappear.

How to Check: Is My Site Blocking ChatGPT Right Now?

Screenshot of workspace Balzac

Broad User-agent: * rules cause most “is my site blocking chatgpt” surprises. You think you blocked a training bot, but you actually blocked OAI-SearchBot, which controls whether ChatGPT search features can fetch and cite your pages. The fastest fix is a quick robots.txt audit focused on the OpenAI search crawler, not just robots.txt gptbot.

  1. Open your robots.txt in a browser. Visit https://yourdomain.com/robots.txt (root domain, not a subfolder). If you run multiple hosts (www and non-www, subdomains), check each one. Crawlers read the robots.txt for the exact host they fetch.
  2. Confirm it returns HTTP 200. Robots.txt needs to load cleanly. A redirect chain, a 403, or a 404 can change crawler behavior and can also mean you are not serving the file you think you are. Cloudflare notes it prepends its managed robots.txt only when the origin robots.txt returns HTTP 200, which is a good reminder to verify the status code end-to-end (Cloudflare managed robots.txt).
  3. Search for the exact user-agent names that control ChatGPT visibility. Look for a group that starts with User-agent: OAI-SearchBot. If you see Disallow: / under that group, you opted out of ChatGPT search answers. OpenAI documents OAI-SearchBot as the crawler used to surface sites in ChatGPT search features (Overview of OpenAI Crawlers).
  4. Check whether your global rules accidentally block search bots. If you do not have an OAI-SearchBot group, your User-agent: * group applies. Scan for patterns like Disallow: /, Disallow: /*?, or broad filetype blocks that might catch HTML pages.
  5. Separate “block GPTBot” from “block ChatGPT search.” A User-agent: GPTBot group with Disallow: / only signals training opt-out. OpenAI says GPTBot and OAI-SearchBot settings are independent (Overview of OpenAI Crawlers).

If you want a quick verdict without manually simulating robots matching, run your domain through Balzac’s free checker: AI crawler checker. It is a fast way to spot the common failure modes: an OAI-SearchBot block, an overbroad User-agent: * rule, or a robots.txt file that does not load as expected.

What To Do If You Find A Block

Fixing the issue usually means adding an explicit User-agent: OAI-SearchBot group that allows the sections you want cited, then tightening your training blocks under User-agent: GPTBot. If your robots.txt already looks correct but citations still do not appear, check your CDN or WAF settings next. A firewall rule can block OAI-SearchBot even when robots.txt allows it.

Step-by-Step: Block GPTBot but Allow OAI-SearchBot

Most robots.txt gptbot fixes fail for one reason: people block GPTBot (training) and accidentally block OAI-SearchBot (ChatGPT search) in the same file. OpenAI treats these controls separately. Its Overview of OpenAI Crawlers says the robots.txt setting for OAI-SearchBot is independent from GPTBot, so you can block training while still allowing search crawling.

Use the steps below to block GPTBot while keeping OAI-SearchBot able to fetch pages for AI answers.

  1. Open your live robots.txt: visit https://yourdomain.com/robots.txt and copy the current contents into a safe place.
  2. Find your default group: look for User-agent: *. If it contains Disallow: / or broad wildcard blocks, assume it may block OAI-SearchBot unless you add a specific allow group.
  3. Add an explicit OAI-SearchBot group that allows crawling where you want citations. Start permissive, then tighten later if needed.
  4. Add a GPTBot group that blocks training access. The simplest opt-out is Disallow: /.
  5. Publish robots.txt and validate it: confirm the URL still returns HTTP 200, then watch your server logs for requests from OAI-SearchBot/1.4 and GPTBot/1.4 (OpenAI shows these example user-agent strings in its crawler documentation).

Here is a minimal robots.txt example that allows ChatGPT search crawling but blocks training:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot Disallow: /

User-agent: * Allow: /

If you need to keep some areas off-limits for all bots, put those restrictions in User-agent: *, then add narrower Allow exceptions inside the OAI-SearchBot group. Remember the matching rules from RFC 9309: the specific user-agent group wins, then the longest matching path wins, and Allow wins ties.

Free tool: the robots.txt generator builds a file with your sitemap and blocked paths, plus presets for blocking AI training bots while keeping AI search.

After You Publish, Confirm You Did Not Block OAI-SearchBot

Robots.txt edits are easy to ship and easy to get subtly wrong. If everything looks right in robots.txt but OAI-SearchBot never hits your site, check your CDN or WAF next. Cloudflare Bot Management, custom firewall rules, and other bot controls can block AI crawlers even when your robots.txt allows them.

How Long Until ChatGPT Search Sees Your Robots.txt Changes?

If you just updated robots.txt gptbot rules and you are refreshing results to see if ChatGPT “noticed,” expect a delay. OpenAI says it can take about 24 hours after a robots.txt update for its search systems to adjust (source: Overview of OpenAI Crawlers).

That timing matters because robots.txt changes often look correct in a browser immediately, while OAI-SearchBot (the crawler that feeds ChatGPT search features) still uses the previous policy until its systems reprocess your file. During that window, you can see confusing behavior: old fetch patterns in logs, missing citations, or a page that starts showing again after you already “fixed” it.

What to Monitor During the 24-Hour Window

You cannot force OpenAI to re-read your robots.txt on demand, but you can verify the pieces you control and watch for signals that the new rules took effect.

  1. Confirm the robots.txt you edited is the one bots receive. Re-check /robots.txt on the exact host that serves your content (www vs non-www, and any relevant subdomains). Make sure it still returns HTTP 200 end-to-end, including through your CDN.
  2. Watch server logs for OAI-SearchBot fetches. Look for requests with an OAI-SearchBot user-agent string. OpenAI’s examples include OAI-SearchBot/1.4, and OpenAI says it may add a robots.txt marker when fetching robots.txt (source: Overview of OpenAI Crawlers). If you only see GPTBot activity after your change, you may have fixed training access while still blocking ChatGPT search.
  3. Check for “allowed but blocked” conflicts. If you changed a broad User-agent: * rule, confirm you did not accidentally disallow key directories (for example, /blog/ or /resources/) that OAI-SearchBot needs for citations.
  4. Verify CDN and WAF behavior did not change. A firewall rule can block OAI-SearchBot regardless of robots.txt. If logs show no hits at all, that points to network blocking, not robots matching.
  5. Re-test with an automated checker after a day. If you want a quick confirmation that your current policy allows the right bots, re-run Balzac’s free AI crawler checker after the adjustment window, then compare the result with what you see in logs.

If ChatGPT search visibility still does not recover after that period, treat it as a configuration issue, not a waiting game. The next step is to re-validate your OAI-SearchBot group and then audit CDN or WAF bot controls.

The Contrarian Gotcha: Your CDN or Firewall Can Block AI Bots Anyway

Your robots.txt can be perfect and you can still be asking “is my site blocking chatgpt?” because robots.txt is only one gate. Your CDN, WAF, or bot protection can block OAI-SearchBot (and other AI crawlers) at the HTTP layer, so the crawler never gets a chance to read the rules you wrote.

This is why “robots.txt gptbot” fixes sometimes change nothing. You might allow OAI-SearchBot and block GPTBot correctly, but a firewall rule still returns a 403, a JavaScript challenge, or a hard block to the bot’s IPs or user-agent.

Where CDN And WAF Blocks Usually Hide

Start by assuming your edge layer can override your intent. Common places to look:

  • Cloudflare Bot Management and firewall rules: a rule that blocks “AI bots,” “unknown bots,” or specific user-agents can stop OAI-SearchBot even when robots.txt allows it.
  • Managed robots.txt features: Cloudflare has a managed robots.txt setting that generates and maintains robots.txt directives for known AI crawlers. Cloudflare also notes robots.txt compliance is voluntary, so this setting signals intent but does not technically prevent access.
  • Origin protection: security plugins, load balancers, or an origin firewall can block data center traffic patterns that crawlers often resemble.
  • Rate limiting: aggressive rate limits can throttle crawlers until they give up, especially on large sites.

Cloudflare’s managed robots.txt behavior can also confuse debugging. Cloudflare says it will prepend its managed robots.txt before your origin robots.txt when the origin robots.txt returns HTTP 200 (Cloudflare managed robots.txt). If your origin serves a 404 or 403 for /robots.txt, you may be looking at a different file than you think.

Cloudflare’s example managed robots.txt explicitly disallows several training crawlers, including User-agent: GPTBot with Disallow: /, plus ClaudeBot, Google-Extended, and Applebot-Extended (Cloudflare managed robots.txt). That can be what you want for training opt-out, but it does not guarantee your search bots are allowed.

If you suspect an edge block, validate it like this:

  1. Check server logs at the edge and origin for requests with OAI-SearchBot/1.4 or GPTBot/1.4 user-agent strings (OpenAI lists these examples in its Overview of OpenAI Crawlers).
  2. Look at the HTTP status codes those requests receive. A 403 or challenge page means your WAF blocked before robots.txt logic mattered.
  3. Audit allowlists in your CDN/WAF for OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, and Bingbot, based on what your logs show.

Should You Use llms.txt in 2026?

Edge blocks usually come from robots.txt or a firewall. llms.txt is different: it is a separate, community-driven proposal for a plain-text guide to a site, and it does not replace ai crawlers robots.txt rules. If your question is “is my site blocking chatgpt,” llms.txt is not the first file to check.

Llms.txt is typically a Markdown file placed at /llms.txt that lists the pages you consider most useful. Think of it as a “publisher note” for tools that choose to read it, not an access control system. Google says Google Search ignores it (what Google says, with sources), and the crawler documentation from OpenAI, Anthropic and Perplexity, read on October 6, 2026, does not say their crawlers use it.

What Llms.txt Can and Can’t Do

Llms.txt can:

  • List preferred URLs, such as canonical docs pages, pricing pages, or a press page, for anyone or any tool that chooses to read the file.
  • Reduce ambiguity for humans auditing your AI visibility setup, because it centralizes your intent in one place.

Llms.txt can’t:

  • Force crawling or indexing. Search and answer crawlers still decide what to fetch, and they still follow robots.txt and network controls.
  • Influence Google Search. Google says its Search ignores the file and that it will neither harm nor help your visibility or rankings.
  • Guarantee citations. No assistant promises to cite a site just because llms.txt exists.
  • Override robots.txt. If you disallow OAI-SearchBot in robots.txt, you opted out of ChatGPT search answers regardless of what llms.txt says (see OpenAI’s Overview of OpenAI Crawlers).

If you want to influence whether ChatGPT can retrieve and cite pages, focus on the bots that matter. OpenAI documents that OAI-SearchBot is used for ChatGPT’s search features, while GPTBot is for training collection, and their robots.txt settings are independent (Overview of OpenAI Crawlers). That distinction matters more than any llms.txt hint.

Use llms.txt, if at all, as low-risk housekeeping if you have the bandwidth. Keep it short, point to your best canonical sources, and avoid treating it like a switch for “block GPTBot” or “allow ChatGPT.” Robots.txt remains the control plane for crawlers, and your CDN or WAF remains the control plane for actual access.

Free tool: the llms.txt generator creates a file in the llmstxt.org format from your key pages.

FAQ: OAI-SearchBot, ChatGPT-User, and Blocking GPTBot

Robots.txt is the control plane for crawlers, but the names matter. Most confusion comes from mixing up “search/answer” bots with “training” bots, then assuming a robots.txt gptbot change affects ChatGPT visibility. It usually does not.

Common Questions About OAI-SearchBot, ChatGPT-User, and GPTBot

Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI says GPTBot is used to crawl content that may be used for training, and disallowing it indicates your content should not be used for training. OpenAI also says the robots.txt settings for GPTBot and OAI-SearchBot are independent. Source: Overview of OpenAI Crawlers.

Which OpenAI bot actually controls whether I can be cited in ChatGPT search answers?
OpenAI documents OAI-SearchBot as the crawler used to surface websites in ChatGPT’s search features. OpenAI also states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Source: Overview of OpenAI Crawlers.

Does ChatGPT-User follow robots.txt?
Sometimes, robots.txt rules may not apply. OpenAI says ChatGPT-User is not used for automatic web crawling, and robots.txt rules may not apply because the actions are initiated by a user. It also is not used to determine whether content may appear in Search. Source: Overview of OpenAI Crawlers.

What should I do if I want to opt out of ChatGPT search results?
Target OAI-SearchBot in robots.txt. OpenAI instructs webmasters to use OAI-SearchBot for managing Search opt-outs and automatic crawl. Practically, that means adding an explicit group for OAI-SearchBot and disallowing the paths you do not want fetched (or Disallow: / if you want a full opt-out). Source: Overview of OpenAI Crawlers.

I allowed OAI-SearchBot, but I still think my site is blocking ChatGPT. What now?
Assume something upstream is blocking access. A CDN or WAF can return a 403, a bot challenge, or rate limits that prevent crawling even when robots.txt allows it. Check edge firewall rules, bot protection settings, and origin logs for requests that look like OAI-SearchBot/1.4 (OpenAI’s documented example user-agent string). Source: Overview of OpenAI Crawlers.

How fast do robots.txt changes take effect for ChatGPT search?
OpenAI says it can take about 24 hours after a robots.txt update for its search systems to adjust. If you changed rules and nothing happened yet, wait out that window, then verify again using logs and a policy check. Source: Overview of OpenAI Crawlers.

If you want a practical next step, test your current rules against the exact OpenAI user-agents, especially OAI-SearchBot, then fix the mismatch you find. The fastest way to get a clear yes or no is Balzac’s free AI crawler checker, then confirm the result in your server logs.

Sources

← All articles

Free AI visibility check

Is ChatGPT citing your site?

See if ChatGPT search and Google AI Overviews cite you for your top searches, and who they cite instead. About a minute, no account.

Want content like this on autopilot?

Balzac researches, writes and publishes articles like this one to your site. Every week.

3 free articles · No card needed