99 Popular Sites' robots.txt vs 16 AI Crawlers (Oct 2026)

We took 99 popular websites and checked each one’s public robots.txt for 16 AI crawlers. This post reports what we found, how we measured it, and the full per-site data. It is a snapshot taken on 2026-10-06, and it describes what the files say, nothing more.

Headline findings

Percentages below are of the 93 sites whose robots.txt we could read, not of all 99. Of the 99 sites, 93 had a readable robots.txt, 1 (khanacademy.org) had no robots.txt, and 5 could not be read reliably (godaddy.com, linkedin.com, mayoclinic.org, stackoverflow.com, udemy.com). We left those six out of every percentage.

  • 55 of 93 sites (59%) block none of the 16 crawlers. The other 38 block at least one.
  • 37 of 93 (40%) block at least one training crawler (GPTBot, ClaudeBot, CCBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent or Bytespider).
  • 23 of 93 (25%) block at least one AI search crawler (OAI-SearchBot, Claude-SearchBot or PerplexityBot). Training crawlers are blocked more often than search crawlers.
  • 14 sites block GPTBot but allow OAI-SearchBot: airbnb.com, bloomberg.com, canva.com, ebay.com, figma.com, forbes.com, glassdoor.com, healthline.com, medium.com, reuters.com, techcrunch.com, tripadvisor.com, webmd.com, zalando.com. Only 1 site did the reverse (facebook.com blocks OAI-SearchBot but not GPTBot).
  • 11 sites block all three AI search crawlers: amazon.com, cnn.com, imdb.com, instagram.com, nytimes.com, pinterest.com, quora.com, reddit.com, tiktok.com, x.com, yelp.com.
  • 1 site blocks all 16 crawlers, including Googlebot and Bingbot: reddit.com.
  • llms.txt: 29 of 99 sites publish one. 24 of those 29 are in our software and AI companies group. See the llms.txt section below.

By category

We assigned each site to one of five groups ourselves. The assignment is approximate and some sites could fit in more than one group, so read the rows as rough patterns rather than precise industry statistics. Percentages are of readable robots.txt files within each group; the llms.txt column counts all listed sites in the group.

CategoryReadableBlocks GPTBotBlocks OAI-SearchBotBlocks ClaudeBotBlocks any training botHas llms.txt
News and media1911 (58%)4 (21%)17 (89%)18 (95%)0 of 20
Social and community96 (67%)7 (78%)6 (67%)7 (78%)0 of 11
Shopping and travel156 (40%)2 (13%)6 (40%)6 (40%)4 of 15
Software and AI companies442 (5%)0 (0%)2 (5%)3 (7%)24 of 45
Other61 (17%)0 (0%)1 (17%)3 (50%)1 of 8

Group sizes are small (6 to 44 readable sites), so a single site moves a percentage by several points.

By crawler

How many of the 93 readable sites block each of the 16 crawlers. A site counts as blocking a crawler if its robots.txt disallows path / for that user-agent (see the method below). “Training”, “Search” and “User-triggered” are our labels based on how each operator describes the crawler; Amazonbot is listed as “Other”.

CrawlerOperatorRoleSites blocking
BytespiderByteDanceTraining33 (35%)
ClaudeBotAnthropicTraining32 (34%)
CCBotCommon CrawlTraining31 (33%)
Applebot-ExtendedAppleTraining30 (32%)
Meta-ExternalAgentMetaTraining28 (30%)
GPTBotOpenAITraining26 (28%)
AmazonbotAmazonOther23 (25%)
Google-ExtendedGoogleTraining23 (25%)
PerplexityBotPerplexitySearch22 (24%)
Claude-SearchBotAnthropicSearch18 (19%)
Claude-UserAnthropicUser-triggered18 (19%)
Perplexity-UserPerplexityUser-triggered18 (19%)
ChatGPT-UserOpenAIUser-triggered15 (16%)
OAI-SearchBotOpenAISearch13 (14%)
BingbotMicrosoftClassic search1 (1%)
GooglebotGoogleClassic search1 (1%)

Googlebot and Bingbot were included as a reference point for classic search. 1 site blocked Googlebot and 1 blocked Bingbot.

llms.txt

29 of 99 sites (29%) serve an llms.txt file. Most are software companies: 24 of the 45 in our software and AI companies group, against 4 in shopping and travel, 1 in other, and none in news and media or social and community.

A caution on interpreting this. llms.txt is a proposal, not a standard, and no major AI search engine promises to read it. Publishing one is not the same as being cited, and it does not change what robots.txt allows. We report the count because it was cheap to check, not because it shows a site is better prepared.

Method

  • The free AI crawler checker read each site’s public robots.txt on 2026-10-06, following the parsing rules in RFC 9309.
  • We tested path / only. A crawler is counted as blocked when the rules that apply to it disallow /. Rules that only cover other paths were not tested.
  • The checker reads robots.txt rules only. A block at a CDN or firewall level is invisible to it, so a site counted as “not blocking” could still refuse these crawlers in practice.
  • Sites can serve different robots.txt files to different visitors, for example by country or by user-agent. We saw one version per site.
  • This is a snapshot. Files change, and some of these may already differ.
  • We chose the sites for popularity and spread across industries, not at random, so the results do not describe the web as a whole.
  • Six sites are excluded from percentages: khanacademy.org returned no robots.txt, and godaddy.com, linkedin.com, mayoclinic.org, stackoverflow.com, udemy.com could not be read reliably.

What to do with this

A robots.txt that blocks training crawlers but allows search crawlers is a choice some sites make: 14 of the 93 readable sites here block GPTBot while allowing OAI-SearchBot. It is one option among several, and this study does not say which option is right for any site.

The practical step is to know what your own file says. Run your domain through the AI crawler checker to see which of the 16 crawlers your robots.txt blocks today. If you want to change it, the robots.txt generator can draft a file, and the robots.txt checker can test specific URLs against your rules. If you use a CDN or firewall, check its bot settings separately, since robots.txt does not show them.

Download the data

Download the data as CSV. Columns: domain, robots_txt (found, missing or unreachable), llms_txt (yes or no), blocked_bots (semicolon-separated).

Appendix: all 99 sites

“Blocked” shows the number of the 16 crawlers blocked, followed by their names. “n/a” means robots.txt could not be evaluated for that site.

Domainrobots.txtBlocked crawlersllms.txt
adobe.comRead0yes
ahrefs.comRead0no
airbnb.comRead3: GPTBot, ClaudeBot, Applebot-Extendedno
amazon.comRead12: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Meta-ExternalAgent, CCBot, Bytespiderno
anthropic.comRead0no
asana.comRead0yes
atlassian.comRead0yes
bbc.comRead12: OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
bestbuy.comRead0no
bloomberg.comRead9: GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
booking.comRead0no
canva.comRead6: GPTBot, ClaudeBot, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespiderno
cloudflare.comRead0yes
cnn.comRead13: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, CCBot, Bytespider, Amazonbotno
copy.aiRead0no
coursera.orgRead1: Meta-ExternalAgentyes
craigslist.orgRead0no
digitalocean.comRead0no
dropbox.comRead0yes
ebay.comRead8: GPTBot, ClaudeBot, PerplexityBot, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
etsy.comRead0yes
expedia.comRead0yes
facebook.comRead8: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, Perplexity-User, Meta-ExternalAgent, CCBot, Bytespiderno
figma.comRead5: GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespiderno
forbes.comRead8: GPTBot, ClaudeBot, PerplexityBot, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
github.comRead1: Bytespideryes
gitlab.comRead0no
glassdoor.comRead7: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Bytespider, Amazonbotno
godaddy.comCould not be readn/ano
grammarly.comRead0no
healthline.comRead7: GPTBot, ClaudeBot, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
heroku.comRead0yes
hubspot.comRead0yes
huggingface.coRead0no
ikea.comRead0no
imdb.comRead13: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespiderno
indeed.comRead0no
instagram.comRead14: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
intercom.comRead0yes
investopedia.comRead11: Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
jasper.aiRead0yes
khanacademy.orgNo robots.txtn/ano
linkedin.comCould not be readn/ano
mailchimp.comRead0no
mayoclinic.orgCould not be readn/ano
medium.comRead6: GPTBot, ClaudeBot, Applebot-Extended, Meta-ExternalAgent, Bytespider, Amazonbotno
mistral.aiRead0yes
monday.comRead0yes
moz.comRead0no
nerdwallet.comRead1: Google-Extendedno
netflix.comRead6: Perplexity-User, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
netlify.comRead0yes
nike.comRead0no
notion.soRead1: Amazonbotyes
nytimes.comRead13: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespiderno
openai.comRead0no
paypal.comRead0yes
perplexity.aiRead0no
pinterest.comRead14: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
quora.comRead9: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Applebot-Extended, Bytespiderno
reddit.comRead16: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Googlebot, Google-Extended, Bingbot, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
reuters.comRead12: GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
revolut.comRead0no
salesforce.comRead0no
semrush.comRead0yes
sendgrid.comRead0no
shopify.comRead0yes
slack.comRead0yes
spotify.comRead0no
squarespace.comRead0yes
stackoverflow.comCould not be readn/ano
stripe.comRead0yes
substack.comRead0no
surferseo.comRead0no
target.comRead0yes
techcrunch.comRead7: ChatGPT-User, GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Bytespiderno
theguardian.comRead9: Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
theverge.comRead11: ChatGPT-User, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespiderno
tiktok.comRead13: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespiderno
tripadvisor.comRead8: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
udemy.comCould not be readn/ano
vercel.comRead0yes
walmart.comRead0no
washingtonpost.comRead7: ClaudeBot, PerplexityBot, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
webflow.comRead0yes
webmd.comRead3: GPTBot, ClaudeBot, CCBotno
wikipedia.orgRead0no
wired.comRead11: Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
wise.comRead0no
wix.comRead0yes
wordpress.comRead0yes
writesonic.comRead0yes
wsj.comRead11: Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
x.comRead14: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
yelp.comRead14: OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
youtube.comRead0no
zalando.comRead8: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider, Amazonbotno
zendesk.comRead0no
zoom.usRead0no
← All articles

Free AI visibility check

Is ChatGPT citing your site?

See if ChatGPT search and Google AI Overviews cite you for your top searches, and who they cite instead. About a minute, no account.

Want content like this on autopilot?

Balzac researches, writes and publishes articles like this one to your site. Every week.

3 free articles · No card needed