How Many Popular Sites List a Sitemap in robots.txt? We Checked 99 (Oct 2026)
A Sitemap: line in robots.txt is the standard way to tell every crawler where your sitemap lives. We checked how many popular sites do it, and whether the sitemaps they point to actually load. This is a snapshot from 2026-10-06 of what the files say, nothing more.
Headline findings
We used the same 99 sites as our AI crawler study. Seven of them had no readable robots.txt in this run (details in the method), so the percentages below are of the 92 sites whose robots.txt we could read.
- 71 of 92 sites (77%) list at least one sitemap in robots.txt. 21 do not.
- 64 of those 71 (90%) point to at least one sitemap that loads. The other 7 (forbes.com, instagram.com, pinterest.com, revolut.com, tripadvisor.com, wikipedia.org, x.com) list sitemaps that our checker could not load or parse. Some of that may be bot protection on their side rather than a broken file, so read it as “not loadable from our checker”, not as “broken”.
- 38 of the 71 list more than one sitemap line. The most any site lists is 434.
- 7 of the 21 sites without a
Sitemap:line still serve a sitemap at the default/sitemap.xml. The checker looks there when robots.txt names none. The other 14 (amazon.com, craigslist.org, etsy.com, expedia.com, facebook.com, github.com, glassdoor.com, imdb.com, indeed.com, quora.com, reddit.com, webmd.com, yelp.com, zalando.com) have neither a listed sitemap nor one at that default path. They may publish sitemaps elsewhere, for example through Search Console only, which we cannot see. - Across all sitemap URLs checked: 117 loaded, 10 were unreachable, 4 returned not found and 3 could not be parsed.
Why it matters
Google and Bing do not need a robots.txt line to find a sitemap you submit in their webmaster tools. The line matters for everything else: smaller search engines, AI crawlers and any tool that reads robots.txt first. It is one line, it costs nothing, and the common failure is a sitemap line that points to a URL that no longer exists. Of the 4 sitemap URLs that returned not found here, 3 were seasonal pages on one site, a reminder to remove entries when you retire them.
What to do with this
- Open your own robots.txt and check there is a
Sitemap:line with the full, absolute URL. - Open each listed URL in a browser and make sure it loads and lists your current pages.
- Run your domain through the free robots.txt and sitemap checker, which reads the file and loads every sitemap for you, and the AI crawler checker to see which AI bots your robots.txt allows.
- If you need a robots.txt, the robots.txt generator builds one with your sitemap line included.
Method
- The free robots.txt and sitemap checker read each site’s public robots.txt on 2026-10-06 and loaded every sitemap URL listed in it, plus the default
/sitemap.xmlwhen none was listed. - A sitemap counts as loaded when it returned a valid sitemap or sitemap index. “Unreachable” means our checker got no usable answer, which includes bot walls. “Invalid” means it answered but we could not parse it.
- We counted lines that start with
Sitemap:. We did not follow sitemaps listed inside other files beyond the index level the checker reads. - Seven sites are excluded from the percentages: khanacademy.org and linkedin.com returned no robots.txt in this run, godaddy.com, mayoclinic.org, stackoverflow.com and udemy.com could not be reached, and washingtonpost.com’s request failed in this run. washingtonpost.com was readable in the earlier crawler study, which shows how much a single request can differ.
- We chose the sites for popularity and spread across industries, not at random, so the results do not describe the web as a whole. This is one request per site at one moment.
Download the data
Download the data as CSV. Columns: domain, robots_txt (found, missing, unreachable or not_checked), sitemap_lines_in_robots, sitemaps_loaded_ok, sitemaps_not_loaded.