Google robots.txt rules

Google robots.txt rules,
as documented

Where robots.txt must live, how errors are handled, which fields Google supports and how groups and wildcards match, in Google's own words. Verified from Google's pages on 2026-10-07.

Google says a robots.txt file is a text file containing rules about which crawlers may access which parts of a site. This page lists what Google documents about how it reads robots.txt, with the wording taken from its page.

What Google documents

TopicWhat Google saysSource
What it is “A robots.txt file is a text file containing rules about which crawlers may access which parts of a site.” Google docs
Verified 2026-10-07
Location “You must place the robots.txt file in the top-level directory of a site, on a supported protocol.” Google docs
Verified 2026-10-07
Subdirectories “Crawlers don't check for robots.txt files in subdirectories.” Google docs
Verified 2026-10-07
Scope of validity “The rules listed in the robots.txt file apply only to the host, protocol, and port number where the robots.txt file is hosted.” Google docs
Verified 2026-10-07
Case “The URL for the robots.txt file is (like other URLs) case-sensitive.” Google docs
Verified 2026-10-07
2xx responses “HTTP status codes that signal success prompt Google's crawlers to process the robots.txt file as provided by the server.” Google docs
Verified 2026-10-07
Redirects “Google follows at least five redirect hops as defined by RFC 1945 and then stops and treats it as a 404 for the robots.txt file.” Google docs
Verified 2026-10-07
Logical redirects “Google doesn't follow logical redirects in robots.txt files (frames, JavaScript, or meta refresh-type redirects).” Google docs
Verified 2026-10-07
4xx responses “Google's crawlers treat all 4xx errors, except 429 , as if a valid robots.txt file didn't exist. This means that Google assumes that there are no crawl restrictions.” Google docs
Verified 2026-10-07
401 and 403 “Don't use 401 and 403 status codes for limiting the crawl rate.” Google docs
Verified 2026-10-07
5xx responses “For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file.”Google: if it cannot fetch a new version, for the next 30 days it uses the last good version while still trying to fetch a new one. If there is no cached version it assumes there are no crawl restrictions. Google docs
Verified 2026-10-07
Caching “Google generally caches the contents of robots.txt file for up to 24 hours” Google docs
Verified 2026-10-07
File format “The robots.txt file must be a UTF-8 encoded plain text file and the lines must be separated by CR , CR/LF , or LF .” Google docs
Verified 2026-10-07
Size limit “Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored.” Google docs
Verified 2026-10-07
Syntax “A valid robots.txt line consists of a field, a colon, and a value.” Google docs
Verified 2026-10-07
Supported fields “Google supports the following fields (other fields such as crawl-delay aren't supported)”Google lists user-agent, allow, disallow and sitemap. Google docs
Verified 2026-10-07
Paths are case-sensitive “The path value must start with / to designate the root and the value is case-sensitive.” Google docs
Verified 2026-10-07
Disallow does not block indexing “Google can't index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet.” Google docs
Verified 2026-10-07
Sitemap field “The sitemap field isn't tied to any specific user agent and may be followed by all crawlers, provided it isn't disallowed for crawling.” Google docs
Verified 2026-10-07
One group per crawler “Only one group is valid for a particular crawler.” Google docs
Verified 2026-10-07
Most specific group “Google's crawlers determine the correct group of rules by finding in the robots.txt file the group with the most specific user agent that matches the crawler's user agent.” Google docs
Verified 2026-10-07
Group order “The order of the groups within the robots.txt file is irrelevant.” Google docs
Verified 2026-10-07
Specific and global groups “User agent specific groups and global groups ( * ) are not combined.” Google docs
Verified 2026-10-07
Wildcards “* designates 0 or more instances of any valid character. $ designates the end of the URL.” Google docs
Verified 2026-10-07
Path matching “Matches any path that starts with /fish . Note that the matching is case-sensitive.” Google docs
Verified 2026-10-07

Read on 2026-10-07. Google's page showed "Last updated 2026-08-31 UTC". Notes under a quote are close paraphrases of the same page. This page covers only what Google documents; other search engines may differ.

Examples

Google's own example: all crawlers are kept out of a directory that Googlebot needs for rendering, plus a sitemap line. Then its grouping example, where googlebot-news merges its two groups and the global group stays separate.

# This robots.txt file controls crawling of URLs under https://example.com.
# All crawlers are disallowed to crawl files in the "includes" directory, such
# as .css, .js, but Google needs them for rendering, so Googlebot is allowed
# to crawl them.
User-agent: *
Disallow: /includes/

User-agent: Googlebot
Allow: /includes/

Sitemap: https://example.com/sitemap.xml

# Grouping: googlebot-news gets /fish and /shrimp, everyone else /carrots
user-agent: googlebot-news
disallow: /fish

user-agent: *
disallow: /carrots

user-agent: googlebot-news
disallow: /shrimp

Status codes beyond robots.txt: Google HTTP status codes guide.

Test a file with the free robots.txt checker (live site) or the robots.txt tester (pasted text, any URL), compare two versions with the robots.txt diff checker, write one with the robots.txt generator, and see AI crawler user agents, Google's sitemap rules and Google's robots meta tag directives. To verify that a request really is from Googlebot, see Google crawlers and verifying Googlebot.

Rules to remember

  • Google says robots.txt must sit at the top level of the host and applies only to that host, protocol and port.
  • Google says it treats 4xx responses (except 429) as no robots.txt, and handles 5xx responses with a 12-hour stop and a 30-day fallback to the last good copy.
  • Google says Disallow controls crawling, not indexing: the URL may still be indexed without a snippet.

How Google describes crawl budget for very large sites: Google crawl budget.

Search Essentials technical requirements: Google Search Essentials, as documented.

Questions

Where does robots.txt have to be?

Google says it must be in the top-level directory of a site on a supported protocol, and that crawlers don't check for robots.txt files in subdirectories. The rules apply only to the host, protocol and port number where the file is hosted.

What happens if my robots.txt returns a 404 or 403?

Google says it treats all 4xx errors except 429 as if a valid robots.txt file didn't exist, so it assumes there are no crawl restrictions. It also says not to use 401 and 403 to limit the crawl rate.

What happens if robots.txt returns a 5xx error?

Google says that for the first 12 hours it stops crawling the site but keeps trying to fetch the file, then for the next 30 days it uses the last good version while still trying to fetch a new one. With no cached version it assumes there are no crawl restrictions.

Does Google support crawl-delay?

No. Google says it supports the user-agent, allow, disallow and sitemap fields and that other fields such as crawl-delay aren't supported.

How big can a robots.txt file be?

Google enforces a limit of 500 kibibytes (KiB) and ignores content after it. It suggests consolidating rules or placing excluded material in a separate directory.

Does Disallow keep a page out of Google?

Not by itself. Google says it can't index the content of pages disallowed for crawling, but it may still index the URL and show it in search results without a snippet. For blocking indexing it points to its separate page on how to block indexing.

Which user-agent group does Google follow?

Google says only one group is valid for a particular crawler: the group with the most specific user agent that matches. The order of groups doesn't matter, and groups for a specific user agent and the global (*) group are not combined.

Can I use wildcards in paths?

Yes, a limited form. Google says * designates 0 or more instances of any valid character and $ designates the end of the URL, and that path values are case-sensitive.

How current is this page?

Every row was read from Google's page on 2026-10-07, which showed "Last updated 2026-08-31 UTC". Google can change its documentation, so treat the linked page as the final word.

Sources

Let Balzac write the articles.

Type your website. Balzac finds the searches you can win, writes articles built to rank on Google and get cited by ChatGPT, and publishes them for you.

3 free articles · No card needed