Google crawlers and verifying Googlebot

Google crawlers,
and how to verify them

How Google says to check that a request really comes from Google, and how its crawlers behave. Verified from Google's pages on 2026-10-07.

Google describes its crawlers and fetchers as programs it uses to perform actions for its products, either automatically or triggered by user request. This page lists what Google documents about telling them apart, verifying a request in your logs, and the technical properties of its crawlers, with the wording taken from Google's pages.

What Google documents

TopicWhat Google saysSource
Why verify “You can verify if a request to your server really is from Google.”Google: this is useful if you are concerned that spammers or other troublemakers are accessing your site while claiming to be from Google. Google docs
Verified 2026-10-07
Three categories “Google's crawlers and fetchers fall into three categories:”Google names them common crawlers (such as Googlebot), special-case crawlers (such as AdsBot) and user-triggered fetchers (such as Google Site Verifier). Google docs
Verified 2026-10-07
Common crawlers “They always respect robots.txt rules for automatic crawls.”Google's description: the common crawlers used for Google's products, such as Googlebot. Google docs
Verified 2026-10-07
Special-case crawlers “These crawlers or fetchers may or may not respect robots.txt rules.”Google's overview gives an example: AdsBot ignores the global robots.txt user agent (*) with the ad publisher's permission. Google docs
Verified 2026-10-07
User-triggered fetchers “Because the fetch was requested by a user, these fetchers ignore robots.txt rules.” Google docs
Verified 2026-10-07
How a crawler identifies itself “Google's crawlers identify themselves in three ways:”Google lists them as the HTTP user-agent request header, the source IP address of the request, and the reverse DNS hostname of the source IP. Google docs
Verified 2026-10-07
Two methods “Manually: For one-off lookups, use command line tools. This method is sufficient for most use cases.”Google's other method: Automatically, for large scale lookups. Google docs
Verified 2026-10-07
Step 1: reverse DNS “Run a reverse DNS lookup on the accessing IP address from your logs, using the host command.” Google docs
Verified 2026-10-07
Step 2: check the domain “Verify that the domain name is either googlebot.com, google.com, or googleusercontent.com.” Google docs
Verified 2026-10-07
Step 3: forward DNS “Run a forward DNS lookup on the domain name retrieved in step 1 using the host command on the retrieved domain name.” Google docs
Verified 2026-10-07
Step 4: compare “Verify that it's the same as the original accessing IP address from your logs.” Google docs
Verified 2026-10-07
Automatic matching “Automatically: For large scale lookups, use an automatic solution to match a crawler's IP address against the list of published Google IP addresses.”Google publishes separate JSON lists for common crawlers, special crawlers, user-triggered fetchers and user-triggered agents. Google docs
Verified 2026-10-07
CIDR format “Note that the IP addresses in the JSON files are represented in CIDR format.” Google docs
Verified 2026-10-07
Other Google IPs “For other Google IP addresses from where your site may be accessed (for example, Apps Scripts), match the accessing IP address against the general list of Google IP addresses.” Google docs
Verified 2026-10-07
Many IP addresses “Therefore, your logs may show visits from several IP addresses.”Google: its clients are distributed across many datacenters across the world, so they are located near the sites they might access. Google docs
Verified 2026-10-07
Where requests come from “Google egresses primarily from IP addresses in the United States.”Google: if it detects that a site is blocking requests from the United States, it may attempt to crawl from IP addresses located in other countries. Google docs
Verified 2026-10-07
File size “By default, Google's crawlers and fetchers only crawl the first 15MB of a file, and any content beyond this limit is ignored.”Google: individual projects may set different limits, and a crawler like Googlebot may have a smaller limit (for example, 2MB). Google docs
Verified 2026-10-07
Protocols “The default protocol version used by Google's crawlers is HTTP/1.1”Google: its crawlers support HTTP/1.1 and HTTP/2, and crawling over HTTP/2 brings no Google-product specific benefit to the site, such as a ranking boost in Google Search. Google docs
Verified 2026-10-07
Opting out of HTTP/2 “To opt out from crawling over HTTP/2, instruct the server that's hosting your site to respond with a 421 HTTP status code when Google attempts to access your site over HTTP/2.” Google docs
Verified 2026-10-07
Content encodings “Google's crawlers and fetchers support the following content encodings (compressions): gzip, deflate, and Brotli (br).” Google docs
Verified 2026-10-07
Crawl rate “If your site is having trouble keeping up with Google's crawling requests, you can reduce the crawl rate.”Google also notes that sending the inappropriate HTTP response code to its crawlers may affect how your site appears in Google products. Google docs
Verified 2026-10-07
ETag and Last-Modified “If both ETag and Last-Modified response header fields are present in the HTTP response, Google's crawlers use the ETag value as required by the HTTP standard.”Google: other HTTP caching directives aren't supported, and individual crawlers may or may not make use of caching. Google docs
Verified 2026-10-07

Read on 2026-10-07. Google's verification page showed "Last updated 2026-03-20 UTC" and its crawler overview showed "Last updated 2026-06-12 UTC". Both pages are now served under developers.google.com/crawling/docs/crawlers-fetchers/ (the older /search/docs/crawling-indexing/ addresses redirect there). Notes under a quote are close paraphrases of the same page. This page covers only what Google documents.

Examples

Google's own first example of the manual check, run on an IP address from a log. The reverse lookup returns a googlebot.com name, and the forward lookup on that name returns the original IP address.

host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1

Check which AI bots your robots.txt lets in with the free AI crawler checker and read the live file with the robots.txt and sitemap checker. The AI crawler checker reads your robots.txt rules; it does not verify the IP address of a request. To read what the requests claiming to be Googlebot got back from your server, paste access-log lines into the Googlebot log analyzer (claimed hits only, not verified). For the user agents of AI companies, see AI crawler user agents; for the rules Google applies to robots.txt, see Google robots.txt rules.

Rules to remember

  • Google says the manual check is a reverse DNS lookup on the IP address from your logs, a check of the domain name, then a forward DNS lookup that must return the same IP address.
  • Google says common crawlers always respect robots.txt rules for automatic crawls, while special-case crawlers may or may not and user-triggered fetchers ignore them.
  • Google says that for large scale lookups you match the IP address against its published lists, which are in CIDR format.

Crawling at scale on big sites: Google crawl budget.

How Google treats response codes: Google HTTP status codes guide.

The robots.txt matching rules: Google robots.txt rules.

Questions

How do I check that a request really comes from Googlebot?

Google's manual method has four steps: run a reverse DNS lookup on the IP address from your logs, verify the domain name is googlebot.com, google.com or googleusercontent.com, run a forward DNS lookup on that name, and verify it returns the original IP address. Google says this method is sufficient for most use cases.

Is the user agent string enough to identify Googlebot?

Google says its crawlers identify themselves in three ways: the user-agent request header, the source IP address and the reverse DNS hostname of that IP. Its verification page exists because spammers may claim to be from Google, so the DNS or IP check is what confirms a request.

How do I verify at scale?

Google says to use an automatic solution that matches the crawler's IP address against its published lists of IP ranges (JSON files in CIDR format), with separate lists for common crawlers, special crawlers, user-triggered fetchers and user-triggered agents.

Do all Google crawlers obey robots.txt?

No. Google says common crawlers such as Googlebot always respect robots.txt rules for automatic crawls, special-case crawlers such as AdsBot may or may not, and user-triggered fetchers ignore robots.txt because the fetch was requested by a user.

How much of a page does Google crawl?

Google says that by default its crawlers and fetchers only crawl the first 15MB of a file and ignore anything beyond that, but individual projects may set different limits (a crawler like Googlebot may have a smaller one, for example 2MB).

How current is this page?

Every row was read from Google's pages on 2026-10-07. The verification page showed a last-updated date of 2026-03-20 UTC and the overview page showed 2026-06-12 UTC. Google can change its documentation, so treat the linked pages as the final word.

Sources

Let Balzac write the articles.

Type your website. Balzac finds the searches you can win, writes articles built to rank on Google and get cited by ChatGPT, and publishes them for you.

3 free articles · No card needed