Google crawl budget
Google crawl budget,
as documented
Who needs to worry about crawl budget, what sets it, and the practices Google lists, in Google's own words. Verified from Google's page on 2026-10-07.
Google's crawl budget guide is written for very large and frequently updated sites. Google says that if your site does not have a large number of pages that change rapidly, or your pages seem to be crawled the same day they are published, you do not need to read it. This page lists what the guide says, with quotes taken from Google's page.
What Google says about crawl budget
| Topic | What Google says | Source |
|---|---|---|
| Who it is for | “If your site doesn't have a large number of pages that change rapidly, or if your pages seem to be crawled the same day that they are published, you don't need to read this guide.”Google: for Google Search specifically, keeping your sitemap up to date and checking the Page Indexing report regularly is adequate. | Google docs Verified 2026-10-07 |
| Size thresholds | “The numbers given here are a rough estimate to help you classify your site. These are not exact thresholds.”Google's examples: 1 million+ unique pages that change about weekly, 10,000+ unique pages that change daily, or a large portion of URLs classified in Search Console as Discovered - currently not indexed. | Google docs Verified 2026-10-07 |
| What a site is | “Google's crawling infrastructure defines a site as a unique hostname.”Google: https://www.example.com/ and https://code.example.com/ are treated as separate sites with separate crawl budgets. | Google docs Verified 2026-10-07 |
| Crawl budget defined | “the set of URLs that Google can and wants to crawl.”Google: this combines the crawl capacity limit and crawl demand. | Google docs Verified 2026-10-07 |
| Crawled is not indexed | “not every page that is crawled will necessarily be indexed.”Google: after crawling, each page must be evaluated, consolidated, and assessed to determine its suitability for the index. | Google docs Verified 2026-10-07 |
| Crawl capacity limit | “Google wants to crawl your site without overwhelming your servers.”Google: the limit (also known as hostload) covers the time your server spends holding connections open for Google, and every site starts with the same conservative default that can adjust over time. | Google docs Verified 2026-10-07 |
| Crawl health | “If the site slows down (latency increases or response times become longer), or responds with server errors (5xx HTTP status codes) or rate-limiting signals (such as HTTP 429), the limit goes down and Google crawls less.”Google: if the site responds consistently and response times stay stable or improve, the limit goes up. | Google docs Verified 2026-10-07 |
| Crawl demand | “For Googlebot, demand varies based on a site's size, update frequency, page quality, and relevance, compared to other sites.”Google's factors you can influence: perceived inventory, popularity and staleness. | Google docs Verified 2026-10-07 |
| Perceived inventory | “This is the factor that you can positively control the most.”Google: without guidance, it tries to crawl all or most of the URLs it knows about, so many duplicate or unwanted URLs waste crawling time. | Google docs Verified 2026-10-07 |
| Site moves | “site-wide events like site moves may trigger an increase in crawl demand in order to reprocess the content under the new URLs.”Google: the crawl capacity limit is shared across all crawlers, so high demand from one crawler can reduce capacity for others. | Google docs Verified 2026-10-07 |
| Consolidate duplicates | “Eliminate duplicate content to focus crawling on unique content rather than unique URLs.”Part of Google's "manage your URL inventory" practice. | Google docs Verified 2026-10-07 |
| Block with robots.txt | “Blocking URLs with robots.txt prevents Google from crawling them, and significantly decreases the chance the URLs will be processed by other Google systems (such as getting indexed by Google Search).”Google's examples: infinite scrolling pages that duplicate information on linked pages, or differently sorted versions of the same page, when you cannot consolidate them. | Google docs Verified 2026-10-07 |
| Not noindex | “Google will still request, but then drop the page when it sees a noindex meta tag or header in the HTTP response, wasting crawling time.”Google: do not use noindex to manage crawl budget. | Google docs Verified 2026-10-07 |
| robots.txt is not a temporary reallocation | “Don't use robots.txt to temporarily reallocate crawl budget for other pages”Google: use it for pages or resources you do not want crawled at all. It will not shift the freed budget to other pages unless Google is already hitting your site's crawl capacity limit. | Google docs Verified 2026-10-07 |
| Removed pages | “a 404 status code is a strong signal not to crawl that URL again.”Google: for permanently removed pages return 404 or 410. Blocked URLs stay in the crawl queue much longer and are recrawled when the block is removed. | Google docs Verified 2026-10-07 |
| Soft 404s | “soft 404 pages will continue to be crawled, and waste your budget.”Google: check the Page Indexing report for soft 404 errors. | Google docs Verified 2026-10-07 |
| Sitemaps | “Google reads your sitemap regularly, so be sure to include all the content that you want Google to crawl.”Google: if your site includes updated content, it recommends including the <lastmod> tag. | Google docs Verified 2026-10-07 |
| Redirect chains | “Avoid long redirect chains”Google: they have a negative effect on crawling. | Google docs Verified 2026-10-07 |
| Page efficiency | “If Google can load and render your pages faster, we might be able to read more content from your site.”Google: optimize server response times and resources. | Google docs Verified 2026-10-07 |
| HTTP caching | “Support 304 (Not Modified) HTTP status codes.”Google: if a page has not changed since the last crawl, a 304 tells Google to reuse the cached version, saving bandwidth and resources. | Google docs Verified 2026-10-07 |
| More server resources | “If your site can't be crawled because of server capacity on your end (for example, you're getting Hostload exceeded in the URL inspection tool), add more server resources if that makes sense for your business.”One of Google's two ways to increase crawl budget. | Google docs Verified 2026-10-07 |
| Content quality | “Optimize your content's quality for the Google product you're targeting”Google: for Google Search, the factors it names are popularity, overall user value, content uniqueness, and serving capacity. | Google docs Verified 2026-10-07 |
How to apply it
A checklist built only from the practices in Google's guide.
- Decide whether it applies: compare your site with the size and change-rate examples above, and if it does not fit, Google's own advice is a current sitemap and a regular look at the Page Indexing report.
- Find wasted crawling: look for duplicate URLs, differently sorted pages and other URLs you do not want crawled, and check the Page Indexing report for soft 404 errors.
- Consolidate duplicates where you can; where you cannot, block unimportant URLs with robots.txt rather than noindex.
- Return 404 or 410 for permanently removed pages instead of blocking them.
- Keep sitemaps current with all the content you want crawled, and add
for updated content. - Shorten redirect chains, and make pages quicker to load, including support for 304 responses.
- Check your site's availability during crawling, and add server resources if server capacity is the limit.
Questions
Do I need to worry about crawl budget?
Google says that if your site does not have a large number of pages that change rapidly, or your pages seem to be crawled the same day they are published, you do not need to read its guide. It describes the size examples as rough estimates, not exact thresholds.
What decides my crawl budget?
Google defines it as the set of URLs that Google can and wants to crawl, shaped by two elements: the crawl capacity limit and crawl demand.
Should I use noindex to save crawl budget?
Google says no. It will still request the page and then drop it when it sees the noindex, which wastes crawling time. It points to robots.txt for pages you do not want crawled at all, and 404 or 410 for permanently removed pages.
Does blocking pages in robots.txt give other pages more budget?
Google says it will not shift the newly available crawl budget to other pages unless it is already hitting your site's crawl capacity limit.
How do I get more crawl budget?
Google lists two ways: add more server resources if server capacity is stopping your site being crawled, and optimize your content's quality for the Google product you are targeting.
Does a crawled page get indexed?
Not necessarily. Google says that after crawling, each page must be evaluated, consolidated, and assessed to determine its suitability for the index.
How current is this page?
Every quote was read from Google's page on 2026-10-07, which showed "Last updated 2026-07-22 UTC". Google can change its documentation, so treat the linked page as the final word.
Sources
- Managing crawl budget for large sites (Google Search Central), read on 2026-10-07
Let Balzac write the articles.
Type your website. Balzac finds the searches you can win, writes articles built to rank on Google and get cited by ChatGPT, and publishes them for you.
3 free articles · No card needed