# Crawl budget

Crawl budget is how many URLs Googlebot is willing and able to fetch from your site in a given period. For a small site, it rarely limits you. It starts to matter when you run into thousands of URLs or create near-infinite URL variants.

Updated 2026-09-13 · Source: https://porteur.ai/glossary/crawl-budget

## When crawl budget matters

Google balances how fast your server can handle requests with how much demand there is to crawl your URLs. On sites with a few hundred pages, you will not hit the limit.

It becomes a constraint when you have thousands of URLs: faceted filters, calendars, session IDs, infinite pagination, or programmatic pages. Then, Google may not reach or refresh important pages often enough.

> Rule: if your site has under about a thousand indexable URLs, a typical threshold not a measured limit, and loads fast, fix content and links first. Do not spend time “optimising crawl budget”.

## How to check crawl activity

1. **Open the Crawl stats report** In Search Console, go to Settings, then Crawl stats. Check total crawl requests, average response time, and by-URL patterns. As of 2026, this is the best high-level view.
2. **Scan response codes** Look for spikes in 404, soft 404, 301 chains, or 5xx. A week full of 404s or long redirect chains wastes fetches.
3. **Group by URL patterns** Expand Host status and crawl requests by URL pattern. Spot runs like /blog?page=, /collections?sort=, /calendar/archives/day, or ?utm= parameters.
4. **Confirm in server logs** If you can, sample web server logs. Filter user agents that contain Googlebot. Tally hits by path and status. This shows exactly what is fetched and how your server responds.
5. **Watch rendering cost** Heavy JavaScript can slow responses. In Crawl stats, a rising average response time can reduce crawl rate. Keep important pages fast.

A healthy small site shows most crawl requests hitting indexable pages, few errors, and steady response times. Your /pricing, /features, and top articles are fetched regularly.

## What wastes crawl budget

- Endless URL variants: ?sort=, ?color=, ?page=1432, session IDs, or tracking parameters like ?utm_source=.
- Duplicate paths with minor changes: /product, /product/, /product?view=grid, mixed http and https, or both www and non-www.
- Soft 404s and long redirect chains: old URLs that 302 to 301 to 200, or empty category pages that return 200.
- Orphaned archives and infinite calendars: /blog/2014/archive or /events/archive/day that no one needs crawled.
- Large assets on HTML requests: slow TTFB and heavy server processing reduce how much Google pulls in a day.

> Blocking everything in robots.txt does not fix duplicates you actually want indexed. Disallow only low-value patterns after you have a way to reach the important ones.

## What to do on a growing site

- Make a clean XML sitemap: only canonical, indexable URLs. Update it when you publish or remove pages.
- Use rel=canonical on duplicates: for example, point /product?color=red to /product.
- Noindex low-value pages: tag filters, thin archives, internal search. Keep them crawlable until Google sees the noindex, then consider disallow.
- Prune parameters at the source: avoid linking to ?utm= on-site. Strip session IDs. Keep one URL per page in your internal links.
- Fix 404s and redirects: update links to the final URL. Return 410 for dead pages you do not want recrawled.
- Keep pages fast: aim for quick server responses. Large delays reduce how much Google fetches.
- Tighten navigation: link to key listings and products from / and hubs. Remove links to infinite pages like /blog?page=last.

A fixed collection page looks like /collections/shirts with static facets, no crawlable ?sort= permutations, a self-referencing canonical, and it is linked from the main menu.

## “Crawled, currently not indexed” and crawl budget

This status in Search Console is usually about quality or relevance, not just budget. If Google can fetch the page but chooses not to index it, improve the content and links.

If many such pages sit in deep pagination or low-value variants, reduce the bloat with canonicals, noindex, and better internal links to the pages that matter.

## Questions

### What is a crawl budget?

It is the number of URLs Googlebot is willing and able to fetch from your site over time. It depends on your server’s capacity and how much demand Google has to crawl your URLs.

### How do I check my crawl budget?

Use Search Console’s Crawl stats report under Settings. Review total requests, response times, status codes, and URL patterns. For detail, inspect server logs filtered to Googlebot.

### Does crawl budget matter for my small site?

Not until you reach thousands of URLs or create many duplicates with parameters. If you have only hundreds of indexable pages, focus on content, internal links, and speed first.

### Is “Crawled, currently not indexed” a crawl budget issue?

Usually no. It signals quality or duplication. Strengthen the page, consolidate duplicates with canonicals, and link it from strong pages so Google sees it as worth indexing.

### Should I block parameters in robots.txt to save budget?

Disallow low-value patterns only after you have canonical and noindex in place for duplicates. Blocking too early can trap signals on the wrong URLs and make consolidation harder.

## Read next

- [Indexation](https://porteur.ai/glossary/indexation): Indexation means a page is stored in Google’s index and can appear in search. Here is how to check status, read Page indexing, and fix what blocks it.
- [robots.txt](https://porteur.ai/glossary/robots-txt): robots.txt tells crawlers which URLs they may fetch. See what it does not do, how to test it, what to put in it, and how to handle AI bots.
- [XML sitemap](https://porteur.ai/glossary/xml-sitemap): What to put in sitemap.xml, how lastmod works, how to create, validate and submit your XML sitemap, and the traps to avoid.
- [Orphan pages: finding the pages nothing links to](https://porteur.ai/guides/orphan-pages): Find orphan pages on your site, why they happen, how to compare crawl, sitemap and Search Console, and what to do: link, merge, redirect or remove.
- [Canonical URL](https://porteur.ai/glossary/canonical-url): A canonical URL names the original version of a page. Use it to handle parameters and duplicates, avoid split signals and keep the right page indexed.
- [Technical SEO checklist for a small site](https://porteur.ai/guides/technical-seo-checklist): Run these 20 technical SEO checks, in order. Each shows how to check it free and what fixed looks like for a site under 1,000 pages.
- [Crawling](https://porteur.ai/glossary/crawling): Crawling is how Googlebot discovers and fetches your URLs. See what it fetched, fix slow or blocked areas, and know when crawl budget matters.

Paste your homepage URL to see wasted crawl on parameter pages and which sections Google is spending requests on, in about thirty seconds, free. Free check: https://porteur.ai/
