Crawl budget

Crawl budget is how many URLs Googlebot is willing and able to fetch from your site in a given period. For a small site, it rarely limits you. It starts to matter when you run into thousands of URLs or create near-infinite URL variants.

By , founder of Porteur · Updated 13 September 2026 · Markdown

When crawl budget matters

Google balances how fast your server can handle requests with how much demand there is to crawl your URLs. On sites with a few hundred pages, you will not hit the limit.

It becomes a constraint when you have thousands of URLs: faceted filters, calendars, session IDs, infinite pagination, or programmatic pages. Then, Google may not reach or refresh important pages often enough.

How to check crawl activity

  1. Open the Crawl stats report

    In Search Console, go to Settings, then Crawl stats. Check total crawl requests, average response time, and by-URL patterns. As of 2026, this is the best high-level view.

  2. Scan response codes

    Look for spikes in 404, soft 404, 301 chains, or 5xx. A week full of 404s or long redirect chains wastes fetches.

  3. Group by URL patterns

    Expand Host status and crawl requests by URL pattern. Spot runs like /blog?page=, /collections?sort=, /calendar/archives/day, or ?utm= parameters.

  4. Confirm in server logs

    If you can, sample web server logs. Filter user agents that contain Googlebot. Tally hits by path and status. This shows exactly what is fetched and how your server responds.

  5. Watch rendering cost

    Heavy JavaScript can slow responses. In Crawl stats, a rising average response time can reduce crawl rate. Keep important pages fast.

A healthy small site shows most crawl requests hitting indexable pages, few errors, and steady response times. Your /pricing, /features, and top articles are fetched regularly.

What wastes crawl budget

  • Endless URL variants: ?sort=, ?color=, ?page=1432, session IDs, or tracking parameters like ?utm_source=.
  • Duplicate paths with minor changes: /product, /product/, /product?view=grid, mixed http and https, or both www and non-www.
  • Soft 404s and long redirect chains: old URLs that 302 to 301 to 200, or empty category pages that return 200.
  • Orphaned archives and infinite calendars: /blog/2014/archive or /events/archive/day that no one needs crawled.
  • Large assets on HTML requests: slow TTFB and heavy server processing reduce how much Google pulls in a day.

What to do on a growing site

  • Make a clean XML sitemap: only canonical, indexable URLs. Update it when you publish or remove pages.
  • Use rel=canonical on duplicates: for example, point /product?color=red to /product.
  • Noindex low-value pages: tag filters, thin archives, internal search. Keep them crawlable until Google sees the noindex, then consider disallow.
  • Prune parameters at the source: avoid linking to ?utm= on-site. Strip session IDs. Keep one URL per page in your internal links.
  • Fix 404s and redirects: update links to the final URL. Return 410 for dead pages you do not want recrawled.
  • Keep pages fast: aim for quick server responses. Large delays reduce how much Google fetches.
  • Tighten navigation: link to key listings and products from / and hubs. Remove links to infinite pages like /blog?page=last.

A fixed collection page looks like /collections/shirts with static facets, no crawlable ?sort= permutations, a self-referencing canonical, and it is linked from the main menu.

“Crawled, currently not indexed” and crawl budget

This status in Search Console is usually about quality or relevance, not just budget. If Google can fetch the page but chooses not to index it, improve the content and links.

If many such pages sit in deep pagination or low-value variants, reduce the bloat with canonicals, noindex, and better internal links to the pages that matter.

Questions

Sources

Check my site, free

Paste your homepage URL to see wasted crawl on parameter pages and which sections Google is spending requests on, in about thirty seconds, free.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next