Crawl waste on a small site: the fetches your real pages never get

Crawl budget matters past tens of thousands of URLs, which is not you. What matters at your size is where the fetches go: if most of them land on parameters, redirects and dead pages, your new guide waits a week to be seen.

By , founder of Porteur · Updated 15 September 2026 · Markdown

Crawl budget is not your problem

Crawl budget is a real constraint on sites with tens of thousands of URLs. On a site with two hundred pages, Google will happily crawl everything you have.

  • Googlebot discovers URLs from links, sitemaps and past crawls, and fetches them within what your server can take.
  • It slows down when the server is slow or returns errors, and speeds up when it is not.
  • Nothing you can set makes it crawl more. There is no dial.
  • So the question is not how many fetches you get. It is what they are spent on.

What crawl waste is

Crawl waste is every fetch that lands on a URL you do not want indexed, do not maintain, or removed months ago. Each one is a fetch your real pages did not get.

The symptoms are recognisable: a new page takes days to appear, a change to an old page takes a fortnight to show in the result, and the Crawl stats report shows plenty of activity that never touches the pages you care about.

Where it comes from on a small site

SourceWhat it looks likeWhy it happens
Query parameters/pricing?ref=twitter, /guides?utm_source=newsletterCampaign links get shared, indexed and crawled as separate URLs
Filters and sorts/tools?sort=name&view=gridEvery combination is a URL, and they multiply
Session and tracking ids/page?sid=8fe2...An old framework habit that creates a URL per visitor
Old redirectsFetches of URLs that 301 somewhere elseInternal links and the sitemap still point at the old address
Dead URLsRepeated 404s on paths deleted last yearThey are still linked from your own pages or from a stale sitemap
Calendars and archives/events/2019/07/, /blog/page/12/Generated pages that go on forever and hold nothing
Tag and author archives/tags/seo/, /author/founder/A CMS creates them by default, often with one item each
Staging and preview URLsA preview domain that got linked onceIt is crawlable and nobody told it not to be

Notice how many of these you created for a reason that has nothing to do with search, and then forgot.

How to see it in the Crawl stats report

  1. Open Settings, then Crawl stats

    It is available for root-level properties and covers the last ninety days: total requests, download size and average response time.

  2. Read the breakdown by response

    A healthy small site is mostly 200 with some 301 and 304. Lots of 404 or 5xx is the first thing to fix.

  3. Read the breakdown by purpose

    Discovery against Refresh. If almost everything is refresh and your new pages are slow to appear, discovery is starved.

  4. Read the breakdown by file type

    If HTML is a small share of the fetches, something else is eating them: images, scripts, or a feed that changes constantly.

  5. Check the host status

    robots.txt fetch, DNS resolution and server connectivity. A red one means Google slows or stops crawling everything.

  6. Sample the example URLs

    The report shows what was actually fetched. Read twenty and count how many you would want indexed.

That last step is the test. If five of twenty fetched URLs are parameters or dead pages, you have found the work.

The fix for each source

SourceFixNote
Campaign parametersA self-referencing canonical on the clean URL, and never link internally with parametersThe URL Parameters tool was retired in 2022. The canonical is the control
Filters and sortsDecide which combinations have search demand; the rest are disallowed in robots.txt or noindexedDo not do both to the same URL: a blocked page cannot deliver a noindex
Session idsRemove them from URLs entirelyA cookie does the job without creating pages
Old redirectsUpdate internal links and the sitemap to the final URLKeep the redirects themselves. It is the links through them that waste fetches
Dead URLsReturn 404 or 410, remove them from the sitemap and from internal linksA 404 nobody links to costs almost nothing
Calendars and archivesNoindex the empty ones or remove the templateKeep them for people if they help; keep them out of the sitemap
Tag and author archivesNoindex thin ones, or curate a few into real hubsOne tag page with one post is not a page
Staging and previewsBlock with authentication, not only with robots.txtA blocked URL can still be indexed as a bare link if someone links to it

What not to do

  • Do not disallow everything you do not like in robots.txt. A disallowed URL can still be indexed without its content when other sites link to it.
  • Do not disallow a URL you have noindexed. Google has to fetch the page to see the noindex, so blocking it keeps it in the index.
  • Do not delete old redirects to save crawl. You would be throwing away the links that point at the old URLs.
  • Do not add a crawl-delay directive expecting Google to obey. Google ignores it.
  • Do not chase crawl waste before the basics. An unindexable page wastes more than a parameter ever will.

The one-hour audit

  1. Crawl your own site from the home page

    List every internal link that returns a redirect or a 404. Those are yours to fix today.

  2. Open the sitemap and check a sample

    Every URL in it should answer 200 and be canonical to itself. Remove redirects and dead URLs.

  3. Search for your own parameters

    A site: search with inurl: and a parameter name shows what Google has indexed of them.

  4. Read twenty fetched URLs in Crawl stats

    Count how many you want indexed. That ratio is your waste, measured.

  5. Fix the largest source, then stop

    One source per session. Re-read the report in a month rather than tuning weekly.

What good looks like

  • Most fetches are 200 on HTML you would want indexed.
  • The sitemap lists only canonical URLs that answer 200.
  • No internal link passes through a redirect.
  • New pages appear in the index within days, not weeks.
  • The Page indexing report's Not indexed reasons are the ones you chose: alternate pages with a canonical, deliberate noindex, redirects.

At that point crawling is not something you manage. It is something that happens correctly while you write the next page.

Questions

Check my site, free

Paste your URL and the free check reads your site the way a crawler does, in about thirty seconds, and shows the dead ends and redirects it walked into.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next