# Crawl waste on a small site: the fetches your real pages never get

Crawl budget matters past tens of thousands of URLs, which is not you. What matters at your size is where the fetches go: if most of them land on parameters, redirects and dead pages, your new guide waits a week to be seen.

Updated 2026-09-15 · Source: https://porteur.ai/guides/crawl-waste

## Crawl budget is not your problem

Crawl budget is a real constraint on sites with tens of thousands of URLs. On a site with two hundred pages, Google will happily crawl everything you have.

- Googlebot discovers URLs from links, sitemaps and past crawls, and fetches them within what your server can take.
- It slows down when the server is slow or returns errors, and speeds up when it is not.
- Nothing you can set makes it crawl more. There is no dial.
- So the question is not how many fetches you get. It is what they are spent on.

> If you are reading crawl budget advice written for a retailer with a million URLs, stop. The small-site version of that problem is waste, and it has different causes.

## What crawl waste is

Crawl waste is every fetch that lands on a URL you do not want indexed, do not maintain, or removed months ago. Each one is a fetch your real pages did not get.

The symptoms are recognisable: a new page takes days to appear, a change to an old page takes a fortnight to show in the result, and the Crawl stats report shows plenty of activity that never touches the pages you care about.

## Where it comes from on a small site

| Source | What it looks like | Why it happens |
| --- | --- | --- |
| Query parameters | /pricing?ref=twitter, /guides?utm_source=newsletter | Campaign links get shared, indexed and crawled as separate URLs |
| Filters and sorts | /tools?sort=name&view=grid | Every combination is a URL, and they multiply |
| Session and tracking ids | /page?sid=8fe2... | An old framework habit that creates a URL per visitor |
| Old redirects | Fetches of URLs that 301 somewhere else | Internal links and the sitemap still point at the old address |
| Dead URLs | Repeated 404s on paths deleted last year | They are still linked from your own pages or from a stale sitemap |
| Calendars and archives | /events/2019/07/, /blog/page/12/ | Generated pages that go on forever and hold nothing |
| Tag and author archives | /tags/seo/, /author/founder/ | A CMS creates them by default, often with one item each |
| Staging and preview URLs | A preview domain that got linked once | It is crawlable and nobody told it not to be |

Notice how many of these you created for a reason that has nothing to do with search, and then forgot.

## How to see it in the Crawl stats report

1. **Open Settings, then Crawl stats** It is available for root-level properties and covers the last ninety days: total requests, download size and average response time.
2. **Read the breakdown by response** A healthy small site is mostly 200 with some 301 and 304. Lots of 404 or 5xx is the first thing to fix.
3. **Read the breakdown by purpose** Discovery against Refresh. If almost everything is refresh and your new pages are slow to appear, discovery is starved.
4. **Read the breakdown by file type** If HTML is a small share of the fetches, something else is eating them: images, scripts, or a feed that changes constantly.
5. **Check the host status** robots.txt fetch, DNS resolution and server connectivity. A red one means Google slows or stops crawling everything.
6. **Sample the example URLs** The report shows what was actually fetched. Read twenty and count how many you would want indexed.

That last step is the test. If five of twenty fetched URLs are parameters or dead pages, you have found the work.

## The fix for each source

| Source | Fix | Note |
| --- | --- | --- |
| Campaign parameters | A self-referencing canonical on the clean URL, and never link internally with parameters | The URL Parameters tool was retired in 2022. The canonical is the control |
| Filters and sorts | Decide which combinations have search demand; the rest are disallowed in robots.txt or noindexed | Do not do both to the same URL: a blocked page cannot deliver a noindex |
| Session ids | Remove them from URLs entirely | A cookie does the job without creating pages |
| Old redirects | Update internal links and the sitemap to the final URL | Keep the redirects themselves. It is the links through them that waste fetches |
| Dead URLs | Return 404 or 410, remove them from the sitemap and from internal links | A 404 nobody links to costs almost nothing |
| Calendars and archives | Noindex the empty ones or remove the template | Keep them for people if they help; keep them out of the sitemap |
| Tag and author archives | Noindex thin ones, or curate a few into real hubs | One tag page with one post is not a page |
| Staging and previews | Block with authentication, not only with robots.txt | A blocked URL can still be indexed as a bare link if someone links to it |

## What not to do

- Do not disallow everything you do not like in robots.txt. A disallowed URL can still be indexed without its content when other sites link to it.
- Do not disallow a URL you have noindexed. Google has to fetch the page to see the noindex, so blocking it keeps it in the index.
- Do not delete old redirects to save crawl. You would be throwing away the links that point at the old URLs.
- Do not add a crawl-delay directive expecting Google to obey. Google ignores it.
- Do not chase crawl waste before the basics. An unindexable page wastes more than a parameter ever will.

> The order that matters: fix what blocks indexing, then what wastes crawl, then what is merely untidy.

## The one-hour audit

1. **Crawl your own site from the home page** List every internal link that returns a redirect or a 404. Those are yours to fix today.
2. **Open the sitemap and check a sample** Every URL in it should answer 200 and be canonical to itself. Remove redirects and dead URLs.
3. **Search for your own parameters** A site: search with inurl: and a parameter name shows what Google has indexed of them.
4. **Read twenty fetched URLs in Crawl stats** Count how many you want indexed. That ratio is your waste, measured.
5. **Fix the largest source, then stop** One source per session. Re-read the report in a month rather than tuning weekly.

## What good looks like

- Most fetches are 200 on HTML you would want indexed.
- The sitemap lists only canonical URLs that answer 200.
- No internal link passes through a redirect.
- New pages appear in the index within days, not weeks.
- The Page indexing report's Not indexed reasons are the ones you chose: alternate pages with a canonical, deliberate noindex, redirects.

At that point crawling is not something you manage. It is something that happens correctly while you write the next page.

## Questions

### Does crawl budget matter for a small site?

Not really. It becomes a constraint past tens of thousands of URLs. Below that, what matters is where the fetches go: parameters, redirects and dead pages taking the crawl your real pages need.

### Should I block parameters in robots.txt?

Only for patterns you never want crawled at all, such as sorts and session ids. For campaign parameters, a self-referencing canonical on the clean URL is the control, and you should never link internally with parameters.

### Why is my new page taking a week to get indexed?

Usually because nothing links to it, or because the crawl is spent elsewhere. Link it from a hub, add it to the sitemap, and check what the Crawl stats report says is actually being fetched.

### Can I make Google crawl more?

Not directly. A fast, error-free server and pages worth crawling increase it over time. Request indexing in the URL Inspection tool asks for one URL, with a small daily quota, and does not change the overall rate.

### Are 404s wasting my crawl?

A few are harmless and expected. Many repeated 404s mean something still links to them: your own pages, your sitemap, or an external site that deserves a redirect.

## Read next

- [The Crawl stats report: how much Google fetches and where it struggles](https://porteur.ai/guides/search-console-crawl-stats-report): Find Crawl stats in Search Console Settings. Read the four charts, host status and breakdowns. Spot 5xx spikes and wasted crawls. Know what is normal.
- [Site architecture for a small product: shallow, linked, and honest](https://porteur.ai/guides/site-architecture): A site is a graph of links, not a folder tree. Hubs that are real pages, navigation a crawler can follow, and a worked map for two hundred pages.
- [Faceted navigation SEO: which filters to index and how to hide the rest](https://porteur.ai/guides/faceted-navigation-seo): Pick which filters should rank, hide the rest, and keep crawl under control. A practical plan for product sites with examples you can ship.
- [Blocked by robots.txt: what Google can still do with the URL](https://porteur.ai/guides/blocked-by-robots-txt): What “Blocked by robots.txt” and “Indexed, though blocked by robots.txt” mean, how to test a URL, the common accidents, and the clean fixes.
- [Delete or redirect an old page: 301, 404 or 410](https://porteur.ai/guides/delete-or-redirect-old-pages): Decide per page: 301 if a clear equivalent exists, else 404 or 410. Check links and clicks first, avoid home page redirects, and monitor what Google drops.
- [Crawl budget](https://porteur.ai/glossary/crawl-budget): Crawl budget is how much Googlebot crawls your site. It matters on sites with thousands of URLs. Here is how to check it and avoid wasting it.
- [URL parameters](https://porteur.ai/glossary/url-parameters): URL parameters are key=value after a question mark. Handle them to avoid duplicate pages, index bloat and wasted crawl, and to keep clean rankings.

Paste your URL and the free check reads your site the way a crawler does, in about thirty seconds, and shows the dead ends and redirects it walked into. Free check: https://porteur.ai/
