Crawl waste on a small site: the fetches your real pages never get
Crawl budget matters past tens of thousands of URLs, which is not you. What matters at your size is where the fetches go: if most of them land on parameters, redirects and dead pages, your new guide waits a week to be seen.
By Théophile Louvart, founder of Porteur · Updated 15 September 2026 · Markdown
Crawl budget is not your problem
Crawl budget is a real constraint on sites with tens of thousands of URLs. On a site with two hundred pages, Google will happily crawl everything you have.
- Googlebot discovers URLs from links, sitemaps and past crawls, and fetches them within what your server can take.
- It slows down when the server is slow or returns errors, and speeds up when it is not.
- Nothing you can set makes it crawl more. There is no dial.
- So the question is not how many fetches you get. It is what they are spent on.
What crawl waste is
Crawl waste is every fetch that lands on a URL you do not want indexed, do not maintain, or removed months ago. Each one is a fetch your real pages did not get.
The symptoms are recognisable: a new page takes days to appear, a change to an old page takes a fortnight to show in the result, and the Crawl stats report shows plenty of activity that never touches the pages you care about.
Where it comes from on a small site
| Source | What it looks like | Why it happens |
|---|---|---|
| Query parameters | /pricing?ref=twitter, /guides?utm_source=newsletter | Campaign links get shared, indexed and crawled as separate URLs |
| Filters and sorts | /tools?sort=name&view=grid | Every combination is a URL, and they multiply |
| Session and tracking ids | /page?sid=8fe2... | An old framework habit that creates a URL per visitor |
| Old redirects | Fetches of URLs that 301 somewhere else | Internal links and the sitemap still point at the old address |
| Dead URLs | Repeated 404s on paths deleted last year | They are still linked from your own pages or from a stale sitemap |
| Calendars and archives | /events/2019/07/, /blog/page/12/ | Generated pages that go on forever and hold nothing |
| Tag and author archives | /tags/seo/, /author/founder/ | A CMS creates them by default, often with one item each |
| Staging and preview URLs | A preview domain that got linked once | It is crawlable and nobody told it not to be |
Notice how many of these you created for a reason that has nothing to do with search, and then forgot.
How to see it in the Crawl stats report
Open Settings, then Crawl stats
It is available for root-level properties and covers the last ninety days: total requests, download size and average response time.
Read the breakdown by response
A healthy small site is mostly 200 with some 301 and 304. Lots of 404 or 5xx is the first thing to fix.
Read the breakdown by purpose
Discovery against Refresh. If almost everything is refresh and your new pages are slow to appear, discovery is starved.
Read the breakdown by file type
If HTML is a small share of the fetches, something else is eating them: images, scripts, or a feed that changes constantly.
Check the host status
robots.txt fetch, DNS resolution and server connectivity. A red one means Google slows or stops crawling everything.
Sample the example URLs
The report shows what was actually fetched. Read twenty and count how many you would want indexed.
That last step is the test. If five of twenty fetched URLs are parameters or dead pages, you have found the work.
The fix for each source
| Source | Fix | Note |
|---|---|---|
| Campaign parameters | A self-referencing canonical on the clean URL, and never link internally with parameters | The URL Parameters tool was retired in 2022. The canonical is the control |
| Filters and sorts | Decide which combinations have search demand; the rest are disallowed in robots.txt or noindexed | Do not do both to the same URL: a blocked page cannot deliver a noindex |
| Session ids | Remove them from URLs entirely | A cookie does the job without creating pages |
| Old redirects | Update internal links and the sitemap to the final URL | Keep the redirects themselves. It is the links through them that waste fetches |
| Dead URLs | Return 404 or 410, remove them from the sitemap and from internal links | A 404 nobody links to costs almost nothing |
| Calendars and archives | Noindex the empty ones or remove the template | Keep them for people if they help; keep them out of the sitemap |
| Tag and author archives | Noindex thin ones, or curate a few into real hubs | One tag page with one post is not a page |
| Staging and previews | Block with authentication, not only with robots.txt | A blocked URL can still be indexed as a bare link if someone links to it |
What not to do
- Do not disallow everything you do not like in robots.txt. A disallowed URL can still be indexed without its content when other sites link to it.
- Do not disallow a URL you have noindexed. Google has to fetch the page to see the noindex, so blocking it keeps it in the index.
- Do not delete old redirects to save crawl. You would be throwing away the links that point at the old URLs.
- Do not add a crawl-delay directive expecting Google to obey. Google ignores it.
- Do not chase crawl waste before the basics. An unindexable page wastes more than a parameter ever will.
The one-hour audit
Crawl your own site from the home page
List every internal link that returns a redirect or a 404. Those are yours to fix today.
Open the sitemap and check a sample
Every URL in it should answer 200 and be canonical to itself. Remove redirects and dead URLs.
Search for your own parameters
A site: search with inurl: and a parameter name shows what Google has indexed of them.
Read twenty fetched URLs in Crawl stats
Count how many you want indexed. That ratio is your waste, measured.
Fix the largest source, then stop
One source per session. Re-read the report in a month rather than tuning weekly.
What good looks like
- Most fetches are 200 on HTML you would want indexed.
- The sitemap lists only canonical URLs that answer 200.
- No internal link passes through a redirect.
- New pages appear in the index within days, not weeks.
- The Page indexing report's Not indexed reasons are the ones you chose: alternate pages with a canonical, deliberate noindex, redirects.
At that point crawling is not something you manage. It is something that happens correctly while you write the next page.
Questions
Not really. It becomes a constraint past tens of thousands of URLs. Below that, what matters is where the fetches go: parameters, redirects and dead pages taking the crawl your real pages need.
Only for patterns you never want crawled at all, such as sorts and session ids. For campaign parameters, a self-referencing canonical on the clean URL is the control, and you should never link internally with parameters.
Usually because nothing links to it, or because the crawl is spent elsewhere. Link it from a hub, add it to the sitemap, and check what the Crawl stats report says is actually being fetched.
Not directly. A fast, error-free server and pages worth crawling increase it over time. Request indexing in the URL Inspection tool asks for one URL, with a small daily quota, and does not change the overall rate.
A few are harmless and expected. Many repeated 404s mean something still links to them: your own pages, your sitemap, or an external site that deserves a redirect.
Check my site, free
Paste your URL and the free check reads your site the way a crawler does, in about thirty seconds, and shows the dead ends and redirects it walked into.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent
Read next
- GuideThe Crawl stats report: how much Google fetches and where it struggles
- GuideSite architecture for a small product: shallow, linked, and honest
- GuideFaceted navigation SEO: which filters to index and how to hide the rest
- GuideBlocked by robots.txt: what Google can still do with the URL
- GuideDelete or redirect an old page: 301, 404 or 410
- GlossaryCrawl budget
- GlossaryURL parameters
- GuideSEO before you launch: what to get right while the site is still small