# The Crawl stats report: how much Google fetches and where it struggles

You use the Crawl stats report to see how Google fetches your site and where it struggles. This guide shows what to check, what is normal for a small site, and what to fix when the charts look wrong.

Updated 2026-09-14 · Source: https://porteur.ai/guides/search-console-crawl-stats-report

## Where the Crawl stats report hides

Open Google Search Console. Pick a root-level property. In the left menu, click Settings, then Crawl stats.

> You only get Crawl stats for root-level properties. If you verified a URL-prefix, add a Domain property with DNS verification.

The report covers the last 90 days. It shows total crawl requests, total download size and average response time. It also shows host status and request breakdowns.

## First read: four charts that tell the story

- Total crawl requests: volume over time. A steady line is normal on a stable site.
- Total download size: bytes fetched. A jump can mean heavier pages or many large assets.
- Average response time: how fast your server answered Google. A climb means your server, code or network slowed.
- Host status: recent checks for robots.txt fetch, DNS resolution and server connectivity.

Scan the 90‑day trends. Mark spikes and drops. Then open each breakdown to find what changed: response codes, file types, purpose and Googlebot type.

## Host status: what a red light means

Host status checks three basics: Google could fetch robots.txt, resolve DNS and connect to your server. If one shows red, Google had recent failures.

| Check | What a failure suggests | What to do |
| --- | --- | --- |
| robots.txt fetch | robots.txt unreachable or too slow | Serve robots.txt fast at /robots.txt. Avoid heavy rewrites. Do not block Google unless you mean it. |
| DNS resolution | Nameserver or DNS record trouble | Confirm A/AAAA and CNAME records. Check nameserver uptime. Reduce TTL while fixing. |
| Server connectivity | TCP timeouts, drops, firewall blocks | Check hosting health, firewalls and rate limits. Review load balancer rules. Restore capacity. |

Fix the red check first. Google slows crawling when it cannot fetch robots.txt or when your host fails to answer. When green again, crawling ramps back by itself.

## Breakdowns: find where crawl time is spent

Open each breakdown. You get requests by response code, by file type, by purpose and by Googlebot type. Compare the timelines to your spikes or drops.

| Breakdown | What normal looks like on a small site | Warning signs |
| --- | --- | --- |
| Response codes | Mostly 200 and not modified responses. Some 404 from typos and removals. | 5xx spikes. Many 404 from broken links. 301 or 302 dominating for weeks. |
| File types | HTML, CSS and JS lead. Images, fonts and JSON as needed. | Huge image or video share you did not ship. Many JS fetches from one route. |
| Purpose (Discovery vs Refresh) | A mix. Refresh leads on stable content. Discovery rises after new pages or a sitemap update. | Discovery high on parameter URLs. Refresh stuck on old redirects. |
| Googlebot type | Smartphone leads on most sites. Some Desktop, Image and Video. | Most requests by AdsBot if you do not run ads. Image bot peaking on sprite sheets. |

Click a response code to see example URLs. Copy a few and test them in your browser and server logs. Fix patterns, not single URLs.

## What is normal for a small site

If your site has fewer than tens of thousands of URLs, crawl budget is rarely the problem. Google can fetch and refresh that scale without special help.

- Steady crawl requests with small weekly swings.
- Average response time stable within a narrow band.
- Most requests on HTML and core assets. Few 404s.
- Discovery bumps when you add a section, for example /guides/.

A healthy report for yourproduct.com shows most crawl time on real pages like /, /pricing and /guides/getting-started. Old redirects and querystring pages are low or gone.

## Real problems the report surfaces

- 5xx spikes: server errors waste crawl and stall indexing.
- Most requests on parameters: Google exploring endless URLs like /list?sort= and /list?page=.
- Crawl time on old redirects: long chains or loops after a migration.
- Average response time climbing: slower app, busy database or network issues.
- Total download size surging: oversized images or assets from a release.

> Do not block problem URLs in robots.txt as your first fix. Blocking stops crawling, but it does not clean up indexing of URLs Google already knows.

## Fixing 5xx spikes and slow responses

1. **Confirm the fault window** Match the spike dates to deploys, traffic surges or hosting incidents. Check your uptime monitor and server logs.
2. **Stabilise the host** Add capacity or roll back. Remove rate limits that block Google. Ensure your load balancer and WAF allow Googlebot.
3. **Serve lighter pages while hot** Delay heavy jobs in requests. Cache HTML for anonymous users. Use a CDN for static assets.
4. **Trim expensive endpoints** Audit routes with long times to first byte. Paginate costly listings. Reduce third‑party calls in critical paths.
5. **Retest in Crawl stats** When average response time returns to normal and 5xx fall, Google will crawl more again.

If a page is gone, return 410 or 404, not a 500. Keep your error budget for real faults, not content removals. A clean 404 does not harm crawl health.

## Fixing parameter crawl and wasteful URLs

Parameter sprawl is common. You see many requests for /products?ref=, /blog?utm_source= and /search?page=. These add near‑duplicate pages and infinite spaces.

- Canonical: point parameter variants to the clean URL when the content is the same.
- Noindex: add noindex on thin results, for example filtered lists that no one should land on.
- Internal links: stop linking to tracked URLs. Strip UTM parameters on site links.
- Pagination: cap the highest page you link to. Use rel=next/prev patterns in your UI, not as signals to Google.
- Faceted navigation: only link combinations that have search demand and unique value.
- Robots.txt: disallow only when you are sure you do not need Google to crawl those patterns.

After a week, Crawl stats should show Discovery falling on parameter URLs and Refresh rising on your core pages like /collections/shoes and /blog/how-we-built-search.

## Fixing redirect churn

Redirect chains and loops burn crawl. You see many 301 or 302 in the response breakdown. Typical after a rebrand or during a CMS change.

1. **Map the chains** Pick examples from the report. Use a redirect checker to trace hops from /old to /new.
2. **Flatten to one hop** Redirect old URLs directly to the final destination. Remove intermediate hops.
3. **Fix internal links** Update menus, footers, sitemaps and body links to the final URLs. Do not link to redirected URLs.
4. **Expire stale routes** If content is gone, return 410 or 404. Do not redirect everything to /.
5. **Verify in Crawl stats** Requests should shift from redirects to 200. Average response time often improves too.

## File types, bots and purpose: tune what Google fetches

Large images and heavy JS inflate download size. If those lines jump without a change you planned, a theme or plugin shipped more assets than before.

- Images: compress and resize. Use modern formats. Lazy‑load non‑critical images.
- JS and CSS: remove unused code. Split bundles. Serve with caching and compression.
- Page resources: avoid querystrings on static assets. Set long cache headers and consistent URLs.

Check Googlebot type. Smartphone should lead on most sites. If Desktop is higher, your mobile site may be thin or blocked. Fix parity between mobile and desktop content and internal links.

On purpose, Discovery rising after you ship new pages or submit a sitemap is fine. Discovery rising on junk patterns is not. Refresh should dominate on mature sections like /pricing.

## Use sitemaps and links to guide crawling

Give Google clean paths to what matters. Sitemaps and internal links are the levers you control. The Crawl stats report tells you if Google is following them.

- XML sitemaps: submit a sitemap or index in the Sitemaps report. Google reads lastmod when it is consistently accurate.
- Scope: a sitemap may hold up to 50,000 URLs or 50 MB uncompressed and can only list URLs on its own host or an approved cross‑host setup.
- Coverage: include canonical, indexable URLs only. Exclude redirects, noindex and parameters.
- Internal links: link to pages you want crawled. Avoid orphan pages. Update links after a move.

A fixed site shows Crawl stats requests clustering on the URLs you linked in your navigation and sitemap, not on /?page= or /tag/uncategorised/?sort=oldest.

## When Crawl stats is bad but rankings look fine

Do not chase a flat line. Crawl demand fluctuates. Treat Crawl stats as an operational signal. Only act when you see clear waste or failure patterns.

Cross‑check with the Performance report and the Core Web Vitals report. If clicks and vitals are steady, a short‑term crawl spike may be nothing. Keep notes near deploys and incidents.

## Questions

### How can I see the Crawl stats report in Google Search Console?

Open Search Console, pick a root‑level property, then go to Settings and select Crawl stats. The report shows 90 days of crawl requests, download size and average response time, with host status and breakdowns.

### How do I check if Google has crawled my site?

Use the Crawl stats report to see total crawl requests and the breakdowns. For a single URL, use the URL Inspection tool to fetch the current index status and the last crawl details for that page.

### How to check crawl budget?

On a small site with fewer than tens of thousands of URLs, crawl budget is rarely your bottleneck. Use Crawl stats to spot waste, for example parameters or redirects. Fix those and keep your server fast. If you run a very large site, use sitemaps, strong internal links and fast responses to help Google allocate more crawl.

### How often will Google crawl my site?

It varies by site, section and freshness. Crawl stats shows the pattern over the last 90 days. New or updated pages are often fetched more, then settle into a refresh rhythm. If your server slows or fails, crawling will dip until stability returns.

### What should I do if I see a lot of 404s in Crawl stats?

Check the example URLs and fix the sources. Update internal links and sitemaps to point at live URLs. Return 410 or 404 for content that is truly gone. You do not need to redirect every 404 to the home page.

### Can I reduce crawling of parameter URLs without hurting important pages?

Yes. Canonical to the clean URL when content is the same, add noindex on thin filtered results and stop linking to tracked URLs. Only use robots.txt disallow when you are certain the patterns are safe to block from crawling.

## Read next

- [How to use Google Search Console in ten minutes a week](https://porteur.ai/guides/how-to-use-google-search-console): A quick weekly routine: set four filters, compare 28 days, check pages then queries, and fix three findings, without getting lost in noise.
- [The Sitemaps report: what Success, Has errors and Couldn’t fetch mean](https://porteur.ai/guides/search-console-sitemaps-report): How to submit your sitemap URL, read Success, Has errors and Couldn’t fetch, and fix invalid XML, 404s, cross‑host URLs and missing child sitemaps.
- [How to use the URL Inspection tool in Search Console](https://porteur.ai/guides/url-inspection-tool): Read each panel, run Test live URL, and know when to request indexing. Fix new pages, dropped pages, and canonicals Google ignores.
- [The Search Console performance report, column by column](https://porteur.ai/guides/search-console-performance-report): Every column in the Google Search Console performance report explained with the traps that trip small sites, and how to act on each.
- [Crawl budget](https://porteur.ai/glossary/crawl-budget): Crawl budget is how much Googlebot crawls your site. It matters on sites with thousands of URLs. Here is how to check it and avoid wasting it.
- [Crawled, currently not indexed: what it means and what to do](https://porteur.ai/guides/crawled-currently-not-indexed): What “Crawled, currently not indexed” means in Search Console, how to tell why it happened on your site, and what to change that gets pages indexed.
- [Crawling](https://porteur.ai/glossary/crawling): Crawling is how Googlebot discovers and fetches your URLs. See what it fetched, fix slow or blocked areas, and know when crawl budget matters.

Run yourdomain.com through Porteur to get a free check that reads your site, the searches around it and rivals from a URL in about thirty seconds and shows three findings whole. Free check: https://porteur.ai/
