Crawling

Crawling is Googlebot discovering and fetching your pages so they can be rendered and considered for indexing. For a small site, it is mostly about not blocking Google, staying fast, and linking pages so they get found. Here is what governs it, where to see it, and what to fix.

By , founder of Porteur · Updated 14 September 2026 · Markdown

What crawling is, in plain terms

Crawling is the first of Google’s three stages: crawling, indexing, serving. Googlebot discovers URLs from links, sitemaps and past crawls, then fetches them within your site’s crawl capacity. It renders pages with an evergreen Chromium in a second wave.

If Google cannot crawl a page, it cannot index or rank it. For a 50 to 5,000 URL site, the work is simple: allow access, keep responses healthy, and link every page.

What governs crawling on a small site

  • robots.txt sets what Googlebot may fetch. A Disallow on a folder like /static/ or /blog/ stops crawling there.
  • Speed and errors control pace. Google slows down on 5xx and timeouts, and will back off if your server strains.
  • Your internal links decide what gets found. Orphan pages with no links are often discovered late or never.
  • Sitemaps help discovery, not priority. Submit a clean XML sitemap that lists canonical URLs you want indexed.

On a fixed site, a good crawl looks like frequent fetches of your homepage, key hubs like /blog/ or /docs/, and steady retrievals of new or updated URLs shortly after you publish.

Where to see crawling in Search Console

  1. Check Crawl stats

    Open Settings, then Crawl stats. Review total requests, average response time, and the split by response code and file type.

  2. Spot errors and slowdowns

    Filter for 5xx and timeouts. Look at spikes in response time. Click a host or path to find problem sections, for example /api/ or /images/.

  3. See discovered but not yet crawled

    Open the Page indexing report. Look for “Discovered, currently not indexed”. Google has the URL but has not fetched it yet.

  4. Inspect a URL

    Use the URL Inspection tool for /pricing or a new post. Check crawl status, last crawl date, and if Google can fetch JavaScript and resources.

If new pages sit as discovered for days, add internal links from crawled hubs like the homepage or /guides/ and include them in your sitemap index.

What to do to improve crawling

  • Keep 200 responses fast and stable. Aim for LCP at 2.5 seconds and avoid 5xx during deploys.
  • Link new pages from existing pages. For a post at /guides/getting-started, add links from /guides/ and the homepage.
  • Submit a clean sitemap
  • Fix broken links and loops. Replace links to 404s, and update chains like /old to /older to /new to a direct /old to /new 301.
  • Allow required resources. Let Google fetch CSS, JS and images it needs to render the page.
  • Use simple pagination and parameters. Prevent infinite calendars and endless filters that spawn near-duplicate URLs.

After fixes, watch Crawl stats for a week. You want fewer 5xx, a stable response time, and new URLs moving from discovered to crawled to indexed.

When crawl budget matters

Crawl budget is only a concern for sites with many thousands of URLs. If you run a small SaaS or store, problems are almost always configuration or speed, not budget.

For large catalogues or programmatic pages, reduce low value URLs. Prune faceted combinations, control parameters, and link key listings one click from /categories/.

Questions

Sources

Check my site, free

Want to see how Googlebot reaches your pages today? Paste your URL into Porteur’s free check, it reads your site and nearby searches in about thirty seconds and shows three crawl findings.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next