# Crawling

Crawling is Googlebot discovering and fetching your pages so they can be rendered and considered for indexing. For a small site, it is mostly about not blocking Google, staying fast, and linking pages so they get found. Here is what governs it, where to see it, and what to fix.

Updated 2026-09-14 · Source: https://porteur.ai/glossary/crawling

## What crawling is, in plain terms

Crawling is the first of Google’s three stages: crawling, indexing, serving. Googlebot discovers URLs from links, sitemaps and past crawls, then fetches them within your site’s crawl capacity. It renders pages with an evergreen Chromium in a second wave.

If Google cannot crawl a page, it cannot index or rank it. For a 50 to 5,000 URL site, the work is simple: allow access, keep responses healthy, and link every page.

## What governs crawling on a small site

- robots.txt sets what Googlebot may fetch. A Disallow on a folder like /static/ or /blog/ stops crawling there.
- Speed and errors control pace. Google slows down on 5xx and timeouts, and will back off if your server strains.
- Your internal links decide what gets found. Orphan pages with no links are often discovered late or never.
- Sitemaps help discovery, not priority. Submit a clean XML sitemap that lists canonical URLs you want indexed.

On a fixed site, a good crawl looks like frequent fetches of your homepage, key hubs like /blog/ or /docs/, and steady retrievals of new or updated URLs shortly after you publish.

## Where to see crawling in Search Console

1. **Check Crawl stats** Open Settings, then Crawl stats. Review total requests, average response time, and the split by response code and file type.
2. **Spot errors and slowdowns** Filter for 5xx and timeouts. Look at spikes in response time. Click a host or path to find problem sections, for example /api/ or /images/.
3. **See discovered but not yet crawled** Open the Page indexing report. Look for “Discovered, currently not indexed”. Google has the URL but has not fetched it yet.
4. **Inspect a URL** Use the URL Inspection tool for /pricing or a new post. Check crawl status, last crawl date, and if Google can fetch JavaScript and resources.

If new pages sit as discovered for days, add internal links from crawled hubs like the homepage or /guides/ and include them in your sitemap index.

## What to do to improve crawling

- Keep 200 responses fast and stable. Aim for LCP at 2.5 seconds and avoid 5xx during deploys.
- Link new pages from existing pages. For a post at /guides/getting-started, add links from /guides/ and the homepage.
- Submit a clean sitemap
- Fix broken links and loops. Replace links to 404s, and update chains like /old to /older to /new to a direct /old to /new 301.
- Allow required resources. Let Google fetch CSS, JS and images it needs to render the page.
- Use simple pagination and parameters. Prevent infinite calendars and endless filters that spawn near-duplicate URLs.

After fixes, watch Crawl stats for a week. You want fewer 5xx, a stable response time, and new URLs moving from discovered to crawled to indexed.

> Do not block Googlebot from fetching CSS and JS it needs to render. If it cannot render, it may misread layout, content, or links.

## When crawl budget matters

Crawl budget is only a concern for sites with many thousands of URLs. If you run a small SaaS or store, problems are almost always configuration or speed, not budget.

For large catalogues or programmatic pages, reduce low value URLs. Prune faceted combinations, control parameters, and link key listings one click from /categories/.

## Questions

### What are crawlers in SEO?

Crawlers are automated fetchers like Googlebot that discover and retrieve URLs. They follow links, read sitemaps, and revisit known pages. Their goal is to fetch content so it can be rendered and considered for indexing.

### What is crawling vs indexing?

Crawling is discovery and fetching. Indexing is storing the processed page so it can appear for queries. A page must be crawled before it can be indexed and served, but a crawled page is not guaranteed to be indexed.

### What does crawlability mean in SEO?

Crawlability is how easily a bot can fetch your pages. It depends on access rules, server health, and links that expose all URLs. Good crawlability means Google can reach, fetch and render the content without errors or detours.

### How can I make Google crawl a new page sooner?

Link it from pages Google crawls often, like the homepage or /blog/. Add it to your XML sitemap. Avoid blocking resources. You can also request indexing in Search Console, but internal links usually work fastest.

### Does crawl budget matter for my site?

If you have fewer than many thousands of URLs, focus on fixes, not budget. Remove 5xx, speed up responses, submit a clean sitemap, and add internal links. Budget concerns start when vast URL sets compete for attention.

### Is ChatGPT a web crawler?

No. ChatGPT is not a crawler. A crawler is software that fetches pages from the web. ChatGPT is a model that generates text. They are different systems.

## Read next

- [Indexation](https://porteur.ai/glossary/indexation): Indexation means a page is stored in Google’s index and can appear in search. Here is how to check status, read Page indexing, and fix what blocks it.
- [robots.txt](https://porteur.ai/glossary/robots-txt): robots.txt tells crawlers which URLs they may fetch. See what it does not do, how to test it, what to put in it, and how to handle AI bots.
- [XML sitemap](https://porteur.ai/glossary/xml-sitemap): What to put in sitemap.xml, how lastmod works, how to create, validate and submit your XML sitemap, and the traps to avoid.
- [The Crawl stats report: how much Google fetches and where it struggles](https://porteur.ai/guides/search-console-crawl-stats-report): Find Crawl stats in Search Console Settings. Read the four charts, host status and breakdowns. Spot 5xx spikes and wasted crawls. Know what is normal.
- [Discovered, currently not indexed: why Google has not crawled the page](https://porteur.ai/guides/discovered-currently-not-indexed): Google knows the URL but has not crawled it. Here is how to check why, cut junk URLs, add internal links, fix the sitemap, and set a date.
- [Crawled, currently not indexed: what it means and what to do](https://porteur.ai/guides/crawled-currently-not-indexed): What “Crawled, currently not indexed” means in Search Console, how to tell why it happened on your site, and what to change that gets pages indexed.

Want to see how Googlebot reaches your pages today? Paste your URL into Porteur’s free check, it reads your site and nearby searches in about thirty seconds and shows three crawl findings. Free check: https://porteur.ai/
