Crawling
Crawling is Googlebot discovering and fetching your pages so they can be rendered and considered for indexing. For a small site, it is mostly about not blocking Google, staying fast, and linking pages so they get found. Here is what governs it, where to see it, and what to fix.
By Théophile Louvart, founder of Porteur · Updated 14 September 2026 · Markdown
What crawling is, in plain terms
Crawling is the first of Google’s three stages: crawling, indexing, serving. Googlebot discovers URLs from links, sitemaps and past crawls, then fetches them within your site’s crawl capacity. It renders pages with an evergreen Chromium in a second wave.
If Google cannot crawl a page, it cannot index or rank it. For a 50 to 5,000 URL site, the work is simple: allow access, keep responses healthy, and link every page.
What governs crawling on a small site
- robots.txt sets what Googlebot may fetch. A Disallow on a folder like /static/ or /blog/ stops crawling there.
- Speed and errors control pace. Google slows down on 5xx and timeouts, and will back off if your server strains.
- Your internal links decide what gets found. Orphan pages with no links are often discovered late or never.
- Sitemaps help discovery, not priority. Submit a clean XML sitemap that lists canonical URLs you want indexed.
On a fixed site, a good crawl looks like frequent fetches of your homepage, key hubs like /blog/ or /docs/, and steady retrievals of new or updated URLs shortly after you publish.
Where to see crawling in Search Console
Check Crawl stats
Open Settings, then Crawl stats. Review total requests, average response time, and the split by response code and file type.
Spot errors and slowdowns
Filter for 5xx and timeouts. Look at spikes in response time. Click a host or path to find problem sections, for example /api/ or /images/.
See discovered but not yet crawled
Open the Page indexing report. Look for “Discovered, currently not indexed”. Google has the URL but has not fetched it yet.
Inspect a URL
Use the URL Inspection tool for /pricing or a new post. Check crawl status, last crawl date, and if Google can fetch JavaScript and resources.
If new pages sit as discovered for days, add internal links from crawled hubs like the homepage or /guides/ and include them in your sitemap index.
What to do to improve crawling
- Keep 200 responses fast and stable. Aim for LCP at 2.5 seconds and avoid 5xx during deploys.
- Link new pages from existing pages. For a post at /guides/getting-started, add links from /guides/ and the homepage.
- Submit a clean sitemap
- Fix broken links and loops. Replace links to 404s, and update chains like /old to /older to /new to a direct /old to /new 301.
- Allow required resources. Let Google fetch CSS, JS and images it needs to render the page.
- Use simple pagination and parameters. Prevent infinite calendars and endless filters that spawn near-duplicate URLs.
After fixes, watch Crawl stats for a week. You want fewer 5xx, a stable response time, and new URLs moving from discovered to crawled to indexed.
When crawl budget matters
Crawl budget is only a concern for sites with many thousands of URLs. If you run a small SaaS or store, problems are almost always configuration or speed, not budget.
For large catalogues or programmatic pages, reduce low value URLs. Prune faceted combinations, control parameters, and link key listings one click from /categories/.
Questions
Crawlers are automated fetchers like Googlebot that discover and retrieve URLs. They follow links, read sitemaps, and revisit known pages. Their goal is to fetch content so it can be rendered and considered for indexing.
Crawling is discovery and fetching. Indexing is storing the processed page so it can appear for queries. A page must be crawled before it can be indexed and served, but a crawled page is not guaranteed to be indexed.
Crawlability is how easily a bot can fetch your pages. It depends on access rules, server health, and links that expose all URLs. Good crawlability means Google can reach, fetch and render the content without errors or detours.
Link it from pages Google crawls often, like the homepage or /blog/. Add it to your XML sitemap. Avoid blocking resources. You can also request indexing in Search Console, but internal links usually work fastest.
If you have fewer than many thousands of URLs, focus on fixes, not budget. Remove 5xx, speed up responses, submit a clean sitemap, and add internal links. Budget concerns start when vast URL sets compete for attention.
No. ChatGPT is not a crawler. A crawler is software that fetches pages from the web. ChatGPT is a model that generates text. They are different systems.
Sources
Check my site, free
Want to see how Googlebot reaches your pages today? Paste your URL into Porteur’s free check, it reads your site and nearby searches in about thirty seconds and shows three crawl findings.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent
Read next
- GlossaryIndexation
- Glossaryrobots.txt
- GlossaryXML sitemap
- GuideThe Crawl stats report: how much Google fetches and where it struggles
- GuideDiscovered, currently not indexed: why Google has not crawled the page
- GuideCrawled, currently not indexed: what it means and what to do
- GlossaryCrawl budget
- GuideThe GA4 landing page report: the pages that bring people, and what they do