Web crawler

A web crawler is software that fetches your pages, reads their links and fetches those in turn. You meet them because search, AI systems and SEO tools use them. This page shows what to allow, what to block, and how to see what crawlers did on your site.

By , founder of Porteur · Updated 14 September 2026 · Markdown

What a web crawler does

A crawler, also called a spider or bot, requests a page, reads the HTML, extracts links, then requests those pages. That is how indexes and audits are built.

Search engines run crawlers to build their index, for example Googlebot and Bingbot. AI vendors run crawlers for training and for answers, for example GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. SEO tools crawl to audit your site, for example Screaming Frog and Sitebulb. Archives and researchers run others.

For a small site, crawlers are how you get discovered, measured and, sometimes, over-fetched. Your job is to make the right pages easy to fetch and the rest cheap to skip.

Who to allow and who to slow

  • Allow search engine crawlers to read the pages you want in search, for example yourproduct.com/, /pricing, /guides/getting-started.
  • Allow your chosen SEO crawler when you run an audit, then disallow it when you are done if server capacity is tight.
  • Decide what AI crawlers may fetch. You can allow, throttle or disallow them per path. Set a clear policy and document it in robots.txt.
  • Block obvious junk routes for every crawler, for example /wp-admin/, /cart, infinite calendars and facet combinations.

If your server is small, crawl spikes can slow real users. Most honest crawlers let you set a crawl rate in their consoles or respect robots crawl-delay when they support it.

Robots.txt, sitemaps and control

robots.txt tells crawlers what they may fetch. Honest crawlers obey it. Put it at yourproduct.com/robots.txt and serve it fast with a 200 status code.

  1. Decide your rules

    List the sections to allow and disallow. Keep public content open. Block admin, search results, filters that explode into infinite URLs and test areas.

  2. Write and test robots.txt

    Add User-agent blocks with Allow and Disallow lines. Add separate sections for AI crawlers when you have a policy. Test that key pages are allowed.

  3. Expose your sitemap

    Link to your XML sitemap in robots.txt with a Sitemap line. Sitemaps help discovery, they do not force indexing.

  4. Verify in consoles

    Check Google Search Console and Bing Webmaster Tools for crawl errors and whether robots.txt was fetched and read.

The JavaScript question

A crawler that does not render JavaScript sees only the HTML your server sends. If your content loads only after client-side rendering, some crawlers will miss it.

  • Serve primary content in HTML where you can, for example the product name, price and key copy on /pricing.
  • Provide links in HTML. Do not hide navigation behind click handlers that need JavaScript to build hrefs.
  • If you use hydration or islands, test a page with JavaScript off. If the core text vanishes, add server rendering or a static fallback.
  • For pages that must be JS-only, expect some crawlers and tools to skip or misread them. Plan manual submissions or API feeds where offered.

How to see and measure crawling

  • Search Console Crawl Stats shows fetch volume, response codes and hosts. Watch for spikes, errors and timeouts.
  • URL Inspection shows when Google last crawled a page and if it can index it. Check a sample of key URLs.
  • Server logs show every request. Filter on user agents like Googlebot, Bingbot, GPTBot and your audit tool to spot heavy hitters or blocked paths.
  • Run your own crawl with an SEO crawler to simulate discovery. Fix broken links and orphan pages it finds before search bots waste budget.

A fixed site loads core HTML, links tie related pages together, and robots.txt lets the right bots in and keeps traps closed.

Questions

Sources

Check my site, free

Get a quick read on how bots discover your site with the free check: paste a URL, in about thirty seconds it reports three findings with rival context.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next