# Web crawler

A web crawler is software that fetches your pages, reads their links and fetches those in turn. You meet them because search, AI systems and SEO tools use them. This page shows what to allow, what to block, and how to see what crawlers did on your site.

Updated 2026-09-14 · Source: https://porteur.ai/glossary/web-crawler

## What a web crawler does

A crawler, also called a spider or bot, requests a page, reads the HTML, extracts links, then requests those pages. That is how indexes and audits are built.

Search engines run crawlers to build their index, for example Googlebot and Bingbot. AI vendors run crawlers for training and for answers, for example GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. SEO tools crawl to audit your site, for example Screaming Frog and Sitebulb. Archives and researchers run others.

For a small site, crawlers are how you get discovered, measured and, sometimes, over-fetched. Your job is to make the right pages easy to fetch and the rest cheap to skip.

## Who to allow and who to slow

- Allow search engine crawlers to read the pages you want in search, for example yourproduct.com/, /pricing, /guides/getting-started.
- Allow your chosen SEO crawler when you run an audit, then disallow it when you are done if server capacity is tight.
- Decide what AI crawlers may fetch. You can allow, throttle or disallow them per path. Set a clear policy and document it in robots.txt.
- Block obvious junk routes for every crawler, for example /wp-admin/, /cart, infinite calendars and facet combinations.

If your server is small, crawl spikes can slow real users. Most honest crawlers let you set a crawl rate in their consoles or respect robots crawl-delay when they support it.

## Robots.txt, sitemaps and control

robots.txt tells crawlers what they may fetch. Honest crawlers obey it. Put it at yourproduct.com/robots.txt and serve it fast with a 200 status code.

1. **Decide your rules** List the sections to allow and disallow. Keep public content open. Block admin, search results, filters that explode into infinite URLs and test areas.
2. **Write and test robots.txt** Add User-agent blocks with Allow and Disallow lines. Add separate sections for AI crawlers when you have a policy. Test that key pages are allowed.
3. **Expose your sitemap** Link to your XML sitemap in robots.txt with a Sitemap line. Sitemaps help discovery, they do not force indexing.
4. **Verify in consoles** Check Google Search Console and Bing Webmaster Tools for crawl errors and whether robots.txt was fetched and read.

> Rule: use robots.txt to control crawling, not indexing. Use noindex to keep a fetched page out of search results.

## The JavaScript question

A crawler that does not render JavaScript sees only the HTML your server sends. If your content loads only after client-side rendering, some crawlers will miss it.

- Serve primary content in HTML where you can, for example the product name, price and key copy on /pricing.
- Provide links in HTML. Do not hide navigation behind click handlers that need JavaScript to build hrefs.
- If you use hydration or islands, test a page with JavaScript off. If the core text vanishes, add server rendering or a static fallback.
- For pages that must be JS-only, expect some crawlers and tools to skip or misread them. Plan manual submissions or API feeds where offered.

## How to see and measure crawling

- Search Console Crawl Stats shows fetch volume, response codes and hosts. Watch for spikes, errors and timeouts.
- URL Inspection shows when Google last crawled a page and if it can index it. Check a sample of key URLs.
- Server logs show every request. Filter on user agents like Googlebot, Bingbot, GPTBot and your audit tool to spot heavy hitters or blocked paths.
- Run your own crawl with an SEO crawler to simulate discovery. Fix broken links and orphan pages it finds before search bots waste budget.

A fixed site loads core HTML, links tie related pages together, and robots.txt lets the right bots in and keeps traps closed.

> Warning: blocking CSS or JS in robots.txt can stop Google from rendering your layout correctly for assessment. Only block heavy assets you are sure are safe.

## Questions

### What is a web crawler?

It is software that fetches web pages, reads their links and fetches those in turn. Search engines, AI vendors, SEO tools, archives and researchers use crawlers to discover and analyse content.

### Is a web crawler illegal?

No. Fetching public pages is allowed in general, but a crawler should obey robots.txt and your terms. Abusive crawling that ignores limits or bypasses controls can breach laws or contracts. If a bot harms your site, block it and contact the operator.

### Is ChatGPT a web crawler?

ChatGPT is an AI interface, not a crawler. Some AI vendors run their own crawlers, for example GPTBot and OAI-SearchBot, to fetch content for training or answers. Control them with robots.txt if you have a policy.

### Does WebCrawler still exist?

WebCrawler as a meta search engine still exists as a brand. It is not the same thing as a web crawler program. When people say web crawler, they mean the software that fetches and follows links.

### How do I submit my website to Google?

Publish an XML sitemap and link to it in robots.txt. Add your site to Google Search Console and request indexing for key URLs with URL Inspection. You do not need to submit every page if your internal links are clean.

### How often will crawlers visit my site?

It depends on site health, popularity and change rate. New or small sites may see sporadic visits. Keep pages fast, link them well and update content to earn steadier crawling.

## Read next

- [robots.txt](https://porteur.ai/glossary/robots-txt): robots.txt tells crawlers which URLs they may fetch. See what it does not do, how to test it, what to put in it, and how to handle AI bots.
- [Crawl budget](https://porteur.ai/glossary/crawl-budget): Crawl budget is how much Googlebot crawls your site. It matters on sites with thousands of URLs. Here is how to check it and avoid wasting it.
- [The Crawl stats report: how much Google fetches and where it struggles](https://porteur.ai/guides/search-console-crawl-stats-report): Find Crawl stats in Search Console Settings. Read the four charts, host status and breakdowns. Spot 5xx spikes and wasted crawls. Know what is normal.
- [How to use the URL Inspection tool in Search Console](https://porteur.ai/guides/url-inspection-tool): Read each panel, run Test live URL, and know when to request indexing. Fix new pages, dropped pages, and canonicals Google ignores.
- [JavaScript SEO](https://porteur.ai/glossary/javascript-seo): What JavaScript SEO means, how Googlebot renders, what breaks, how to fix links and content, and how to choose SSR or SSG so pages index fast.
- [Googlebot](https://porteur.ai/glossary/googlebot): Googlebot is Google’s web crawler. Learn its versions, what it obeys, how it renders JavaScript, how to verify it, and what to fix on a small site.
- [The best website crawlers for SEO: desktop, cloud, and free tiers](https://porteur.ai/guides/best-website-crawlers): Choose the right SEO crawler for a small site or a growing one. Screaming Frog, Sitebulb, Lumar, Oncrawl, and the suites’ audits, plus a first crawl plan.

Get a quick read on how bots discover your site with the free check: paste a URL, in about thirty seconds it reports three findings with rival context. Free check: https://porteur.ai/
