Web crawler
A web crawler is software that fetches your pages, reads their links and fetches those in turn. You meet them because search, AI systems and SEO tools use them. This page shows what to allow, what to block, and how to see what crawlers did on your site.
By Théophile Louvart, founder of Porteur · Updated 14 September 2026 · Markdown
What a web crawler does
A crawler, also called a spider or bot, requests a page, reads the HTML, extracts links, then requests those pages. That is how indexes and audits are built.
Search engines run crawlers to build their index, for example Googlebot and Bingbot. AI vendors run crawlers for training and for answers, for example GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. SEO tools crawl to audit your site, for example Screaming Frog and Sitebulb. Archives and researchers run others.
For a small site, crawlers are how you get discovered, measured and, sometimes, over-fetched. Your job is to make the right pages easy to fetch and the rest cheap to skip.
Who to allow and who to slow
- Allow search engine crawlers to read the pages you want in search, for example yourproduct.com/, /pricing, /guides/getting-started.
- Allow your chosen SEO crawler when you run an audit, then disallow it when you are done if server capacity is tight.
- Decide what AI crawlers may fetch. You can allow, throttle or disallow them per path. Set a clear policy and document it in robots.txt.
- Block obvious junk routes for every crawler, for example /wp-admin/, /cart, infinite calendars and facet combinations.
If your server is small, crawl spikes can slow real users. Most honest crawlers let you set a crawl rate in their consoles or respect robots crawl-delay when they support it.
Robots.txt, sitemaps and control
robots.txt tells crawlers what they may fetch. Honest crawlers obey it. Put it at yourproduct.com/robots.txt and serve it fast with a 200 status code.
Decide your rules
List the sections to allow and disallow. Keep public content open. Block admin, search results, filters that explode into infinite URLs and test areas.
Write and test robots.txt
Add User-agent blocks with Allow and Disallow lines. Add separate sections for AI crawlers when you have a policy. Test that key pages are allowed.
Expose your sitemap
Link to your XML sitemap in robots.txt with a Sitemap line. Sitemaps help discovery, they do not force indexing.
Verify in consoles
Check Google Search Console and Bing Webmaster Tools for crawl errors and whether robots.txt was fetched and read.
The JavaScript question
A crawler that does not render JavaScript sees only the HTML your server sends. If your content loads only after client-side rendering, some crawlers will miss it.
- Serve primary content in HTML where you can, for example the product name, price and key copy on /pricing.
- Provide links in HTML. Do not hide navigation behind click handlers that need JavaScript to build hrefs.
- If you use hydration or islands, test a page with JavaScript off. If the core text vanishes, add server rendering or a static fallback.
- For pages that must be JS-only, expect some crawlers and tools to skip or misread them. Plan manual submissions or API feeds where offered.
How to see and measure crawling
- Search Console Crawl Stats shows fetch volume, response codes and hosts. Watch for spikes, errors and timeouts.
- URL Inspection shows when Google last crawled a page and if it can index it. Check a sample of key URLs.
- Server logs show every request. Filter on user agents like Googlebot, Bingbot, GPTBot and your audit tool to spot heavy hitters or blocked paths.
- Run your own crawl with an SEO crawler to simulate discovery. Fix broken links and orphan pages it finds before search bots waste budget.
A fixed site loads core HTML, links tie related pages together, and robots.txt lets the right bots in and keeps traps closed.
Questions
It is software that fetches web pages, reads their links and fetches those in turn. Search engines, AI vendors, SEO tools, archives and researchers use crawlers to discover and analyse content.
No. Fetching public pages is allowed in general, but a crawler should obey robots.txt and your terms. Abusive crawling that ignores limits or bypasses controls can breach laws or contracts. If a bot harms your site, block it and contact the operator.
ChatGPT is an AI interface, not a crawler. Some AI vendors run their own crawlers, for example GPTBot and OAI-SearchBot, to fetch content for training or answers. Control them with robots.txt if you have a policy.
WebCrawler as a meta search engine still exists as a brand. It is not the same thing as a web crawler program. When people say web crawler, they mean the software that fetches and follows links.
Publish an XML sitemap and link to it in robots.txt. Add your site to Google Search Console and request indexing for key URLs with URL Inspection. You do not need to submit every page if your internal links are clean.
It depends on site health, popularity and change rate. New or small sites may see sporadic visits. Keep pages fast, link them well and update content to earn steadier crawling.
Sources
Check my site, free
Get a quick read on how bots discover your site with the free check: paste a URL, in about thirty seconds it reports three findings with rival context.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent
Read next
- Glossaryrobots.txt
- GlossaryCrawl budget
- GuideThe Crawl stats report: how much Google fetches and where it struggles
- GuideHow to use the URL Inspection tool in Search Console
- GlossaryJavaScript SEO
- GlossaryGooglebot
- GuideThe best website crawlers for SEO: desktop, cloud, and free tiers
- GlossaryChatGPT search