# AI crawler

An AI crawler is a bot that fetches your pages for an AI model, to train it or to answer with it. You care because assistants cite, summarise or ignore your site based on what their crawlers can read. This page shows the names to know, what to measure, and what to do on a small site.

Updated 2026-09-13 · Source: https://porteur.ai/glossary/ai-crawler

## What an AI crawler is and why it matters

An AI crawler is a user agent that reads pages for an AI system. It either builds a training corpus or an index used to answer users. Some also fetch a page on demand when a user asks about it.

For a small site, this affects visibility in assistants and summaries. If search or answering bots cannot read you, you will not show in ChatGPT results, Claude results, Perplexity, Bing Copilot or AI Overviews. If training bots can read you, your words may train models.

## Names and tokens to recognise

These are the vendor tokens, as of 2026, and what each is for. They appear in user agents and in robots rules.

- OpenAI: GPTBot, training. OAI-SearchBot, index for ChatGPT search results. ChatGPT-User, on-demand fetch when a user asks.
- Anthropic: ClaudeBot, training and improving models. Claude-SearchBot, search index. Claude-User, on-demand fetch. Older tokens still seen: anthropic-ai, Claude-Web.
- Google: Googlebot, Search and pages AI Overviews cite. Google-Extended, control for Gemini training and grounding, not ranking and not AI Overviews.
- Perplexity: PerplexityBot, index. Perplexity-User, on demand.
- Apple: Applebot, Siri and Spotlight. Applebot-Extended, control for Apple Intelligence training.
- Microsoft: Bingbot, Bing search and Microsoft Copilot.
- Meta: meta-externalagent, training. meta-externalfetcher, fetching.
- Amazon: Amazonbot.
- Cohere: cohere-ai.
- DuckDuckGo: DuckAssistBot.
- Common Crawl: CCBot, public crawl used to train many models.
- ByteDance: Bytespider.

> Robots.txt is honoured voluntarily. Blocking a training crawler does not remove pages already collected.

## How to measure AI crawler access and impact

Use server logs, CDN logs or your WAF. Filter by user agent tokens above. Confirm the reverse DNS if you need to verify a vendor. Track hits, unique URLs fetched and bandwidth by bot.

Check crawl coverage on key pages: yourproduct.com, /pricing, /docs, /guides/getting-started. If ChatGPT-User or Claude-User hit a page, that reflects real user demand. If only training bots hit you, assistants may still miss you in results.

For Google, use Search Console Crawl stats for Googlebot. For assistants, watch referral spikes from chat.openai.com or claude.ai after on-demand fetches, and watch brand mentions that cite your URLs in their answers.

## What to allow or block on a small site

Decide per goal. If you want visibility in assistants, allow search and answering crawlers. If you do not want your words in training sets, block the training controls where offered.

- Allow for visibility: Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, DuckAssistBot.
- Allow for direct answers: ChatGPT-User, Claude-User, Perplexity-User.
- Optional to block for training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, cohere-ai, Bytespider, Amazonbot. Note the trade off if any vendor also uses a token to improve quality beyond training.
- Set rules in robots.txt by user agent. Use llms.txt if you maintain it for finer guidance. Test with a fetch tool and watch logs for changes after deploy.

> Blocking OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot or Googlebot makes your site invisible to those assistants and to their search results.

A fixed setup could be: allow Googlebot, Bingbot, OAI-SearchBot and Claude-SearchBot across the site, allow ChatGPT-User and Claude-User, and disallow GPTBot and Google-Extended in robots.txt.

## Common traps and how to avoid them

- Relying on robots.txt alone. Some bots ignore it. Set CDN rules if a vendor does not honour robots, and verify IPs if you must block.
- Blocking search bots by accident. A wildcard disallow or a typo in user agent can remove you from assistants. Test on staging and then live.
- Expecting past data removal. Blocking now does not revoke content already used in training or indexing.
- Serving broken pages to AI crawlers. They follow links and render JavaScript. Ensure 200 status, canonical tags and core pages are accessible without login.
- Confusing tokens. anthropic-ai and Claude-Web still appear but are older. Keep rules updated to current tokens.

Porteur reads your site and nearby searches from a URL and shows three findings in about thirty seconds, so you can see crawl access issues fast.

## Questions

### Is ChatGPT a web crawler?

No. ChatGPT is a product. It uses several crawlers. OAI-SearchBot builds the index for ChatGPT search results. ChatGPT-User fetches a page on demand when a user asks about it. GPTBot is for model training.

### Is Google a crawler?

Google Search uses Googlebot to crawl the web. Google-Extended is a separate control for Gemini training and grounding. Allowing or blocking Google-Extended does not change your Search ranking or AI Overviews visibility.

### Will blocking training crawlers remove my content from AI models?

No. Blocking now stops future crawling where honoured but does not remove pages already collected or training already done. Check each vendor’s policy for any separate opt-out forms.

### How do I see if these bots are hitting my site?

Read server or CDN logs and filter by user agent tokens like GPTBot, Claude-SearchBot or Perplexity-User. For vendors that publish reverse DNS, verify hostnames to rule out spoofing. Watch for on-demand fetches shortly after users mention your URLs in chats.

### Should I block all AI crawlers on a new site?

If you need assistant visibility, do not block the search and on-demand crawlers. You can block training crawlers if you prefer. Start with a narrow robots.txt, test, then widen access to the sections you want discovered.

## Read next

- [robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow](https://porteur.ai/guides/robots-txt-for-ai-crawlers): Decide which AI crawlers to allow in robots.txt, why it matters, and copy‑paste examples for GPTBot, ClaudeBot, PerplexityBot, Google‑Extended and more.
- [llms.txt: what it is, examples, and how to write yours](https://porteur.ai/guides/llms-txt): llms.txt is a Markdown file at /llms.txt that maps your site for language models. See the format, a worked example, and write yours in twenty minutes.
- [AI crawler access checker](https://porteur.ai/tools/ai-crawler-access-checker): Paste a site or a page and see, crawler by crawler, whether it may read it: root and page, the rule that decides, the text without JavaScript, llms.txt.
- [How to appear in Google’s AI Overviews](https://porteur.ai/guides/how-to-rank-in-ai-overviews): What overviews cite, which queries trigger them, and how to structure pages that get linked. Steps to find your openings and measure them.
- [Generative engine optimization (GEO): how to be named by AI answers](https://porteur.ai/guides/generative-engine-optimization): How to get your site named in AI answers: sources, crawlers, lists, structure, llms.txt, comparisons, and how to measure by asking assistants.
- [AI Overview](https://porteur.ai/glossary/ai-overview): AI Overviews are Google’s generated answers above results. See when they appear, what they cite, how to measure impact, and how to get cited.

See which AI crawlers can read your site today and who else appears on those searches with a free check from your URL in about thirty seconds. Free check: https://porteur.ai/
