AI crawler
An AI crawler is a bot that fetches your pages for an AI model, to train it or to answer with it. You care because assistants cite, summarise or ignore your site based on what their crawlers can read. This page shows the names to know, what to measure, and what to do on a small site.
By Théophile Louvart, founder of Porteur · Updated 13 September 2026 · Markdown
What an AI crawler is and why it matters
An AI crawler is a user agent that reads pages for an AI system. It either builds a training corpus or an index used to answer users. Some also fetch a page on demand when a user asks about it.
For a small site, this affects visibility in assistants and summaries. If search or answering bots cannot read you, you will not show in ChatGPT results, Claude results, Perplexity, Bing Copilot or AI Overviews. If training bots can read you, your words may train models.
Names and tokens to recognise
These are the vendor tokens, as of 2026, and what each is for. They appear in user agents and in robots rules.
- OpenAI: GPTBot, training. OAI-SearchBot, index for ChatGPT search results. ChatGPT-User, on-demand fetch when a user asks.
- Anthropic: ClaudeBot, training and improving models. Claude-SearchBot, search index. Claude-User, on-demand fetch. Older tokens still seen: anthropic-ai, Claude-Web.
- Google: Googlebot, Search and pages AI Overviews cite. Google-Extended, control for Gemini training and grounding, not ranking and not AI Overviews.
- Perplexity: PerplexityBot, index. Perplexity-User, on demand.
- Apple: Applebot, Siri and Spotlight. Applebot-Extended, control for Apple Intelligence training.
- Microsoft: Bingbot, Bing search and Microsoft Copilot.
- Meta: meta-externalagent, training. meta-externalfetcher, fetching.
- Amazon: Amazonbot.
- Cohere: cohere-ai.
- DuckDuckGo: DuckAssistBot.
- Common Crawl: CCBot, public crawl used to train many models.
- ByteDance: Bytespider.
How to measure AI crawler access and impact
Use server logs, CDN logs or your WAF. Filter by user agent tokens above. Confirm the reverse DNS if you need to verify a vendor. Track hits, unique URLs fetched and bandwidth by bot.
Check crawl coverage on key pages: yourproduct.com, /pricing, /docs, /guides/getting-started. If ChatGPT-User or Claude-User hit a page, that reflects real user demand. If only training bots hit you, assistants may still miss you in results.
For Google, use Search Console Crawl stats for Googlebot. For assistants, watch referral spikes from chat.openai.com or claude.ai after on-demand fetches, and watch brand mentions that cite your URLs in their answers.
What to allow or block on a small site
Decide per goal. If you want visibility in assistants, allow search and answering crawlers. If you do not want your words in training sets, block the training controls where offered.
- Allow for visibility: Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, DuckAssistBot.
- Allow for direct answers: ChatGPT-User, Claude-User, Perplexity-User.
- Optional to block for training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, cohere-ai, Bytespider, Amazonbot. Note the trade off if any vendor also uses a token to improve quality beyond training.
- Set rules in robots.txt by user agent. Use llms.txt if you maintain it for finer guidance. Test with a fetch tool and watch logs for changes after deploy.
A fixed setup could be: allow Googlebot, Bingbot, OAI-SearchBot and Claude-SearchBot across the site, allow ChatGPT-User and Claude-User, and disallow GPTBot and Google-Extended in robots.txt.
Common traps and how to avoid them
- Relying on robots.txt alone. Some bots ignore it. Set CDN rules if a vendor does not honour robots, and verify IPs if you must block.
- Blocking search bots by accident. A wildcard disallow or a typo in user agent can remove you from assistants. Test on staging and then live.
- Expecting past data removal. Blocking now does not revoke content already used in training or indexing.
- Serving broken pages to AI crawlers. They follow links and render JavaScript. Ensure 200 status, canonical tags and core pages are accessible without login.
- Confusing tokens. anthropic-ai and Claude-Web still appear but are older. Keep rules updated to current tokens.
Porteur reads your site and nearby searches from a URL and shows three findings in about thirty seconds, so you can see crawl access issues fast.
Questions
No. ChatGPT is a product. It uses several crawlers. OAI-SearchBot builds the index for ChatGPT search results. ChatGPT-User fetches a page on demand when a user asks about it. GPTBot is for model training.
Google Search uses Googlebot to crawl the web. Google-Extended is a separate control for Gemini training and grounding. Allowing or blocking Google-Extended does not change your Search ranking or AI Overviews visibility.
No. Blocking now stops future crawling where honoured but does not remove pages already collected or training already done. Check each vendor’s policy for any separate opt-out forms.
Read server or CDN logs and filter by user agent tokens like GPTBot, Claude-SearchBot or Perplexity-User. For vendors that publish reverse DNS, verify hostnames to rule out spoofing. Watch for on-demand fetches shortly after users mention your URLs in chats.
If you need assistant visibility, do not block the search and on-demand crawlers. You can block training crawlers if you prefer. Start with a narrow robots.txt, test, then widen access to the sections you want discovered.
Check my site, free
See which AI crawlers can read your site today and who else appears on those searches with a free check from your URL in about thirty seconds.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent
Read next
- Guiderobots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow
- Guidellms.txt: what it is, examples, and how to write yours
- Free toolAI crawler access checker
- GuideHow to appear in Google’s AI Overviews
- GuideGenerative engine optimization (GEO): how to be named by AI answers
- GlossaryAI Overview
- Glossaryrobots.txt
- ComparisonAhrefs vs Screaming Frog: an index against a crawler