AI crawler

An AI crawler is a bot that fetches your pages for an AI model, to train it or to answer with it. You care because assistants cite, summarise or ignore your site based on what their crawlers can read. This page shows the names to know, what to measure, and what to do on a small site.

By , founder of Porteur · Updated 13 September 2026 · Markdown

What an AI crawler is and why it matters

An AI crawler is a user agent that reads pages for an AI system. It either builds a training corpus or an index used to answer users. Some also fetch a page on demand when a user asks about it.

For a small site, this affects visibility in assistants and summaries. If search or answering bots cannot read you, you will not show in ChatGPT results, Claude results, Perplexity, Bing Copilot or AI Overviews. If training bots can read you, your words may train models.

Names and tokens to recognise

These are the vendor tokens, as of 2026, and what each is for. They appear in user agents and in robots rules.

  • OpenAI: GPTBot, training. OAI-SearchBot, index for ChatGPT search results. ChatGPT-User, on-demand fetch when a user asks.
  • Anthropic: ClaudeBot, training and improving models. Claude-SearchBot, search index. Claude-User, on-demand fetch. Older tokens still seen: anthropic-ai, Claude-Web.
  • Google: Googlebot, Search and pages AI Overviews cite. Google-Extended, control for Gemini training and grounding, not ranking and not AI Overviews.
  • Perplexity: PerplexityBot, index. Perplexity-User, on demand.
  • Apple: Applebot, Siri and Spotlight. Applebot-Extended, control for Apple Intelligence training.
  • Microsoft: Bingbot, Bing search and Microsoft Copilot.
  • Meta: meta-externalagent, training. meta-externalfetcher, fetching.
  • Amazon: Amazonbot.
  • Cohere: cohere-ai.
  • DuckDuckGo: DuckAssistBot.
  • Common Crawl: CCBot, public crawl used to train many models.
  • ByteDance: Bytespider.

How to measure AI crawler access and impact

Use server logs, CDN logs or your WAF. Filter by user agent tokens above. Confirm the reverse DNS if you need to verify a vendor. Track hits, unique URLs fetched and bandwidth by bot.

Check crawl coverage on key pages: yourproduct.com, /pricing, /docs, /guides/getting-started. If ChatGPT-User or Claude-User hit a page, that reflects real user demand. If only training bots hit you, assistants may still miss you in results.

For Google, use Search Console Crawl stats for Googlebot. For assistants, watch referral spikes from chat.openai.com or claude.ai after on-demand fetches, and watch brand mentions that cite your URLs in their answers.

What to allow or block on a small site

Decide per goal. If you want visibility in assistants, allow search and answering crawlers. If you do not want your words in training sets, block the training controls where offered.

  • Allow for visibility: Googlebot, Bingbot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Applebot, DuckAssistBot.
  • Allow for direct answers: ChatGPT-User, Claude-User, Perplexity-User.
  • Optional to block for training: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, cohere-ai, Bytespider, Amazonbot. Note the trade off if any vendor also uses a token to improve quality beyond training.
  • Set rules in robots.txt by user agent. Use llms.txt if you maintain it for finer guidance. Test with a fetch tool and watch logs for changes after deploy.

A fixed setup could be: allow Googlebot, Bingbot, OAI-SearchBot and Claude-SearchBot across the site, allow ChatGPT-User and Claude-User, and disallow GPTBot and Google-Extended in robots.txt.

Common traps and how to avoid them

  • Relying on robots.txt alone. Some bots ignore it. Set CDN rules if a vendor does not honour robots, and verify IPs if you must block.
  • Blocking search bots by accident. A wildcard disallow or a typo in user agent can remove you from assistants. Test on staging and then live.
  • Expecting past data removal. Blocking now does not revoke content already used in training or indexing.
  • Serving broken pages to AI crawlers. They follow links and render JavaScript. Ensure 200 status, canonical tags and core pages are accessible without login.
  • Confusing tokens. anthropic-ai and Claude-Web still appear but are older. Keep rules updated to current tokens.

Porteur reads your site and nearby searches from a URL and shows three findings in about thirty seconds, so you can see crawl access issues fast.

Questions

Check my site, free

See which AI crawlers can read your site today and who else appears on those searches with a free check from your URL in about thirty seconds.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next