robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow

You decide what AI can fetch, index and train on your pages. This guide shows who the crawlers are, what they do, the trade‑offs for a small product, and exact robots.txt rules.

By , founder of Porteur · Updated 13 September 2026 · Markdown

The short answer for a small product site

If you need visibility in assistants, allow search and answering bots. If you do not want your content used to train models, block the training bots.

  • Allow: Googlebot, Bingbot, PerplexityBot, OAI-SearchBot, Claude-SearchBot, Applebot, DuckAssistBot. These feed search and assistant answers.
  • Allow: the on‑demand fetchers that load your page when a user asks, like ChatGPT-User, Claude-User, Perplexity-User. That is how assistants quote your content.
  • Consider blocking: GPTBot, ClaudeBot, CCBot, meta-externalagent, Applebot-Extended, Google-Extended and other training controls. This reduces training use, not past copies.
  • Do not block your core Search bots unless you accept lower search traffic. Blocking Googlebot or Bingbot hides you from Search and Microsoft Copilot.

You can also split by path. For example, allow /docs and /blog to assistants, and block /app and /pricing from training. The examples below show both.

Who the crawlers are and what they do

These are the tokens you will see in robots.txt and logs, and what each is for as of 2026. Use exact names for User-agent lines.

VendorTokenWhat it is used for
OpenAIGPTBotCrawling for model training
OpenAIOAI-SearchBotIndexing for ChatGPT search results
OpenAIChatGPT-UserFetches a page when a user asks about it
AnthropicClaudeBotCrawling for training and to improve models
AnthropicClaude-SearchBotSearch index for Claude answers
AnthropicClaude-UserFetch on user request
Anthropicanthropic-ai, Claude-WebOlder tokens still seen
GoogleGooglebotGoogle Search, and the pages AI Overviews cite
GoogleGoogle-ExtendedControl for Gemini training and grounding. Does not change Search ranking or AI Overviews
PerplexityPerplexityBotIndex for Perplexity answers
PerplexityPerplexity-UserFetch on user request
AppleApplebotSiri and Spotlight search
AppleApplebot-ExtendedControl for Apple Intelligence training
Common CrawlCCBotOpen crawl used to train many models
ByteDanceBytespiderByteDance crawler
Metameta-externalagentMeta training crawler
Metameta-externalfetcherMeta fetcher
AmazonAmazonbotAmazon crawler
Coherecohere-aiCohere crawler
DuckDuckGoDuckAssistBotCrawler for DuckAssist answers
MicrosoftBingbotBing Search and Microsoft Copilot

Robots.txt basics you will actually use

  • Place robots.txt at the root, for example yourproduct.com/robots.txt.
  • Use a group per crawler when you need different rules. User-agent lines match tokens literally.
  • Disallow blocks crawl of a path. Allow lets a path through within a blocked parent. Keep it simple. Avoid wildcards unless needed.
  • Add your XML sitemap at the end with a full URL. That helps Search, not AI training rules.
  • Rules are read per bot. A block on GPTBot does not affect Googlebot. Be explicit.

Copy and paste robots.txt examples

Start with a clean base. Then add per‑bot groups you want to block or allow. Replace yourproduct.com with your domain and paths you use.

# Base: allow everything to normal search, block nothing by default
User-agent: *
Disallow:

# Sitemaps
Sitemap: https://yourproduct.com/sitemap.xml

Block the common training crawlers. Keep assistants and search open. This is a common small site setup.

# Block training crawlers
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Bytespider
Disallow: /

# Training controls from search vendors
User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

Allow on‑demand user fetchers and search. You usually do not need to add groups to allow. They inherit the base allow. Shown here for clarity only.

# Assistants and search remain allowed via the base group
# Examples of bots that will be allowed:
# Googlebot, Bingbot, PerplexityBot, OAI-SearchBot, Claude-SearchBot,
# ChatGPT-User, Claude-User, Perplexity-User, Applebot, DuckAssistBot

Block AI across a sensitive path, like your app area, while keeping docs and blog open. This mixes Allow on subpaths with a parent Disallow for selected bots.

# Example: block training bots from everything under /, but allow /docs and /blog to be crawled by search
User-agent: GPTBot
Disallow: /
Allow: /docs/
Allow: /blog/

User-agent: ClaudeBot
Disallow: /
Allow: /docs/
Allow: /blog/

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

# Leave search and answer bots on the base allow

Block all AI and assistant activity except classic Search. Only use this if you do not want to appear in assistants at all. You will still be on Google Search and Bing if you keep those open.

# Allow classic search
User-agent: Googlebot
Disallow:

User-agent: Bingbot
Disallow:

# Block assistant ecosystem
User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

User-agent: DuckAssistBot
Disallow: /

# Block broader training sources
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Bytespider
Disallow: /

Only block assistants and training on your pricing, and leave docs open. This is a common SaaS call when you want pricing clicked on your site, not summarised.

# Block assistants and training on /pricing only
User-agent: GPTBot
Disallow: /pricing

User-agent: ClaudeBot
Disallow: /pricing

User-agent: CCBot
Disallow: /pricing

User-agent: OAI-SearchBot
Disallow: /pricing

User-agent: Claude-SearchBot
Disallow: /pricing

User-agent: PerplexityBot
Disallow: /pricing

User-agent: Perplexity-User
Disallow: /pricing

User-agent: ChatGPT-User
Disallow: /pricing

User-agent: Google-Extended
Disallow: /pricing

User-agent: Applebot-Extended
Disallow: /pricing

# Let Googlebot, Bingbot, and Applebot crawl everything (base group open)

How to check what you allow today

  1. Fetch your robots.txt

    Open yourproduct.com/robots.txt in a browser. If it 404s, you have no file. Create one at the web root.

  2. Search for AI tokens

    Scan for GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Claude-SearchBot, ChatGPT-User, Perplexity-User, Google-Extended, Applebot-Extended, CCBot. Note each rule.

  3. Test a path with curl

    Run curl -A "GPTBot" https://yourproduct.com/robots.txt and read what that bot sees. Repeat for ClaudeBot and others. The file is static, but this helps you think per bot.

  4. Check server logs

    Look for recent hits with these user agents. If you see ChatGPT-User or Claude-User, a user asked about your page. If you see GPTBot or CCBot, training crawls are running.

  5. Use a tester

    Run a robots tester on a URL and user‑agent to confirm allow or block. Keep in mind not all testers include every AI bot.

  6. Validate after changes

    Update robots.txt, deploy, and fetch it again. Check that your sitemaps still resolve. Log any 403s or 404s caused by a typo.

You can also run an ai crawler access check from a URL to summarise which AI bots your current robots.txt allows or blocks, and on which paths. That saves guessing when the file is long.

Deciding the right trade‑off for your site

Decide this by page type and by business model. You get reach from search and assistants. You lose some control over how your text is reused in models. Pick a line you can explain to a customer and to your team.

  • If you sell a product and want brand demand, allow assistants on /blog and /guides. Block training across the site if that is your policy.
  • If you run paid content, block both training and assistants on the paid area. Keep Googlebot and Bingbot allowed if you rely on Search for free pages.
  • If you compete with generic summaries, keep /pricing and /compare pages out of assistants. Let docs and integration guides be quoted to earn trust.
  • If you publish research or public docs, allow assistants and on‑demand fetchers site‑wide. Consider leaving training open if you want maximum reach.

Write your policy in a short comment at the top of robots.txt. That way, future you remembers why you blocked GPTBot or allowed PerplexityBot on /docs. A clear file reduces random edits later.

What happens after you change robots.txt

  • Bots see the new rules the next time they fetch robots.txt. Many fetch before they crawl. Expect a delay before crawl patterns change.
  • Search rankings do not change because you set Google-Extended. It controls Gemini training and grounding, not Search.
  • If you block a training crawler, old copies remain with the vendor. Robots.txt does not delete historical data.
  • If you block search or answer bots, expect fewer assistant and search referrals. Watch this in Search Console and in logs for user‑agent hits.
  • If you allow on‑demand fetchers, you may see more user‑agent hits but the traffic is bursty and tied to specific user questions.

A fixed setup looks like this: your /guides pages show up in Perplexity answers with your link, while GPTBot and ClaudeBot stop crawling your site. You keep Googlebot and Bingbot open and your search traffic stays stable.

Implementation notes and pitfalls

  • Robots.txt must be plain text. No redirects if you can avoid them. Serve it at 200 OK.
  • User‑agent tokens are case sensitive in practice. Copy them exactly as vendors document them.
  • Do not rely on Crawl-delay. Many bots ignore it. Prefer Disallow on the heavy paths.
  • A Disallow on a folder does not stop an assistant from fetching a single page when a user clicks a cited link. That fetch uses the -User token group if allowed.
  • Do not add noindex in robots.txt. Use a meta robots tag on the page when you need noindex.
  • If you are on a CDN, purge the robots.txt path after edits. Some CDNs cache it longer than you expect.
  • List all sitemaps with full URLs. This helps search bots find updates. It does not change training access.
  • Keep path rules simple. Avoid complex wildcards. Most AI crawlers respect Disallow and Allow, but not all implement pattern quirks the same way.
  • Older Anthropic tokens like anthropic-ai and Claude-Web still appear in robots files. Include them only if you see them in logs or want to be thorough.

Questions

Check my site, free

Paste your URL into the AI crawler access checker to see in about thirty seconds which AI bots your robots.txt allows or blocks and three findings to fix.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next