# robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow

You decide what AI can fetch, index and train on your pages. This guide shows who the crawlers are, what they do, the trade‑offs for a small product, and exact robots.txt rules.

Updated 2026-09-13 · Source: https://porteur.ai/guides/robots-txt-for-ai-crawlers

## The short answer for a small product site

If you need visibility in assistants, allow search and answering bots. If you do not want your content used to train models, block the training bots.

- Allow: Googlebot, Bingbot, PerplexityBot, OAI-SearchBot, Claude-SearchBot, Applebot, DuckAssistBot. These feed search and assistant answers.
- Allow: the on‑demand fetchers that load your page when a user asks, like ChatGPT-User, Claude-User, Perplexity-User. That is how assistants quote your content.
- Consider blocking: GPTBot, ClaudeBot, CCBot, meta-externalagent, Applebot-Extended, Google-Extended and other training controls. This reduces training use, not past copies.
- Do not block your core Search bots unless you accept lower search traffic. Blocking Googlebot or Bingbot hides you from Search and Microsoft Copilot.

You can also split by path. For example, allow /docs and /blog to assistants, and block /app and /pricing from training. The examples below show both.

## Who the crawlers are and what they do

These are the tokens you will see in robots.txt and logs, and what each is for as of 2026. Use exact names for User-agent lines.

| Vendor | Token | What it is used for |
| --- | --- | --- |
| OpenAI | GPTBot | Crawling for model training |
| OpenAI | OAI-SearchBot | Indexing for ChatGPT search results |
| OpenAI | ChatGPT-User | Fetches a page when a user asks about it |
| Anthropic | ClaudeBot | Crawling for training and to improve models |
| Anthropic | Claude-SearchBot | Search index for Claude answers |
| Anthropic | Claude-User | Fetch on user request |
| Anthropic | anthropic-ai, Claude-Web | Older tokens still seen |
| Google | Googlebot | Google Search, and the pages AI Overviews cite |
| Google | Google-Extended | Control for Gemini training and grounding. Does not change Search ranking or AI Overviews |
| Perplexity | PerplexityBot | Index for Perplexity answers |
| Perplexity | Perplexity-User | Fetch on user request |
| Apple | Applebot | Siri and Spotlight search |
| Apple | Applebot-Extended | Control for Apple Intelligence training |
| Common Crawl | CCBot | Open crawl used to train many models |
| ByteDance | Bytespider | ByteDance crawler |
| Meta | meta-externalagent | Meta training crawler |
| Meta | meta-externalfetcher | Meta fetcher |
| Amazon | Amazonbot | Amazon crawler |
| Cohere | cohere-ai | Cohere crawler |
| DuckDuckGo | DuckAssistBot | Crawler for DuckAssist answers |
| Microsoft | Bingbot | Bing Search and Microsoft Copilot |

> Robots.txt is honoured voluntarily. Blocking a training crawler does not remove pages already collected.

## Training versus answering and search

Training crawlers collect pages to improve and train large models. Answering and search crawlers collect pages to index and cite in answers and search results. On‑demand fetchers load a page at the moment a user asks about it.

- Blocking training can limit reuse of your text in future models. It does not claw back what has already been trained.
- Blocking answering and search makes your site invisible to assistants and their search pages. That includes OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot and Googlebot.
- Allowing the on‑demand fetchers lets assistants quote your current page. This can drive referral clicks when a user opens your link from an answer.

For most small sites, the upside of being cited outweighs the cost of some training. If you publish docs, guides or pricing, you likely want assistants to find them. You can still block GPTBot, ClaudeBot and CCBot to limit training while staying visible in results and answers.

## Robots.txt basics you will actually use

- Place robots.txt at the root, for example yourproduct.com/robots.txt.
- Use a group per crawler when you need different rules. User-agent lines match tokens literally.
- Disallow blocks crawl of a path. Allow lets a path through within a blocked parent. Keep it simple. Avoid wildcards unless needed.
- Add your XML sitemap at the end with a full URL. That helps Search, not AI training rules.
- Rules are read per bot. A block on GPTBot does not affect Googlebot. Be explicit.

> Blocking Googlebot or Bingbot reduces your search traffic and can remove you from assistant answers. Only do this if you accept the trade‑off.

## Copy and paste robots.txt examples

Start with a clean base. Then add per‑bot groups you want to block or allow. Replace yourproduct.com with your domain and paths you use.

```txt
# Base: allow everything to normal search, block nothing by default
User-agent: *
Disallow:

# Sitemaps
Sitemap: https://yourproduct.com/sitemap.xml
```

Block the common training crawlers. Keep assistants and search open. This is a common small site setup.

```txt
# Block training crawlers
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Bytespider
Disallow: /

# Training controls from search vendors
User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /
```

Allow on‑demand user fetchers and search. You usually do not need to add groups to allow. They inherit the base allow. Shown here for clarity only.

```txt
# Assistants and search remain allowed via the base group
# Examples of bots that will be allowed:
# Googlebot, Bingbot, PerplexityBot, OAI-SearchBot, Claude-SearchBot,
# ChatGPT-User, Claude-User, Perplexity-User, Applebot, DuckAssistBot
```

Block AI across a sensitive path, like your app area, while keeping docs and blog open. This mixes Allow on subpaths with a parent Disallow for selected bots.

```txt
# Example: block training bots from everything under /, but allow /docs and /blog to be crawled by search
User-agent: GPTBot
Disallow: /
Allow: /docs/
Allow: /blog/

User-agent: ClaudeBot
Disallow: /
Allow: /docs/
Allow: /blog/

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

# Leave search and answer bots on the base allow
```

Block all AI and assistant activity except classic Search. Only use this if you do not want to appear in assistants at all. You will still be on Google Search and Bing if you keep those open.

```txt
# Allow classic search
User-agent: Googlebot
Disallow:

User-agent: Bingbot
Disallow:

# Block assistant ecosystem
User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

User-agent: DuckAssistBot
Disallow: /

# Block broader training sources
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: Bytespider
Disallow: /
```

Only block assistants and training on your pricing, and leave docs open. This is a common SaaS call when you want pricing clicked on your site, not summarised.

```txt
# Block assistants and training on /pricing only
User-agent: GPTBot
Disallow: /pricing

User-agent: ClaudeBot
Disallow: /pricing

User-agent: CCBot
Disallow: /pricing

User-agent: OAI-SearchBot
Disallow: /pricing

User-agent: Claude-SearchBot
Disallow: /pricing

User-agent: PerplexityBot
Disallow: /pricing

User-agent: Perplexity-User
Disallow: /pricing

User-agent: ChatGPT-User
Disallow: /pricing

User-agent: Google-Extended
Disallow: /pricing

User-agent: Applebot-Extended
Disallow: /pricing

# Let Googlebot, Bingbot, and Applebot crawl everything (base group open)
```

## How to check what you allow today

1. **Fetch your robots.txt** Open yourproduct.com/robots.txt in a browser. If it 404s, you have no file. Create one at the web root.
2. **Search for AI tokens** Scan for GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Claude-SearchBot, ChatGPT-User, Perplexity-User, Google-Extended, Applebot-Extended, CCBot. Note each rule.
3. **Test a path with curl** Run curl -A "GPTBot" https://yourproduct.com/robots.txt and read what that bot sees. Repeat for ClaudeBot and others. The file is static, but this helps you think per bot.
4. **Check server logs** Look for recent hits with these user agents. If you see ChatGPT-User or Claude-User, a user asked about your page. If you see GPTBot or CCBot, training crawls are running.
5. **Use a tester** Run a robots tester on a URL and user‑agent to confirm allow or block. Keep in mind not all testers include every AI bot.
6. **Validate after changes** Update robots.txt, deploy, and fetch it again. Check that your sitemaps still resolve. Log any 403s or 404s caused by a typo.

You can also run an ai crawler access check from a URL to summarise which AI bots your current robots.txt allows or blocks, and on which paths. That saves guessing when the file is long.

## Deciding the right trade‑off for your site

Decide this by page type and by business model. You get reach from search and assistants. You lose some control over how your text is reused in models. Pick a line you can explain to a customer and to your team.

- If you sell a product and want brand demand, allow assistants on /blog and /guides. Block training across the site if that is your policy.
- If you run paid content, block both training and assistants on the paid area. Keep Googlebot and Bingbot allowed if you rely on Search for free pages.
- If you compete with generic summaries, keep /pricing and /compare pages out of assistants. Let docs and integration guides be quoted to earn trust.
- If you publish research or public docs, allow assistants and on‑demand fetchers site‑wide. Consider leaving training open if you want maximum reach.

Write your policy in a short comment at the top of robots.txt. That way, future you remembers why you blocked GPTBot or allowed PerplexityBot on /docs. A clear file reduces random edits later.

## What happens after you change robots.txt

- Bots see the new rules the next time they fetch robots.txt. Many fetch before they crawl. Expect a delay before crawl patterns change.
- Search rankings do not change because you set Google-Extended. It controls Gemini training and grounding, not Search.
- If you block a training crawler, old copies remain with the vendor. Robots.txt does not delete historical data.
- If you block search or answer bots, expect fewer assistant and search referrals. Watch this in Search Console and in logs for user‑agent hits.
- If you allow on‑demand fetchers, you may see more user‑agent hits but the traffic is bursty and tied to specific user questions.

A fixed setup looks like this: your /guides pages show up in Perplexity answers with your link, while GPTBot and ClaudeBot stop crawling your site. You keep Googlebot and Bingbot open and your search traffic stays stable.

## Implementation notes and pitfalls

- Robots.txt must be plain text. No redirects if you can avoid them. Serve it at 200 OK.
- User‑agent tokens are case sensitive in practice. Copy them exactly as vendors document them.
- Do not rely on Crawl-delay. Many bots ignore it. Prefer Disallow on the heavy paths.
- A Disallow on a folder does not stop an assistant from fetching a single page when a user clicks a cited link. That fetch uses the -User token group if allowed.
- Do not add noindex in robots.txt. Use a meta robots tag on the page when you need noindex.
- If you are on a CDN, purge the robots.txt path after edits. Some CDNs cache it longer than you expect.
- List all sitemaps with full URLs. This helps search bots find updates. It does not change training access.
- Keep path rules simple. Avoid complex wildcards. Most AI crawlers respect Disallow and Allow, but not all implement pattern quirks the same way.
- Older Anthropic tokens like anthropic-ai and Claude-Web still appear in robots files. Include them only if you see them in logs or want to be thorough.

## Questions

### What is an AI crawler?

It is a bot that fetches web pages for an AI system. Some build a search index for answers. Others collect text to train models. On‑demand bots fetch a page only when a user asks about it.

### Is ChatGPT a web crawler?

ChatGPT is not a crawler. OpenAI runs several bots. GPTBot crawls for training. OAI-SearchBot builds a search index. ChatGPT-User fetches a page when a user asks about it in ChatGPT.

### Do web crawlers still exist?

Yes. Googlebot and Bingbot crawl for classic search. Newer crawlers add AI uses. PerplexityBot and Claude-SearchBot build answer indexes. GPTBot and ClaudeBot crawl for training. Common Crawl still runs CCBot.

### Should I block Common Crawl?

If you do not want your content in open training corpora, block CCBot. Many models train on Common Crawl. Blocking it does not remove past snapshots.

### Does Google-Extended affect my rankings?

No. It controls whether your content is used in Gemini training and grounding. It does not change your Google Search ranking or whether AI Overviews cite your page.

### What are the main AI platforms to consider in robots.txt?

Start with OpenAI, Anthropic, Perplexity, Google, Apple and Common Crawl. Then add Meta, ByteDance and Microsoft. Include Amazon, Cohere and DuckDuckGo if you see them in logs.

## Read next

- [robots.txt](https://porteur.ai/glossary/robots-txt): robots.txt tells crawlers which URLs they may fetch. See what it does not do, how to test it, what to put in it, and how to handle AI bots.
- [AI crawler access checker](https://porteur.ai/tools/ai-crawler-access-checker): Paste a site or a page and see, crawler by crawler, whether it may read it: root and page, the rule that decides, the text without JavaScript, llms.txt.
- [llms.txt: what it is, examples, and how to write yours](https://porteur.ai/guides/llms-txt): llms.txt is a Markdown file at /llms.txt that maps your site for language models. See the format, a worked example, and write yours in twenty minutes.
- [robots.txt tester](https://porteur.ai/tools/robots-txt-tester): Paste a site, paths and a crawler: what its robots.txt allows by the rules Google reads with, the rule that decided each, why it won, and the file mended.
- [How to appear in Google’s AI Overviews](https://porteur.ai/guides/how-to-rank-in-ai-overviews): What overviews cite, which queries trigger them, and how to structure pages that get linked. Steps to find your openings and measure them.
- [AI crawler](https://porteur.ai/glossary/ai-crawler): An AI crawler is a bot that fetches your pages for an AI. See the main bots, how to spot them in logs, and how to allow or block them well.
- [Claude’s web search: how it fetches, cites and what to allow](https://porteur.ai/guides/claude-web-search-and-your-site): How Claude searches and cites, which Anthropic tokens to allow in robots.txt, and how to check if Claude names your site in answers.
- [robots.txt generator](https://porteur.ai/tools/robots-txt-generator): Build a robots.txt that says which AI crawlers may train on your pages and which may cite them, with the sitemap line and the paths closed to everyone.

Paste your URL into the AI crawler access checker to see in about thirty seconds which AI bots your robots.txt allows or blocks and three findings to fix. Free check: https://porteur.ai/
