robots.txt generator
Choose, vendor by vendor, whether the crawler that trains models and the crawler that answers questions may read your site. The file writes itself, with the sitemap line and the paths closed to everyone, and a plain sentence for every block.
By Théophile Louvart, founder of Porteur · Updated 14 September 2026 · Markdown
| Vendor | Training crawler | Answering and search crawler |
|---|---|---|
| OpenAI (ChatGPT) | ||
| Anthropic (Claude) | ||
| Google (Gemini and Search) | ||
| Microsoft (Bing and Copilot) | none published | |
| Perplexity | none published | |
| Apple | ||
| Meta | ||
| Common Crawl | none published | |
| ByteDance | none published | |
| Amazon | none published | |
| Cohere | none published | |
| DuckDuckGo | none published |
Ticked means allowed. A vendor honours these lines voluntarily; a training crawler blocked today does not give back what it collected yesterday.
User-agent: * Disallow: /api/ Disallow: /admin/ User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: meta-externalagent Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / Sitemap: https://yourproduct.com/sitemap.xml
- User-agent: *
- Everyone may read the site except /api/, /admin/.
- GPTBot
- GPTBot is refused: GPTBot collects pages for training.
- OAI-SearchBot, ChatGPT-User
- OAI-SearchBot and ChatGPT-User follow the rules for everyone: OAI-SearchBot indexes for ChatGPT search; ChatGPT-User fetches a page a user asks about.
- ClaudeBot
- ClaudeBot is refused: ClaudeBot crawls for training and to improve models.
- Claude-SearchBot, Claude-User
- Claude-SearchBot and Claude-User follow the rules for everyone: Claude-SearchBot indexes for Claude web search; Claude-User fetches on request.
- Google-Extended
- Google-Extended is refused: Google-Extended controls use in Gemini training and grounding; it changes nothing in Search or AI Overviews.
- Googlebot
- Googlebot follow the rules for everyone: Googlebot is Google Search, and the pages AI Overviews and AI Mode cite.
- Bingbot
- Bingbot follow the rules for everyone: Bingbot is Bing Search, Microsoft Copilot, Yahoo and part of DuckDuckGo.
- PerplexityBot, Perplexity-User
- PerplexityBot and Perplexity-User follow the rules for everyone: PerplexityBot indexes for Perplexity answers; Perplexity-User fetches on request.
- Applebot-Extended
- Applebot-Extended is refused: Applebot-Extended controls Apple Intelligence training.
- Applebot
- Applebot follow the rules for everyone: Applebot serves Siri and Spotlight search.
- meta-externalagent
- meta-externalagent is refused: meta-externalagent crawls for training.
- meta-externalfetcher
- meta-externalfetcher follow the rules for everyone: meta-externalfetcher fetches pages for Meta products.
- CCBot
- CCBot is refused: CCBot builds the open corpus many models train on.
- Bytespider
- Bytespider is refused: Bytespider is ByteDance’s crawler.
- Amazonbot
- Amazonbot is refused: Amazonbot crawls for Alexa and Amazon services.
- cohere-ai
- cohere-ai is refused: cohere-ai crawls for Cohere models.
- DuckAssistBot
- DuckAssistBot follow the rules for everyone: DuckAssistBot fetches pages for DuckAssist answers.
What it makes
A robots.txt file, ready to save at the root of the site. It opens with the block for everyone (the paths you close, /api/ and /admin/ by default), then one block per AI crawler you refuse, then the sitemap line. Crawlers you allow get no block of their own: they follow the rules for everyone, which is how the file works.
The table lists the vendors that publish crawler tokens as of 2026, in two columns: the crawler that collects pages to train models, and the crawler that indexes for search or fetches a page to answer a question. Most sites want the second and not the first. Three presets set the whole table at once: open to all; answers and search yes, training no; closed to AI. Every switch can then be changed by hand.
| Vendor | Trains on pages | Answers or searches with pages |
|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot (ChatGPT search index), ChatGPT-User (fetch on request) |
| Anthropic | ClaudeBot | Claude-SearchBot (Claude web search), Claude-User (fetch on request) |
| Google-Extended (Gemini training and grounding) | Googlebot (Search, AI Overviews, AI Mode) | |
| Microsoft | none published | Bingbot (Bing, Copilot, Yahoo, part of DuckDuckGo) |
| Perplexity | none published | PerplexityBot, Perplexity-User |
| Apple | Applebot-Extended | Applebot (Siri, Spotlight) |
| Meta | meta-externalagent | meta-externalfetcher |
| Common Crawl, ByteDance, Amazon, Cohere | CCBot, Bytespider, Amazonbot, cohere-ai | none published |
| DuckDuckGo | none published | DuckAssistBot |
Why it matters
The two kinds of crawler do opposite things for a site. A training crawler takes the pages into a corpus; nothing comes back, and the vendor honours the refusal voluntarily. An answering crawler is how a product gets named when someone asks ChatGPT, Claude, Perplexity or Copilot what to use: block it and the site is invisible to that assistant. A file written in a hurry often blocks both, because the tokens look alike.
Two tokens are not AI crawlers at all. Googlebot is Google Search, and the pages AI Overviews and AI Mode cite come from the same index; Bingbot is Bing, Microsoft Copilot, Yahoo and part of DuckDuckGo. Refusing either removes the site from a search engine. The generator warns in red when a stance would do that, and the closed preset keeps both open on purpose.
How to use it
Start from a preset
Answers and search yes, training no is the stance most product sites want: cited by assistants, not fed to their training. Open to all is the stance of a site that wants to be everywhere. Closed to AI keeps only the two search engines.
Adjust a vendor
Tick or untick the two boxes on its row. Ticked means allowed. The file and the explanations under it rewrite as you click.
Close the paths nobody should crawl
One per line: the API, the admin, a staging path, a search results page. Do not close the CSS and JavaScript the pages need to render, and do not close the sitemap.
Give the sitemap URL
The absolute address. Crawlers that read robots.txt find the sitemap there before Search Console ever tells them.
Copy, save, test
Save the file as robots.txt at the root of the site, so it answers at yourproduct.com/robots.txt. Then paste a URL into the robots.txt tester on this site to check what each crawler may read.
A worked example. yourproduct.com wants to be cited by every assistant and trained on by none. The preset writes a block for everyone that closes /api/ and /admin/, then Disallow: / blocks for GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, Bytespider, Amazonbot and cohere-ai, and the sitemap line. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot get no block and follow the rules for everyone. Twelve lines of intent, written in a minute.
How to read the result
Under the file, one sentence per block says what it does in plain words. A red sentence means the stance removes the site from a search engine: read it twice before saving. The rest say which crawler is refused and why the vendor uses it, or that a crawler follows the rules for everyone.
User-agent: *
Disallow: /api/
Disallow: /admin/
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
Sitemap: https://yourproduct.com/sitemap.xmlHow a crawler reads it: it looks for a block with its own token; if one exists, it follows only that block; if none does, it follows the block for everyone. That is why an allowed crawler needs no block, and why a refused one gets Disallow: / and nothing else.
Limits
- Voluntary. The vendors listed say they honour robots.txt; a crawler that does not is stopped by a firewall, not by this file.
- Not retroactive. A training crawler refused today keeps what it collected before.
- Not a lock. robots.txt is public and only asks; a page that must not be read needs authentication or must not be published.
- Tokens change. The list is the one the vendors document as of 2026; check the vendor's page when a new assistant appears.
- Not a noindex. A page blocked here can still appear in Google as a bare URL if other pages link to it; use a noindex tag to keep a page out of results.
Questions
Blocking GPTBot keeps future pages out of OpenAI's training and changes nothing about ChatGPT search, which uses OAI-SearchBot. Most product sites block the first and allow the second. It is a choice about training, not about visibility.
No. Google-Extended controls Gemini training and grounding. AI Overviews and AI Mode cite pages from Google's Search index, read by Googlebot; the only controls there are the snippet directives.
A crawler follows its own block if one exists and the block for everyone otherwise. An allowed crawler is meant to follow the rules for everyone, so writing a block for it would only repeat them.
At the root of the host, so it answers at https://yourproduct.com/robots.txt. A file in a subfolder is not read. Each subdomain has its own file.
No. The file asks; it does not enforce. Crawlers that ignore it are handled at the server or the CDN, by user agent or by address.
Check my site, free
The file says who may read the site. The free check says who does: which assistants and rivals are named on your searches, and which pages Google shows without anyone clicking.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent
Read next
- Free toolrobots.txt tester
- Guiderobots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow
- Glossaryrobots.txt
- GlossaryAI crawler
- Free toolAI crawler access checker
- Free toolllms.txt generator
- GuideHow ChatGPT search picks and cites sources
- GuideBlocked by robots.txt: what Google can still do with the URL