robots.txt generator

Choose, vendor by vendor, whether the crawler that trains models and the crawler that answers questions may read your site. The file writes itself, with the sitemap line and the paths closed to everyone, and a plain sentence for every block.

By , founder of Porteur · Updated 14 September 2026 · Markdown

VendorTraining crawlerAnswering and search crawler
OpenAI (ChatGPT)
Anthropic (Claude)
Google (Gemini and Search)
Microsoft (Bing and Copilot)none published
Perplexitynone published
Apple
Meta
Common Crawlnone published
ByteDancenone published
Amazonnone published
Coherenone published
DuckDuckGonone published

Ticked means allowed. A vendor honours these lines voluntarily; a training crawler blocked today does not give back what it collected yesterday.

Save it as robots.txt at the root of the site, then test a URL with the robots.txt tester.
User-agent: *
Disallow: /api/
Disallow: /admin/

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: meta-externalagent
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: cohere-ai
Disallow: /

Sitemap: https://yourproduct.com/sitemap.xml
User-agent: *
Everyone may read the site except /api/, /admin/.
GPTBot
GPTBot is refused: GPTBot collects pages for training.
OAI-SearchBot, ChatGPT-User
OAI-SearchBot and ChatGPT-User follow the rules for everyone: OAI-SearchBot indexes for ChatGPT search; ChatGPT-User fetches a page a user asks about.
ClaudeBot
ClaudeBot is refused: ClaudeBot crawls for training and to improve models.
Claude-SearchBot, Claude-User
Claude-SearchBot and Claude-User follow the rules for everyone: Claude-SearchBot indexes for Claude web search; Claude-User fetches on request.
Google-Extended
Google-Extended is refused: Google-Extended controls use in Gemini training and grounding; it changes nothing in Search or AI Overviews.
Googlebot
Googlebot follow the rules for everyone: Googlebot is Google Search, and the pages AI Overviews and AI Mode cite.
Bingbot
Bingbot follow the rules for everyone: Bingbot is Bing Search, Microsoft Copilot, Yahoo and part of DuckDuckGo.
PerplexityBot, Perplexity-User
PerplexityBot and Perplexity-User follow the rules for everyone: PerplexityBot indexes for Perplexity answers; Perplexity-User fetches on request.
Applebot-Extended
Applebot-Extended is refused: Applebot-Extended controls Apple Intelligence training.
Applebot
Applebot follow the rules for everyone: Applebot serves Siri and Spotlight search.
meta-externalagent
meta-externalagent is refused: meta-externalagent crawls for training.
meta-externalfetcher
meta-externalfetcher follow the rules for everyone: meta-externalfetcher fetches pages for Meta products.
CCBot
CCBot is refused: CCBot builds the open corpus many models train on.
Bytespider
Bytespider is refused: Bytespider is ByteDance’s crawler.
Amazonbot
Amazonbot is refused: Amazonbot crawls for Alexa and Amazon services.
cohere-ai
cohere-ai is refused: cohere-ai crawls for Cohere models.
DuckAssistBot
DuckAssistBot follow the rules for everyone: DuckAssistBot fetches pages for DuckAssist answers.

What it makes

A robots.txt file, ready to save at the root of the site. It opens with the block for everyone (the paths you close, /api/ and /admin/ by default), then one block per AI crawler you refuse, then the sitemap line. Crawlers you allow get no block of their own: they follow the rules for everyone, which is how the file works.

The table lists the vendors that publish crawler tokens as of 2026, in two columns: the crawler that collects pages to train models, and the crawler that indexes for search or fetches a page to answer a question. Most sites want the second and not the first. Three presets set the whole table at once: open to all; answers and search yes, training no; closed to AI. Every switch can then be changed by hand.

VendorTrains on pagesAnswers or searches with pages
OpenAIGPTBotOAI-SearchBot (ChatGPT search index), ChatGPT-User (fetch on request)
AnthropicClaudeBotClaude-SearchBot (Claude web search), Claude-User (fetch on request)
GoogleGoogle-Extended (Gemini training and grounding)Googlebot (Search, AI Overviews, AI Mode)
Microsoftnone publishedBingbot (Bing, Copilot, Yahoo, part of DuckDuckGo)
Perplexitynone publishedPerplexityBot, Perplexity-User
AppleApplebot-ExtendedApplebot (Siri, Spotlight)
Metameta-externalagentmeta-externalfetcher
Common Crawl, ByteDance, Amazon, CohereCCBot, Bytespider, Amazonbot, cohere-ainone published
DuckDuckGonone publishedDuckAssistBot

Why it matters

The two kinds of crawler do opposite things for a site. A training crawler takes the pages into a corpus; nothing comes back, and the vendor honours the refusal voluntarily. An answering crawler is how a product gets named when someone asks ChatGPT, Claude, Perplexity or Copilot what to use: block it and the site is invisible to that assistant. A file written in a hurry often blocks both, because the tokens look alike.

Two tokens are not AI crawlers at all. Googlebot is Google Search, and the pages AI Overviews and AI Mode cite come from the same index; Bingbot is Bing, Microsoft Copilot, Yahoo and part of DuckDuckGo. Refusing either removes the site from a search engine. The generator warns in red when a stance would do that, and the closed preset keeps both open on purpose.

How to use it

  1. Start from a preset

    Answers and search yes, training no is the stance most product sites want: cited by assistants, not fed to their training. Open to all is the stance of a site that wants to be everywhere. Closed to AI keeps only the two search engines.

  2. Adjust a vendor

    Tick or untick the two boxes on its row. Ticked means allowed. The file and the explanations under it rewrite as you click.

  3. Close the paths nobody should crawl

    One per line: the API, the admin, a staging path, a search results page. Do not close the CSS and JavaScript the pages need to render, and do not close the sitemap.

  4. Give the sitemap URL

    The absolute address. Crawlers that read robots.txt find the sitemap there before Search Console ever tells them.

  5. Copy, save, test

    Save the file as robots.txt at the root of the site, so it answers at yourproduct.com/robots.txt. Then paste a URL into the robots.txt tester on this site to check what each crawler may read.

A worked example. yourproduct.com wants to be cited by every assistant and trained on by none. The preset writes a block for everyone that closes /api/ and /admin/, then Disallow: / blocks for GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, Bytespider, Amazonbot and cohere-ai, and the sitemap line. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot get no block and follow the rules for everyone. Twelve lines of intent, written in a minute.

How to read the result

Under the file, one sentence per block says what it does in plain words. A red sentence means the stance removes the site from a search engine: read it twice before saving. The rest say which crawler is refused and why the vendor uses it, or that a crawler follows the rules for everyone.

User-agent: *
Disallow: /api/
Disallow: /admin/

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

Sitemap: https://yourproduct.com/sitemap.xml

How a crawler reads it: it looks for a block with its own token; if one exists, it follows only that block; if none does, it follows the block for everyone. That is why an allowed crawler needs no block, and why a refused one gets Disallow: / and nothing else.

Limits

  • Voluntary. The vendors listed say they honour robots.txt; a crawler that does not is stopped by a firewall, not by this file.
  • Not retroactive. A training crawler refused today keeps what it collected before.
  • Not a lock. robots.txt is public and only asks; a page that must not be read needs authentication or must not be published.
  • Tokens change. The list is the one the vendors document as of 2026; check the vendor's page when a new assistant appears.
  • Not a noindex. A page blocked here can still appear in Google as a bare URL if other pages link to it; use a noindex tag to keep a page out of results.

Questions

Check my site, free

The file says who may read the site. The free check says who does: which assistants and rivals are named on your searches, and which pages Google shows without anyone clicking.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next