# robots.txt generator

Choose, vendor by vendor, whether the crawler that trains models and the crawler that answers questions may read your site. The file writes itself, with the sitemap line and the paths closed to everyone, and a plain sentence for every block.

Updated 2026-09-14 · Source: https://porteur.ai/tools/robots-txt-generator

## What it makes

A robots.txt file, ready to save at the root of the site. It opens with the block for everyone (the paths you close, /api/ and /admin/ by default), then one block per AI crawler you refuse, then the sitemap line. Crawlers you allow get no block of their own: they follow the rules for everyone, which is how the file works.

The table lists the vendors that publish crawler tokens as of 2026, in two columns: the crawler that collects pages to train models, and the crawler that indexes for search or fetches a page to answer a question. Most sites want the second and not the first. Three presets set the whole table at once: open to all; answers and search yes, training no; closed to AI. Every switch can then be changed by hand.

| Vendor | Trains on pages | Answers or searches with pages |
| --- | --- | --- |
| OpenAI | GPTBot | OAI-SearchBot (ChatGPT search index), ChatGPT-User (fetch on request) |
| Anthropic | ClaudeBot | Claude-SearchBot (Claude web search), Claude-User (fetch on request) |
| Google | Google-Extended (Gemini training and grounding) | Googlebot (Search, AI Overviews, AI Mode) |
| Microsoft | none published | Bingbot (Bing, Copilot, Yahoo, part of DuckDuckGo) |
| Perplexity | none published | PerplexityBot, Perplexity-User |
| Apple | Applebot-Extended | Applebot (Siri, Spotlight) |
| Meta | meta-externalagent | meta-externalfetcher |
| Common Crawl, ByteDance, Amazon, Cohere | CCBot, Bytespider, Amazonbot, cohere-ai | none published |
| DuckDuckGo | none published | DuckAssistBot |

## Why it matters

The two kinds of crawler do opposite things for a site. A training crawler takes the pages into a corpus; nothing comes back, and the vendor honours the refusal voluntarily. An answering crawler is how a product gets named when someone asks ChatGPT, Claude, Perplexity or Copilot what to use: block it and the site is invisible to that assistant. A file written in a hurry often blocks both, because the tokens look alike.

Two tokens are not AI crawlers at all. Googlebot is Google Search, and the pages AI Overviews and AI Mode cite come from the same index; Bingbot is Bing, Microsoft Copilot, Yahoo and part of DuckDuckGo. Refusing either removes the site from a search engine. The generator warns in red when a stance would do that, and the closed preset keeps both open on purpose.

> Google-Extended is the one that confuses people: it controls whether pages may be used for Gemini training and grounding, and changes nothing in Search or AI Overviews. Blocking it costs nothing in Google Search; it does not keep a site out of AI Overviews either.

## How to use it

1. **Start from a preset** Answers and search yes, training no is the stance most product sites want: cited by assistants, not fed to their training. Open to all is the stance of a site that wants to be everywhere. Closed to AI keeps only the two search engines.
2. **Adjust a vendor** Tick or untick the two boxes on its row. Ticked means allowed. The file and the explanations under it rewrite as you click.
3. **Close the paths nobody should crawl** One per line: the API, the admin, a staging path, a search results page. Do not close the CSS and JavaScript the pages need to render, and do not close the sitemap.
4. **Give the sitemap URL** The absolute address. Crawlers that read robots.txt find the sitemap there before Search Console ever tells them.
5. **Copy, save, test** Save the file as robots.txt at the root of the site, so it answers at yourproduct.com/robots.txt. Then paste a URL into the robots.txt tester on this site to check what each crawler may read.

A worked example. yourproduct.com wants to be cited by every assistant and trained on by none. The preset writes a block for everyone that closes /api/ and /admin/, then Disallow: / blocks for GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, meta-externalagent, CCBot, Bytespider, Amazonbot and cohere-ai, and the sitemap line. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot get no block and follow the rules for everyone. Twelve lines of intent, written in a minute.

## How to read the result

Under the file, one sentence per block says what it does in plain words. A red sentence means the stance removes the site from a search engine: read it twice before saving. The rest say which crawler is refused and why the vendor uses it, or that a crawler follows the rules for everyone.

```text
User-agent: *
Disallow: /api/
Disallow: /admin/

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

Sitemap: https://yourproduct.com/sitemap.xml
```

How a crawler reads it: it looks for a block with its own token; if one exists, it follows only that block; if none does, it follows the block for everyone. That is why an allowed crawler needs no block, and why a refused one gets Disallow: / and nothing else.

## Limits

- Voluntary. The vendors listed say they honour robots.txt; a crawler that does not is stopped by a firewall, not by this file.
- Not retroactive. A training crawler refused today keeps what it collected before.
- Not a lock. robots.txt is public and only asks; a page that must not be read needs authentication or must not be published.
- Tokens change. The list is the one the vendors document as of 2026; check the vendor's page when a new assistant appears.
- Not a noindex. A page blocked here can still appear in Google as a bare URL if other pages link to it; use a noindex tag to keep a page out of results.

## Questions

### Should I block GPTBot?

Blocking GPTBot keeps future pages out of OpenAI's training and changes nothing about ChatGPT search, which uses OAI-SearchBot. Most product sites block the first and allow the second. It is a choice about training, not about visibility.

### Does blocking Google-Extended remove me from AI Overviews?

No. Google-Extended controls Gemini training and grounding. AI Overviews and AI Mode cite pages from Google's Search index, read by Googlebot; the only controls there are the snippet directives.

### Why does an allowed crawler get no block?

A crawler follows its own block if one exists and the block for everyone otherwise. An allowed crawler is meant to follow the rules for everyone, so writing a block for it would only repeat them.

### Where does the file go?

At the root of the host, so it answers at https://yourproduct.com/robots.txt. A file in a subfolder is not read. Each subdomain has its own file.

### Will this stop scrapers that ignore robots.txt?

No. The file asks; it does not enforce. Crawlers that ignore it are handled at the server or the CDN, by user agent or by address.

## Read next

- [robots.txt tester](https://porteur.ai/tools/robots-txt-tester): Paste a site, paths and a crawler: what its robots.txt allows by the rules Google reads with, the rule that decided each, why it won, and the file mended.
- [robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow](https://porteur.ai/guides/robots-txt-for-ai-crawlers): Decide which AI crawlers to allow in robots.txt, why it matters, and copy‑paste examples for GPTBot, ClaudeBot, PerplexityBot, Google‑Extended and more.
- [robots.txt](https://porteur.ai/glossary/robots-txt): robots.txt tells crawlers which URLs they may fetch. See what it does not do, how to test it, what to put in it, and how to handle AI bots.
- [AI crawler](https://porteur.ai/glossary/ai-crawler): An AI crawler is a bot that fetches your pages for an AI. See the main bots, how to spot them in logs, and how to allow or block them well.
- [AI crawler access checker](https://porteur.ai/tools/ai-crawler-access-checker): Paste a site or a page and see, crawler by crawler, whether it may read it: root and page, the rule that decides, the text without JavaScript, llms.txt.
- [llms.txt generator](https://porteur.ai/tools/llms-txt-generator): Write an llms.txt from your site or from a few fields, in sections, and check the one a site already has against the proposal, link by link.
- [How ChatGPT search picks and cites sources](https://porteur.ai/guides/how-chatgpt-search-cites-sources): See how ChatGPT search finds pages and cites them, what your page needs to be picked, and how to check if your site is named.
- [Blocked by robots.txt: what Google can still do with the URL](https://porteur.ai/guides/blocked-by-robots-txt): What “Blocked by robots.txt” and “Indexed, though blocked by robots.txt” mean, how to test a URL, the common accidents, and the clean fixes.

The file says who may read the site. The free check says who does: which assistants and rivals are named on your searches, and which pages Google shows without anyone clicking. Free check: https://porteur.ai/
