# robots.txt

robots.txt is a plain text file at yourdomain.com/robots.txt that tells crawlers which URLs they may fetch. It is a crawl allow and disallow list, not an index delete button. For a small site, it keeps junk and admin paths out of crawl while leaving product and docs free to fetch.

Updated 2026-09-13 · Source: https://porteur.ai/glossary/robots-txt

## What robots.txt does and does not do

It controls crawling, not indexing. Disallow stops compliant bots from fetching a URL. If a page is already known from links or a sitemap, some engines may still index the URL without content.

It is public, voluntary and per bot. Each crawler chooses to honour it. Major vendors listed below do, as of 2026.

> You cannot set noindex in robots.txt. Use a meta robots noindex tag or an X Robots Tag header on the URL. Disallow alone will not remove it from results.

## How to check yours

1. **Open the file** Visit yoursite.com/robots.txt. It should load fast with HTTP 200 and Content Type text/plain.
2. **Confirm scope and syntax** Look for User-agent, Disallow, Allow, and Sitemap lines. Wildcards * and $ are supported by Google and Bing. Keep comments short and clear.
3. **Test a URL** In Search Console, use URL inspection on for example /guides/getting-started. Check Crawl allowed and the referencing rule. Repeat for a private path like /admin.
4. **Check server rules** If a URL must be excluded from indexing, put noindex on the page or via X Robots Tag. Keep robots.txt for crawl control only.

If you run a staging or demo subdomain, either password protect it at the server, or return noindex on every page. A robots.txt Disallow on its own is not enough.

## A simple robots.txt for a small site

```txt
# Allow all crawlers to fetch public pages
User-agent: *
Allow: /
Disallow: /admin
Disallow: /cart

# Point crawlers to your sitemap
Sitemap: https://yourproduct.com/sitemap.xml

# Block selected AI training crawlers
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /

# Allow search and answering bots you rely on
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /
```

This keeps admin and cart out of crawl, lets search engines fetch everything else, and declines selected training use. Replace the domains and paths with your own.

## AI crawlers and tokens to know, as of 2026

Blocking a training crawler does not remove pages already collected. Vendors honour robots.txt voluntarily. If you block answering and search crawlers, your pages will not appear in assistants.

- OpenAI: GPTBot, for model training. OAI-SearchBot, for search indexing used in ChatGPT. ChatGPT-User, fetches a page when a user asks.
- Anthropic: ClaudeBot, for training and improving models. Claude-SearchBot, for search indexing. Claude-User, per user fetch. anthropic-ai and Claude-Web are older tokens still seen.
- Google: Googlebot, for Search and sources in AI Overviews. Google-Extended, a control for Gemini training and grounding, not ranking or AI Overviews.
- Perplexity: PerplexityBot, index. Perplexity-User, on demand.
- Apple: Applebot, for Siri and Spotlight. Applebot-Extended, a control for Apple Intelligence training.
- Common Crawl: CCBot, broad web crawl used by many models.
- ByteDance: Bytespider.
- Meta: meta-externalagent for training and meta-externalfetcher.
- Amazon: Amazonbot.
- Cohere: cohere-ai.
- DuckDuckGo: DuckAssistBot.
- Microsoft: Bingbot, for Bing Search and Microsoft Copilot.

## Common traps and how to avoid them

- Disallowing all with User-agent: * then later Allowing a few paths. Many bots follow the most specific rule. Keep public Allow simple. Avoid blanket Disallow on production.
- Blocking CSS or JS folders. That can hurt rendering in Search. Do not Disallow assets needed to render /pricing or /docs.
- Using robots.txt to hide private data. It advertises the paths. Use authentication and proper access controls instead.
- Forgetting the Sitemap line. It helps discovery. Keep it absolute, for example https://yourproduct.com/sitemap.xml.
- Serving 404, 403 or HTML from /robots.txt. Always return a text/plain file with HTTP 200.

## Questions

### What is robots.txt used for?

To tell crawlers which URLs they may fetch. You scope rules per user agent and path. It manages crawl load and keeps non public areas out of fetch.

### Does robots.txt still work?

Yes, as of 2026 major crawlers honour it. It is a voluntary standard. It controls crawling, not indexing, and each vendor may handle edge cases differently.

### Is ignoring robots.txt illegal?

robots.txt is a technical standard, not a law. Legal risk depends on jurisdiction and use. For your site, write clear rules and assume good actors will comply.

### How do I fix a “blocked by robots.txt” error?

Find the matching rule in robots.txt, then remove or relax it. If the URL must be indexed, also ensure there is no noindex. Re test with URL inspection.

### Should I block AI crawlers?

Decide per business goal. Blocking training crawlers like GPTBot and Google Extended can reduce model use of your content. Blocking search and answering bots like OAI SearchBot, Claude SearchBot, Bingbot or Googlebot makes you invisible to assistants.

### Can I noindex a page via robots.txt?

No. robots.txt cannot set noindex. Put a meta robots noindex on the page, or send an X Robots Tag header. Leave crawl allowed until the tag is seen.

## Read next

- [robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow](https://porteur.ai/guides/robots-txt-for-ai-crawlers): Decide which AI crawlers to allow in robots.txt, why it matters, and copy‑paste examples for GPTBot, ClaudeBot, PerplexityBot, Google‑Extended and more.
- [robots.txt tester](https://porteur.ai/tools/robots-txt-tester): Paste a site, paths and a crawler: what its robots.txt allows by the rules Google reads with, the rule that decided each, why it won, and the file mended.
- [AI crawler access checker](https://porteur.ai/tools/ai-crawler-access-checker): Paste a site or a page and see, crawler by crawler, whether it may read it: root and page, the rule that decides, the text without JavaScript, llms.txt.
- [llms.txt generator](https://porteur.ai/tools/llms-txt-generator): Write an llms.txt from your site or from a few fields, in sections, and check the one a site already has against the proposal, link by link.
- [AI Overview](https://porteur.ai/glossary/ai-overview): AI Overviews are Google’s generated answers above results. See when they appear, what they cite, how to measure impact, and how to get cited.
- [Noindex](https://porteur.ai/glossary/noindex): Noindex tells search engines not to index a page. Use it for pages that should never rank, and avoid adding it to pages that earn.
- [robots.txt generator](https://porteur.ai/tools/robots-txt-generator): Build a robots.txt that says which AI crawlers may train on your pages and which may cite them, with the sitemap line and the paths closed to everyone.

Paste your URL to check how your robots.txt handles search and AI bots, and which rivals are more open, in about thirty seconds, free. Free check: https://porteur.ai/
