robots.txt

robots.txt is a plain text file at yourdomain.com/robots.txt that tells crawlers which URLs they may fetch. It is a crawl allow and disallow list, not an index delete button. For a small site, it keeps junk and admin paths out of crawl while leaving product and docs free to fetch.

By , founder of Porteur · Updated 13 September 2026 · Markdown

What robots.txt does and does not do

It controls crawling, not indexing. Disallow stops compliant bots from fetching a URL. If a page is already known from links or a sitemap, some engines may still index the URL without content.

It is public, voluntary and per bot. Each crawler chooses to honour it. Major vendors listed below do, as of 2026.

How to check yours

  1. Open the file

    Visit yoursite.com/robots.txt. It should load fast with HTTP 200 and Content Type text/plain.

  2. Confirm scope and syntax

    Look for User-agent, Disallow, Allow, and Sitemap lines. Wildcards * and $ are supported by Google and Bing. Keep comments short and clear.

  3. Test a URL

    In Search Console, use URL inspection on for example /guides/getting-started. Check Crawl allowed and the referencing rule. Repeat for a private path like /admin.

  4. Check server rules

    If a URL must be excluded from indexing, put noindex on the page or via X Robots Tag. Keep robots.txt for crawl control only.

If you run a staging or demo subdomain, either password protect it at the server, or return noindex on every page. A robots.txt Disallow on its own is not enough.

A simple robots.txt for a small site

# Allow all crawlers to fetch public pages
User-agent: *
Allow: /
Disallow: /admin
Disallow: /cart

# Point crawlers to your sitemap
Sitemap: https://yourproduct.com/sitemap.xml

# Block selected AI training crawlers
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /

# Allow search and answering bots you rely on
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /

This keeps admin and cart out of crawl, lets search engines fetch everything else, and declines selected training use. Replace the domains and paths with your own.

AI crawlers and tokens to know, as of 2026

Blocking a training crawler does not remove pages already collected. Vendors honour robots.txt voluntarily. If you block answering and search crawlers, your pages will not appear in assistants.

  • OpenAI: GPTBot, for model training. OAI-SearchBot, for search indexing used in ChatGPT. ChatGPT-User, fetches a page when a user asks.
  • Anthropic: ClaudeBot, for training and improving models. Claude-SearchBot, for search indexing. Claude-User, per user fetch. anthropic-ai and Claude-Web are older tokens still seen.
  • Google: Googlebot, for Search and sources in AI Overviews. Google-Extended, a control for Gemini training and grounding, not ranking or AI Overviews.
  • Perplexity: PerplexityBot, index. Perplexity-User, on demand.
  • Apple: Applebot, for Siri and Spotlight. Applebot-Extended, a control for Apple Intelligence training.
  • Common Crawl: CCBot, broad web crawl used by many models.
  • ByteDance: Bytespider.
  • Meta: meta-externalagent for training and meta-externalfetcher.
  • Amazon: Amazonbot.
  • Cohere: cohere-ai.
  • DuckDuckGo: DuckAssistBot.
  • Microsoft: Bingbot, for Bing Search and Microsoft Copilot.

Common traps and how to avoid them

  • Disallowing all with User-agent: * then later Allowing a few paths. Many bots follow the most specific rule. Keep public Allow simple. Avoid blanket Disallow on production.
  • Blocking CSS or JS folders. That can hurt rendering in Search. Do not Disallow assets needed to render /pricing or /docs.
  • Using robots.txt to hide private data. It advertises the paths. Use authentication and proper access controls instead.
  • Forgetting the Sitemap line. It helps discovery. Keep it absolute, for example https://yourproduct.com/sitemap.xml.
  • Serving 404, 403 or HTML from /robots.txt. Always return a text/plain file with HTTP 200.

Questions

Sources

Check my site, free

Paste your URL to check how your robots.txt handles search and AI bots, and which rivals are more open, in about thirty seconds, free.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next