robots.txt
robots.txt is a plain text file at yourdomain.com/robots.txt that tells crawlers which URLs they may fetch. It is a crawl allow and disallow list, not an index delete button. For a small site, it keeps junk and admin paths out of crawl while leaving product and docs free to fetch.
By Théophile Louvart, founder of Porteur · Updated 13 September 2026 · Markdown
What robots.txt does and does not do
It controls crawling, not indexing. Disallow stops compliant bots from fetching a URL. If a page is already known from links or a sitemap, some engines may still index the URL without content.
It is public, voluntary and per bot. Each crawler chooses to honour it. Major vendors listed below do, as of 2026.
How to check yours
Open the file
Visit yoursite.com/robots.txt. It should load fast with HTTP 200 and Content Type text/plain.
Confirm scope and syntax
Look for User-agent, Disallow, Allow, and Sitemap lines. Wildcards * and $ are supported by Google and Bing. Keep comments short and clear.
Test a URL
In Search Console, use URL inspection on for example /guides/getting-started. Check Crawl allowed and the referencing rule. Repeat for a private path like /admin.
Check server rules
If a URL must be excluded from indexing, put noindex on the page or via X Robots Tag. Keep robots.txt for crawl control only.
If you run a staging or demo subdomain, either password protect it at the server, or return noindex on every page. A robots.txt Disallow on its own is not enough.
A simple robots.txt for a small site
# Allow all crawlers to fetch public pages
User-agent: *
Allow: /
Disallow: /admin
Disallow: /cart
# Point crawlers to your sitemap
Sitemap: https://yourproduct.com/sitemap.xml
# Block selected AI training crawlers
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
# Allow search and answering bots you rely on
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /This keeps admin and cart out of crawl, lets search engines fetch everything else, and declines selected training use. Replace the domains and paths with your own.
AI crawlers and tokens to know, as of 2026
Blocking a training crawler does not remove pages already collected. Vendors honour robots.txt voluntarily. If you block answering and search crawlers, your pages will not appear in assistants.
- OpenAI: GPTBot, for model training. OAI-SearchBot, for search indexing used in ChatGPT. ChatGPT-User, fetches a page when a user asks.
- Anthropic: ClaudeBot, for training and improving models. Claude-SearchBot, for search indexing. Claude-User, per user fetch. anthropic-ai and Claude-Web are older tokens still seen.
- Google: Googlebot, for Search and sources in AI Overviews. Google-Extended, a control for Gemini training and grounding, not ranking or AI Overviews.
- Perplexity: PerplexityBot, index. Perplexity-User, on demand.
- Apple: Applebot, for Siri and Spotlight. Applebot-Extended, a control for Apple Intelligence training.
- Common Crawl: CCBot, broad web crawl used by many models.
- ByteDance: Bytespider.
- Meta: meta-externalagent for training and meta-externalfetcher.
- Amazon: Amazonbot.
- Cohere: cohere-ai.
- DuckDuckGo: DuckAssistBot.
- Microsoft: Bingbot, for Bing Search and Microsoft Copilot.
Common traps and how to avoid them
- Disallowing all with User-agent: * then later Allowing a few paths. Many bots follow the most specific rule. Keep public Allow simple. Avoid blanket Disallow on production.
- Blocking CSS or JS folders. That can hurt rendering in Search. Do not Disallow assets needed to render /pricing or /docs.
- Using robots.txt to hide private data. It advertises the paths. Use authentication and proper access controls instead.
- Forgetting the Sitemap line. It helps discovery. Keep it absolute, for example https://yourproduct.com/sitemap.xml.
- Serving 404, 403 or HTML from /robots.txt. Always return a text/plain file with HTTP 200.
Questions
To tell crawlers which URLs they may fetch. You scope rules per user agent and path. It manages crawl load and keeps non public areas out of fetch.
Yes, as of 2026 major crawlers honour it. It is a voluntary standard. It controls crawling, not indexing, and each vendor may handle edge cases differently.
robots.txt is a technical standard, not a law. Legal risk depends on jurisdiction and use. For your site, write clear rules and assume good actors will comply.
Find the matching rule in robots.txt, then remove or relax it. If the URL must be indexed, also ensure there is no noindex. Re test with URL inspection.
Decide per business goal. Blocking training crawlers like GPTBot and Google Extended can reduce model use of your content. Blocking search and answering bots like OAI SearchBot, Claude SearchBot, Bingbot or Googlebot makes you invisible to assistants.
No. robots.txt cannot set noindex. Put a meta robots noindex on the page, or send an X Robots Tag header. Leave crawl allowed until the tag is seen.
Sources
Check my site, free
Paste your URL to check how your robots.txt handles search and AI bots, and which rivals are more open, in about thirty seconds, free.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent