robots.txt tester

Fetch a site's robots.txt, test any number of paths against one crawler with the rules Google reads by, and read why each verdict came out as it did: the rule that won, the rules it beat, and the group it came from. The warnings name the mistakes, and a mended copy of the file fixes the ones that can be fixed without changing what it allows.

By , founder of Porteur · Updated 14 September 2026 · Markdown

Crawler

What it checks

The tool fetches /robots.txt, parses it into groups, one per set of User-agent lines with their Allow and Disallow rules, and tests each path you give against the crawler you pick from the list (Googlebot, Googlebot-Image, Bingbot, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot, AhrefsBot, everyone, or a token you type). The rules are the ones Google documents: the group that names the crawler wins over the group for everyone; among matching rules the longest path wins; on a tie Allow wins; an asterisk matches anything and a dollar sign anchors the end. For each path the answer comes with the rule that decided it, every rule that matched with its length, and a sentence on why the winner won. It lists the sitemaps the file declares and the mistakes it sees, and writes a mended copy: the byte-order mark dropped, a relative Sitemap line made absolute, a path without its leading slash given one, a noindex line commented out, and a Sitemap line added when there is none.

Why one file deserves a test

robots.txt is the only file on a site that can make a page disappear from search with one line, and it is edited rarely, by a framework default, a host template or someone who left. "Disallow: /" shipped from staging has taken whole sites out of Google for weeks. A blocked CSS folder makes Google render the page without its layout and judge it on that. A rule written for one crawler with a typo in its name applies to nobody. None of these show in a browser, which is why the file needs a reader.

How to read the result

  • May fetch or may not fetch for each path, with the rule that decided it, or with no rule matching, which means allowed.
  • Why says which group answered (the crawler's own, or the group for everyone) and how the winner won: the only match, the longest path, or Allow on a tie. The matching rules are listed with their lengths under it.
  • No robots.txt means everything is allowed to everyone. Fine for most small sites; add one when you have something to keep out.
  • Groups shows which crawlers have their own rules, with the count of Allow and Disallow lines. A group with a misspelt crawler name is a group for nobody.
  • The warnings name the mistakes: Disallow: / for everyone, blocked scripts or styles, a relative Sitemap line, Crawl-delay (which Google ignores), a file past 500 KB, rules before any User-agent, a noindex line.
  • The mended robots.txt is the file with what can be fixed without changing what it allows, with the list of changes, to copy.

The mistakes it warns about

What it seesWhy it mattersWhat to do
Disallow: / for every crawlerNothing may be crawled; the site leaves the index.Remove it on the live site; keep it on staging only.
A blocked css, js or static folderGoogle renders pages; without styles and scripts it judges a broken page.Allow them, or narrow the rule to what is really private.
Rules before any User-agentThey apply to nobody.Move them under the group they were meant for.
A byte-order mark at the startSome readers see a malformed first line.Save the file as UTF-8 without BOM.
Sitemap: /sitemap.xmlA relative path; crawlers ignore it.Write the full address: Sitemap: https://yourproduct.com/sitemap.xml.
Crawl-delay: 10Google ignores it; Bing and Yandex honour it.Leave it if you want to slow Bing; set Google's rate in Search Console.
Disallow: private/No leading slash; matches nothing.Disallow: /private/
A file over 500 KBGoogle reads the first 500 KB and stops.Shorten it; robots.txt is for folders, not for every page.
No Sitemap lineCrawlers find the sitemap only if told or by guessing the root.Add Sitemap: with the full address.

Limits

The tester reads what the file says. A crawler that ignores robots.txt is not stopped by it; a server that blocks by user agent or by firewall is not read here; a page kept out of the index with a noindex tag is a different mechanism, and a page blocked in robots.txt can still appear in results as an address without a snippet if other sites link to it. To keep a page out of the index, let it be crawled and mark it noindex.

Questions

Check my site, free

One path against one crawler here. The free check reads the whole site: what a crawler finds on every page, what Google shows of it, and what the AI answers quote on your searches.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next