# robots.txt tester

Fetch a site's robots.txt, test any number of paths against one crawler with the rules Google reads by, and read why each verdict came out as it did: the rule that won, the rules it beat, and the group it came from. The warnings name the mistakes, and a mended copy of the file fixes the ones that can be fixed without changing what it allows.

Updated 2026-09-14 · Source: https://porteur.ai/tools/robots-txt-tester

## What it checks

The tool fetches /robots.txt, parses it into groups, one per set of User-agent lines with their Allow and Disallow rules, and tests each path you give against the crawler you pick from the list (Googlebot, Googlebot-Image, Bingbot, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot, AhrefsBot, everyone, or a token you type). The rules are the ones Google documents: the group that names the crawler wins over the group for everyone; among matching rules the longest path wins; on a tie Allow wins; an asterisk matches anything and a dollar sign anchors the end. For each path the answer comes with the rule that decided it, every rule that matched with its length, and a sentence on why the winner won. It lists the sitemaps the file declares and the mistakes it sees, and writes a mended copy: the byte-order mark dropped, a relative Sitemap line made absolute, a path without its leading slash given one, a noindex line commented out, and a Sitemap line added when there is none.

## Why one file deserves a test

robots.txt is the only file on a site that can make a page disappear from search with one line, and it is edited rarely, by a framework default, a host template or someone who left. "Disallow: /" shipped from staging has taken whole sites out of Google for weeks. A blocked CSS folder makes Google render the page without its layout and judge it on that. A rule written for one crawler with a typo in its name applies to nobody. None of these show in a browser, which is why the file needs a reader.

## How to read the result

- **May fetch** or **may not fetch** for each path, with the rule that decided it, or with no rule matching, which means allowed.
- **Why** says which group answered (the crawler's own, or the group for everyone) and how the winner won: the only match, the longest path, or Allow on a tie. The matching rules are listed with their lengths under it.
- **No robots.txt** means everything is allowed to everyone. Fine for most small sites; add one when you have something to keep out.
- **Groups** shows which crawlers have their own rules, with the count of Allow and Disallow lines. A group with a misspelt crawler name is a group for nobody.
- **The warnings** name the mistakes: Disallow: / for everyone, blocked scripts or styles, a relative Sitemap line, Crawl-delay (which Google ignores), a file past 500 KB, rules before any User-agent, a noindex line.
- **The mended robots.txt** is the file with what can be fixed without changing what it allows, with the list of changes, to copy.

## The mistakes it warns about

| What it sees | Why it matters | What to do |
| --- | --- | --- |
| Disallow: / for every crawler | Nothing may be crawled; the site leaves the index. | Remove it on the live site; keep it on staging only. |
| A blocked css, js or static folder | Google renders pages; without styles and scripts it judges a broken page. | Allow them, or narrow the rule to what is really private. |
| Rules before any User-agent | They apply to nobody. | Move them under the group they were meant for. |
| A byte-order mark at the start | Some readers see a malformed first line. | Save the file as UTF-8 without BOM. |
| Sitemap: /sitemap.xml | A relative path; crawlers ignore it. | Write the full address: Sitemap: https://yourproduct.com/sitemap.xml. |
| Crawl-delay: 10 | Google ignores it; Bing and Yandex honour it. | Leave it if you want to slow Bing; set Google's rate in Search Console. |
| Disallow: private/ | No leading slash; matches nothing. | Disallow: /private/ |
| A file over 500 KB | Google reads the first 500 KB and stops. | Shorten it; robots.txt is for folders, not for every page. |
| No Sitemap line | Crawlers find the sitemap only if told or by guessing the root. | Add Sitemap: with the full address. |

## Limits

The tester reads what the file says. A crawler that ignores robots.txt is not stopped by it; a server that blocks by user agent or by firewall is not read here; a page kept out of the index with a noindex tag is a different mechanism, and a page blocked in robots.txt can still appear in results as an address without a snippet if other sites link to it. To keep a page out of the index, let it be crawled and mark it noindex.

## Questions

### Does robots.txt remove a page from Google?

Not reliably. It stops the crawl; a page that other sites link to can still be listed by its address. To remove a page from the index, allow the crawl and add a noindex tag, or return a 404 or 410.

### What does Disallow with nothing after it mean?

An empty Disallow allows everything for that group. It is the same as no rule.

### Which rule wins when two match?

The longest matching path. On a tie, Allow wins over Disallow. A group that names the crawler wins over the group for everyone, entirely, not rule by rule.

### Should I block AI crawlers in robots.txt?

Decide per token: the training crawlers and the answering crawlers are different tokens from the same vendors, and blocking the answering ones removes you from their citations. The AI crawler access checker lists them with what each governs.

### What does the mended file change?

Only what can be fixed without changing what the file allows: it drops a byte-order mark, makes a relative Sitemap line absolute, gives a path its leading slash, comments out a noindex line Google does not read, marks Crawl-delay as ignored by Google, and adds a Sitemap line when there is none. A Disallow you did not mean is a decision, not a typo; the tool warns and leaves it.

## Read next

- [robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow](https://porteur.ai/guides/robots-txt-for-ai-crawlers): Decide which AI crawlers to allow in robots.txt, why it matters, and copy‑paste examples for GPTBot, ClaudeBot, PerplexityBot, Google‑Extended and more.
- [Technical SEO checklist for a small site](https://porteur.ai/guides/technical-seo-checklist): Run these 20 technical SEO checks, in order. Each shows how to check it free and what fixed looks like for a site under 1,000 pages.
- [Orphan pages: finding the pages nothing links to](https://porteur.ai/guides/orphan-pages): Find orphan pages on your site, why they happen, how to compare crawl, sitemap and Search Console, and what to do: link, merge, redirect or remove.
- [AI crawler access checker](https://porteur.ai/tools/ai-crawler-access-checker): Paste a site or a page and see, crawler by crawler, whether it may read it: root and page, the rule that decides, the text without JavaScript, llms.txt.
- [Sitemap checker](https://porteur.ai/tools/sitemap-checker): Read a sitemap the way a crawler does: what it lists, how it is dated, whether robots.txt names it, and whether thirty of its addresses answer 200.
- [robots.txt](https://porteur.ai/glossary/robots-txt): robots.txt tells crawlers which URLs they may fetch. See what it does not do, how to test it, what to put in it, and how to handle AI bots.
- [Noindex](https://porteur.ai/glossary/noindex): Noindex tells search engines not to index a page. Use it for pages that should never rank, and avoid adding it to pages that earn.
- [Indexation](https://porteur.ai/glossary/indexation): Indexation means a page is stored in Google’s index and can appear in search. Here is how to check status, read Page indexing, and fix what blocks it.

One path against one crawler here. The free check reads the whole site: what a crawler finds on every page, what Google shows of it, and what the AI answers quote on your searches. Free check: https://porteur.ai/
