robots.txt tester
Fetch a site's robots.txt, test any number of paths against one crawler with the rules Google reads by, and read why each verdict came out as it did: the rule that won, the rules it beat, and the group it came from. The warnings name the mistakes, and a mended copy of the file fixes the ones that can be fixed without changing what it allows.
By Théophile Louvart, founder of Porteur · Updated 14 September 2026 · Markdown
What it checks
The tool fetches /robots.txt, parses it into groups, one per set of User-agent lines with their Allow and Disallow rules, and tests each path you give against the crawler you pick from the list (Googlebot, Googlebot-Image, Bingbot, GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot, AhrefsBot, everyone, or a token you type). The rules are the ones Google documents: the group that names the crawler wins over the group for everyone; among matching rules the longest path wins; on a tie Allow wins; an asterisk matches anything and a dollar sign anchors the end. For each path the answer comes with the rule that decided it, every rule that matched with its length, and a sentence on why the winner won. It lists the sitemaps the file declares and the mistakes it sees, and writes a mended copy: the byte-order mark dropped, a relative Sitemap line made absolute, a path without its leading slash given one, a noindex line commented out, and a Sitemap line added when there is none.
Why one file deserves a test
robots.txt is the only file on a site that can make a page disappear from search with one line, and it is edited rarely, by a framework default, a host template or someone who left. "Disallow: /" shipped from staging has taken whole sites out of Google for weeks. A blocked CSS folder makes Google render the page without its layout and judge it on that. A rule written for one crawler with a typo in its name applies to nobody. None of these show in a browser, which is why the file needs a reader.
How to read the result
- May fetch or may not fetch for each path, with the rule that decided it, or with no rule matching, which means allowed.
- Why says which group answered (the crawler's own, or the group for everyone) and how the winner won: the only match, the longest path, or Allow on a tie. The matching rules are listed with their lengths under it.
- No robots.txt means everything is allowed to everyone. Fine for most small sites; add one when you have something to keep out.
- Groups shows which crawlers have their own rules, with the count of Allow and Disallow lines. A group with a misspelt crawler name is a group for nobody.
- The warnings name the mistakes: Disallow: / for everyone, blocked scripts or styles, a relative Sitemap line, Crawl-delay (which Google ignores), a file past 500 KB, rules before any User-agent, a noindex line.
- The mended robots.txt is the file with what can be fixed without changing what it allows, with the list of changes, to copy.
The mistakes it warns about
| What it sees | Why it matters | What to do |
|---|---|---|
| Disallow: / for every crawler | Nothing may be crawled; the site leaves the index. | Remove it on the live site; keep it on staging only. |
| A blocked css, js or static folder | Google renders pages; without styles and scripts it judges a broken page. | Allow them, or narrow the rule to what is really private. |
| Rules before any User-agent | They apply to nobody. | Move them under the group they were meant for. |
| A byte-order mark at the start | Some readers see a malformed first line. | Save the file as UTF-8 without BOM. |
| Sitemap: /sitemap.xml | A relative path; crawlers ignore it. | Write the full address: Sitemap: https://yourproduct.com/sitemap.xml. |
| Crawl-delay: 10 | Google ignores it; Bing and Yandex honour it. | Leave it if you want to slow Bing; set Google's rate in Search Console. |
| Disallow: private/ | No leading slash; matches nothing. | Disallow: /private/ |
| A file over 500 KB | Google reads the first 500 KB and stops. | Shorten it; robots.txt is for folders, not for every page. |
| No Sitemap line | Crawlers find the sitemap only if told or by guessing the root. | Add Sitemap: with the full address. |
Limits
The tester reads what the file says. A crawler that ignores robots.txt is not stopped by it; a server that blocks by user agent or by firewall is not read here; a page kept out of the index with a noindex tag is a different mechanism, and a page blocked in robots.txt can still appear in results as an address without a snippet if other sites link to it. To keep a page out of the index, let it be crawled and mark it noindex.
Questions
Not reliably. It stops the crawl; a page that other sites link to can still be listed by its address. To remove a page from the index, allow the crawl and add a noindex tag, or return a 404 or 410.
An empty Disallow allows everything for that group. It is the same as no rule.
The longest matching path. On a tie, Allow wins over Disallow. A group that names the crawler wins over the group for everyone, entirely, not rule by rule.
Decide per token: the training crawlers and the answering crawlers are different tokens from the same vendors, and blocking the answering ones removes you from their citations. The AI crawler access checker lists them with what each governs.
Only what can be fixed without changing what the file allows: it drops a byte-order mark, makes a relative Sitemap line absolute, gives a path its leading slash, comments out a noindex line Google does not read, marks Crawl-delay as ignored by Google, and adds a Sitemap line when there is none. A Disallow you did not mean is a decision, not a typo; the tool warns and leaves it.
Check my site, free
One path against one crawler here. The free check reads the whole site: what a crawler finds on every page, what Google shows of it, and what the AI answers quote on your searches.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent