# AI crawler access checker

Paste your site, or one page of it, and see, crawler by crawler, whether it may read it: the training crawlers, the ones that answer questions with citations, the ones a person sends to open a page. The rule that decides each is quoted, the page's text without JavaScript is counted, and a robots.txt that lets the answering crawlers in is there to copy.

Updated 2026-09-14 · Source: https://porteur.ai/tools/ai-crawler-access-checker

## What it checks

Four things. First, robots.txt, read the way Google reads it: the group that names each crawler, or the group for everyone when none does, the longest matching rule, Allow over Disallow on a tie. Every crawler is tested twice, at the site root and at the page you pasted, and the rule that decided each is quoted, with the group it came from: a site can allow the root and block /docs, and the table shows both. Second, the page itself: its status, its title and h1, the robots meta tag and the X-Robots-Tag header (a noindex takes the page out of the search indexes the assistants read; a nosnippet stops Google quoting it), the language, the structured data, and the words in the HTML before any script runs, because the answering crawlers run none. Third, llms.txt and llms-full.txt, present or not. Fourth, a robots.txt block that allows the answering and search crawlers and blocks the training ones, with your sitemap line, to copy.

The crawlers are listed with what each token governs. That is the part people get wrong: the tokens are not interchangeable, and a line written in 2023 to keep a site out of training data has, on some sites, kept it out of every AI answer since.

## What each token governs

| Crawler | Vendor | Governs |
| --- | --- | --- |
| GPTBot | OpenAI | Training the models. Blocking it does not remove you from ChatGPT's answers. |
| OAI-SearchBot | OpenAI | ChatGPT search: the answers that cite pages. Block it and ChatGPT cannot quote you. |
| ChatGPT-User | OpenAI | A person asking ChatGPT to open a page. Rare to block on purpose. |
| ClaudeBot, anthropic-ai | Anthropic | Training. Claude-SearchBot and Claude-User are the search and the browsing halves. |
| PerplexityBot | Perplexity | Search and answers. Perplexity-User is a person asking it to open a page. |
| Google-Extended | Google | Training Gemini. It affects neither Search nor AI Overviews, which use Googlebot. |
| CCBot | Common Crawl | The public crawl many models train on. |
| Applebot-Extended, Bytespider, meta-externalagent, Amazonbot, cohere-ai | Apple, ByteDance, Meta, Amazon, Cohere | Training, mostly; Amazonbot also feeds Alexa's answers. |

## How to read the result

- **The verdict** says whether the page you pasted is open to the crawlers that answer and search, and whether its text is there to read. Blocked to an answering crawler, noindexed, or empty without JavaScript: each is said in a sentence.
- **Site root and This page** are two columns. A crawler allowed at the root and blocked on the page is blocked on the page; the rule under the verdict is the line to change, and the group it came from says whether the crawler has rules of its own or follows the rules for everyone.
- **Text without JavaScript** is the word count in the raw HTML. Under sixty words on a page that looks full in a browser means the text arrives by script, and the answering crawlers read an empty page.
- **Robots directives on the page** and **Snippets** read the meta tag and the header. index, follow is the default and needs no tag; noindex and nosnippet are the ones to notice.
- **A robots.txt that lets the answering crawlers in** is a block to copy: the search and browsing tokens allowed, the training tokens blocked, your sitemap declared. Keep your own Disallow lines for private paths under each group.

## What to change

1. **Decide per token, not per vendor** Allow OAI-SearchBot, Claude-SearchBot and PerplexityBot if you want to be cited; decide on GPTBot, ClaudeBot and Google-Extended separately, on how you feel about training.
2. **Write the groups explicitly** A "User-agent: *" with "Disallow: /" blocks every crawler you have not named, the answering ones included. Name the ones you allow, and keep the star group for the rest.
3. **Serve the text in the HTML** If the home page is built by a script, render it on the server or export it. The check tells you when it reads empty.
4. **Add llms.txt** Ten lines that say what the site is and which pages to read. The generator is in the tools.
5. **Check again after a deploy** Frameworks and hosts ship default robots files; a redeploy can put back a rule you removed.

## Limits

The tool reads what robots.txt says and what the page carries. A crawler that ignores robots.txt is not stopped by it, a firewall that blocks by user agent is not read here, and a page behind a sign-in is read as the sign-in page. The word count is a measure of text in the HTML, not of quality. Whether an assistant then cites the page depends on what the page says and on where else the product is named, which the free check reads.

## Questions

### Should I block GPTBot?

That is a choice about training, not about being found. Blocking GPTBot does not remove your pages from ChatGPT's answers; blocking OAI-SearchBot does. Decide the two separately.

### Does blocking Google-Extended hurt my rankings?

No. Google-Extended governs training for Gemini. Search and AI Overviews use Googlebot, which the token does not touch.

### Why does the tool say my home page is empty when it looks fine to me?

Because you run JavaScript and most AI crawlers do not. The HTML the server sends has almost no text; the script draws the page after. Server rendering or a static export puts the text in the HTML.

### Can an AI crawler ignore robots.txt?

Yes, technically; the file is a convention. The named crawlers of the large vendors honour it, and honouring it is the only thing this check can read.

### Why test the page and not just the root?

Because robots.txt rules are paths. A site that allows everything at the root can block /guides/ or /docs/ two lines later, and the pages an assistant would cite are rarely the home page. The table shows the root and the page you pasted, with the rule that decided each.

## Read next

- [robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow](https://porteur.ai/guides/robots-txt-for-ai-crawlers): Decide which AI crawlers to allow in robots.txt, why it matters, and copy‑paste examples for GPTBot, ClaudeBot, PerplexityBot, Google‑Extended and more.
- [Generative engine optimization (GEO): how to be named by AI answers](https://porteur.ai/guides/generative-engine-optimization): How to get your site named in AI answers: sources, crawlers, lists, structure, llms.txt, comparisons, and how to measure by asking assistants.
- [SEO for ChatGPT: how assistants choose what to cite](https://porteur.ai/guides/seo-for-chatgpt): How ChatGPT with search finds pages and decides what to cite, what to fix on your site, and how to test it each month in under 30 minutes.
- [How to appear in Google’s AI Overviews](https://porteur.ai/guides/how-to-rank-in-ai-overviews): What overviews cite, which queries trigger them, and how to structure pages that get linked. Steps to find your openings and measure them.
- [robots.txt tester](https://porteur.ai/tools/robots-txt-tester): Paste a site, paths and a crawler: what its robots.txt allows by the rules Google reads with, the rule that decided each, why it won, and the file mended.
- [llms.txt generator](https://porteur.ai/tools/llms-txt-generator): Write an llms.txt from your site or from a few fields, in sections, and check the one a site already has against the proposal, link by link.
- [AI crawler](https://porteur.ai/glossary/ai-crawler): An AI crawler is a bot that fetches your pages for an AI. See the main bots, how to spot them in logs, and how to allow or block them well.

This tells you whether the assistants may read you. The free check tells you whether they name you: the AI answers on your searches, where you are quoted, and where a rival is quoted instead. Free check: https://porteur.ai/
