AI crawler access checker
Paste your site, or one page of it, and see, crawler by crawler, whether it may read it: the training crawlers, the ones that answer questions with citations, the ones a person sends to open a page. The rule that decides each is quoted, the page's text without JavaScript is counted, and a robots.txt that lets the answering crawlers in is there to copy.
By Théophile Louvart, founder of Porteur · Updated 14 September 2026 · Markdown
What it checks
Four things. First, robots.txt, read the way Google reads it: the group that names each crawler, or the group for everyone when none does, the longest matching rule, Allow over Disallow on a tie. Every crawler is tested twice, at the site root and at the page you pasted, and the rule that decided each is quoted, with the group it came from: a site can allow the root and block /docs, and the table shows both. Second, the page itself: its status, its title and h1, the robots meta tag and the X-Robots-Tag header (a noindex takes the page out of the search indexes the assistants read; a nosnippet stops Google quoting it), the language, the structured data, and the words in the HTML before any script runs, because the answering crawlers run none. Third, llms.txt and llms-full.txt, present or not. Fourth, a robots.txt block that allows the answering and search crawlers and blocks the training ones, with your sitemap line, to copy.
The crawlers are listed with what each token governs. That is the part people get wrong: the tokens are not interchangeable, and a line written in 2023 to keep a site out of training data has, on some sites, kept it out of every AI answer since.
What each token governs
| Crawler | Vendor | Governs |
|---|---|---|
| GPTBot | OpenAI | Training the models. Blocking it does not remove you from ChatGPT's answers. |
| OAI-SearchBot | OpenAI | ChatGPT search: the answers that cite pages. Block it and ChatGPT cannot quote you. |
| ChatGPT-User | OpenAI | A person asking ChatGPT to open a page. Rare to block on purpose. |
| ClaudeBot, anthropic-ai | Anthropic | Training. Claude-SearchBot and Claude-User are the search and the browsing halves. |
| PerplexityBot | Perplexity | Search and answers. Perplexity-User is a person asking it to open a page. |
| Google-Extended | Training Gemini. It affects neither Search nor AI Overviews, which use Googlebot. | |
| CCBot | Common Crawl | The public crawl many models train on. |
| Applebot-Extended, Bytespider, meta-externalagent, Amazonbot, cohere-ai | Apple, ByteDance, Meta, Amazon, Cohere | Training, mostly; Amazonbot also feeds Alexa's answers. |
How to read the result
- The verdict says whether the page you pasted is open to the crawlers that answer and search, and whether its text is there to read. Blocked to an answering crawler, noindexed, or empty without JavaScript: each is said in a sentence.
- Site root and This page are two columns. A crawler allowed at the root and blocked on the page is blocked on the page; the rule under the verdict is the line to change, and the group it came from says whether the crawler has rules of its own or follows the rules for everyone.
- Text without JavaScript is the word count in the raw HTML. Under sixty words on a page that looks full in a browser means the text arrives by script, and the answering crawlers read an empty page.
- Robots directives on the page and Snippets read the meta tag and the header. index, follow is the default and needs no tag; noindex and nosnippet are the ones to notice.
- A robots.txt that lets the answering crawlers in is a block to copy: the search and browsing tokens allowed, the training tokens blocked, your sitemap declared. Keep your own Disallow lines for private paths under each group.
What to change
Decide per token, not per vendor
Allow OAI-SearchBot, Claude-SearchBot and PerplexityBot if you want to be cited; decide on GPTBot, ClaudeBot and Google-Extended separately, on how you feel about training.
Write the groups explicitly
A "User-agent: *" with "Disallow: /" blocks every crawler you have not named, the answering ones included. Name the ones you allow, and keep the star group for the rest.
Serve the text in the HTML
If the home page is built by a script, render it on the server or export it. The check tells you when it reads empty.
Add llms.txt
Ten lines that say what the site is and which pages to read. The generator is in the tools.
Check again after a deploy
Frameworks and hosts ship default robots files; a redeploy can put back a rule you removed.
Limits
The tool reads what robots.txt says and what the page carries. A crawler that ignores robots.txt is not stopped by it, a firewall that blocks by user agent is not read here, and a page behind a sign-in is read as the sign-in page. The word count is a measure of text in the HTML, not of quality. Whether an assistant then cites the page depends on what the page says and on where else the product is named, which the free check reads.
Questions
That is a choice about training, not about being found. Blocking GPTBot does not remove your pages from ChatGPT's answers; blocking OAI-SearchBot does. Decide the two separately.
No. Google-Extended governs training for Gemini. Search and AI Overviews use Googlebot, which the token does not touch.
Because you run JavaScript and most AI crawlers do not. The HTML the server sends has almost no text; the script draws the page after. Server rendering or a static export puts the text in the HTML.
Yes, technically; the file is a convention. The named crawlers of the large vendors honour it, and honouring it is the only thing this check can read.
Because robots.txt rules are paths. A site that allows everything at the root can block /guides/ or /docs/ two lines later, and the pages an assistant would cite are rarely the home page. The table shows the root and the page you pasted, with the rule that decided each.
Check my site, free
This tells you whether the assistants may read you. The free check tells you whether they name you: the AI answers on your searches, where you are quoted, and where a rival is quoted instead.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent
Read next
- Guiderobots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow
- GuideGenerative engine optimization (GEO): how to be named by AI answers
- GuideSEO for ChatGPT: how assistants choose what to cite
- GuideHow to appear in Google’s AI Overviews
- Free toolrobots.txt tester
- Free toolllms.txt generator
- GlossaryAI crawler
- Glossaryrobots.txt