AI crawler access checker

Paste your site, or one page of it, and see, crawler by crawler, whether it may read it: the training crawlers, the ones that answer questions with citations, the ones a person sends to open a page. The rule that decides each is quoted, the page's text without JavaScript is counted, and a robots.txt that lets the answering crawlers in is there to copy.

By , founder of Porteur · Updated 14 September 2026 · Markdown

What it checks

Four things. First, robots.txt, read the way Google reads it: the group that names each crawler, or the group for everyone when none does, the longest matching rule, Allow over Disallow on a tie. Every crawler is tested twice, at the site root and at the page you pasted, and the rule that decided each is quoted, with the group it came from: a site can allow the root and block /docs, and the table shows both. Second, the page itself: its status, its title and h1, the robots meta tag and the X-Robots-Tag header (a noindex takes the page out of the search indexes the assistants read; a nosnippet stops Google quoting it), the language, the structured data, and the words in the HTML before any script runs, because the answering crawlers run none. Third, llms.txt and llms-full.txt, present or not. Fourth, a robots.txt block that allows the answering and search crawlers and blocks the training ones, with your sitemap line, to copy.

The crawlers are listed with what each token governs. That is the part people get wrong: the tokens are not interchangeable, and a line written in 2023 to keep a site out of training data has, on some sites, kept it out of every AI answer since.

What each token governs

CrawlerVendorGoverns
GPTBotOpenAITraining the models. Blocking it does not remove you from ChatGPT's answers.
OAI-SearchBotOpenAIChatGPT search: the answers that cite pages. Block it and ChatGPT cannot quote you.
ChatGPT-UserOpenAIA person asking ChatGPT to open a page. Rare to block on purpose.
ClaudeBot, anthropic-aiAnthropicTraining. Claude-SearchBot and Claude-User are the search and the browsing halves.
PerplexityBotPerplexitySearch and answers. Perplexity-User is a person asking it to open a page.
Google-ExtendedGoogleTraining Gemini. It affects neither Search nor AI Overviews, which use Googlebot.
CCBotCommon CrawlThe public crawl many models train on.
Applebot-Extended, Bytespider, meta-externalagent, Amazonbot, cohere-aiApple, ByteDance, Meta, Amazon, CohereTraining, mostly; Amazonbot also feeds Alexa's answers.

How to read the result

  • The verdict says whether the page you pasted is open to the crawlers that answer and search, and whether its text is there to read. Blocked to an answering crawler, noindexed, or empty without JavaScript: each is said in a sentence.
  • Site root and This page are two columns. A crawler allowed at the root and blocked on the page is blocked on the page; the rule under the verdict is the line to change, and the group it came from says whether the crawler has rules of its own or follows the rules for everyone.
  • Text without JavaScript is the word count in the raw HTML. Under sixty words on a page that looks full in a browser means the text arrives by script, and the answering crawlers read an empty page.
  • Robots directives on the page and Snippets read the meta tag and the header. index, follow is the default and needs no tag; noindex and nosnippet are the ones to notice.
  • A robots.txt that lets the answering crawlers in is a block to copy: the search and browsing tokens allowed, the training tokens blocked, your sitemap declared. Keep your own Disallow lines for private paths under each group.

What to change

  1. Decide per token, not per vendor

    Allow OAI-SearchBot, Claude-SearchBot and PerplexityBot if you want to be cited; decide on GPTBot, ClaudeBot and Google-Extended separately, on how you feel about training.

  2. Write the groups explicitly

    A "User-agent: *" with "Disallow: /" blocks every crawler you have not named, the answering ones included. Name the ones you allow, and keep the star group for the rest.

  3. Serve the text in the HTML

    If the home page is built by a script, render it on the server or export it. The check tells you when it reads empty.

  4. Add llms.txt

    Ten lines that say what the site is and which pages to read. The generator is in the tools.

  5. Check again after a deploy

    Frameworks and hosts ship default robots files; a redeploy can put back a rule you removed.

Limits

The tool reads what robots.txt says and what the page carries. A crawler that ignores robots.txt is not stopped by it, a firewall that blocks by user agent is not read here, and a page behind a sign-in is read as the sign-in page. The word count is a measure of text in the HTML, not of quality. Whether an assistant then cites the page depends on what the page says and on where else the product is named, which the free check reads.

Questions

Check my site, free

This tells you whether the assistants may read you. The free check tells you whether they name you: the AI answers on your searches, where you are quoted, and where a rival is quoted instead.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next