# Claude’s web search: how it fetches, cites and what to allow

Claude can search the web and cite sources. You want it to fetch your pages when users ask and to list you in answers. This guide shows how Claude searches, which crawlers to allow, and how to write pages it will cite.

Updated 2026-09-14 · Source: https://porteur.ai/guides/claude-web-search-and-your-site

## How Claude web search works today

Claude web search arrived in March 2025. It answers a question, and when it needs the web it searches and reads sources. You see inline citations with links inside the answer.

Two fetch modes matter. Claude-SearchBot crawls to build a search index. Claude-User fetches pages on demand when a user asks about a topic or a URL. ClaudeBot is for training and improving models, not for answering a live question.

Reports suggest Claude’s index uses Brave Search as a base. Anthropic also crawls with Claude-SearchBot. There is no submission form. You get in by being crawlable and by being linked in pages that already rank for the query.

## What Claude cites and why it matters for you

Assistants cite pages that answer the question directly. They prefer clear structure, consistent names for products and people, lists and tables, and figures they can quote. A visible date and an author help disambiguate and build trust.

Claude shows inline citations. If your site is named, users click through from the link. Those visits appear in analytics with referrers like claude.ai or the chat host. You will not get a separate report from Anthropic as of 2026.

For a small site, one named citation on a high intent query can bring meaningful visits. Your job is to publish a page Claude can quote, and to be on the lists and comparisons it already trusts.

## The Anthropic tokens in robots.txt

Three current tokens to know as of 2026. Claude-SearchBot builds Anthropic’s search index. Claude-User fetches a page a user asks about. ClaudeBot crawls for training and model improvement. Older tokens, anthropic-ai and Claude-Web, are still seen in robots files and can be kept for clarity.

```text
# Allow Claude’s answering and search crawlers
User-agent: Claude-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

# Optional: allow or block training crawlers
User-agent: ClaudeBot
Disallow: /

# Legacy tokens some sites still mention
User-agent: anthropic-ai
Allow: /
User-agent: Claude-Web
Allow: /
```

Allow Claude-SearchBot and Claude-User if you want visibility in Claude’s answers. Decide on ClaudeBot based on your training preference. Vendors honour robots.txt voluntarily. Blocking a training crawler does not remove pages already collected.

Keep the file simple. Put these rules at the top of /robots.txt. Test with a plain fetch. Update your allowlist when vendors rename a token. Do not guess a name you have not seen documented.

> Blocking answering and search crawlers makes a site invisible to assistants.

## Training crawls vs answering crawls

Training crawls feed models. Answering crawls fetch or index pages to ground an answer for a user. They are different purposes and often different bots.

For Claude, ClaudeBot is training. Claude-SearchBot is index. Claude-User is on demand fetch. Allow the last two if you want citations and clicks. You can disallow training without losing search visibility in Claude.

Across the field as of 2026, OpenAI separates GPTBot, OAI-SearchBot and ChatGPT-User. Perplexity separates PerplexityBot and Perplexity-User. Google separates Googlebot from Google-Extended. The pattern is similar: answering access is separate from training access. Your robots file should reflect that intent on a vendor by vendor basis.

## Where Claude gets results and how to be there

Claude retrieves from its index and search partners. As reported, Brave Search underpins the index. Claude-SearchBot also crawls. There is no inclusion request. You earn inclusion by being linked and crawlable.

- Publish pages that answer a single question early and plainly. Use headings that match the question.
- Use consistent product and company names. Avoid nicknames on key pages.
- Show a recent date and an author. Keep the page updated.
- Add a short list or table that a model can quote without paraphrase.
- Get listed on independent lists, comparisons and reviews your buyers already read.

Being on the pages that answers already cite is the reliable route. For example, if answers to “best database for small SaaS” often cite a roundup on smalldev.tools, you need a mention there. Then your own “/compare/our-db-vs-x” can be quoted for details and specs.

## Write a page Claude can quote

Start with the buyer’s phrasing. If buyers ask “pricing for YourProduct”, your “/pricing” page should say “YourProduct pricing” in the h1 and in the first sentence. Do not bury the price chart under tabs.

1. **State the answer in the first 2 paragraphs** Put the key figure or choice first. On “/guides/getting-started”, write “Set up YourProduct in 5 minutes. You need an email, a card, and these three steps.”
2. **Name entities consistently** Use the exact product name, not a shorthand. Use the same spelling for models, plans and integrations across pages.
3. **Structure with quotable blocks** Add a list of steps, a bulleted feature list, or a 2-column table of limits. Keep each item concrete and short.
4. **Add a visible date and author** Show “Updated September 2026” and who wrote it. Update the page when details change.
5. **Link the comparisons you want to rank for** Publish “/alternatives/competitor” and “/compare/yourproduct-vs-competitor”. Use honest, sourced tables and clear pro and con lists.

A fixed page looks like “/alternatives/competitor”: an h1 with the exact query, a two paragraph summary, a 6 item pros and cons list, and a table that states plan limits and key specs. Claude can lift lines with confidence and cite you inline.

## Set robots.txt for Claude and other assistants

Treat answering access as the default allow. Then make explicit choices for training crawlers. Keep rules vendor specific. Here is a compact pattern that covers the major assistants as of 2026 without blocking search visibility.

```text
# Claude answering and index
User-agent: Claude-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
# Claude training
User-agent: ClaudeBot
Disallow: /

# OpenAI answering and index
User-agent: ChatGPT-User
Allow: /
User-agent: OAI-SearchBot
Allow: /
# OpenAI training
User-agent: GPTBot
Disallow: /

# Perplexity answering and index
User-agent: Perplexity-User
Allow: /
User-agent: PerplexityBot
Allow: /

# Google Search and Gemini control
User-agent: Googlebot
Allow: /
User-agent: Google-Extended
Disallow: /

# Microsoft Bing and Copilot
User-agent: Bingbot
Allow: /

# Other common crawlers
User-agent: CCBot
Disallow: /
User-agent: Applebot
Allow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: meta-externalagent
Disallow: /
User-agent: meta-externalfetcher
Disallow: /
User-agent: Amazonbot
Allow: /
User-agent: Bytespider
Disallow: /
User-agent: cohere-ai
Disallow: /
User-agent: DuckAssistBot
Allow: /
```

Test with a robots tester and a curl. Watch server logs for the exact user agent string and IP range before you make strict allowlists. Vendors can change user agent formats over time, so review monthly.

## Check what Claude says about your category

You cannot get a citation report from Anthropic as of 2026. So you measure it yourself. Use a fixed list of buyer questions and track who Claude cites over time.

1. **Write 20 buyer questions** Use the wording buyers use. Examples: “best payroll for UK startups”, “how to migrate from X to YourProduct”, “YourProduct pricing”.
2. **Ask Claude on the same cadence** Run the set weekly. Paste the questions as separate prompts. Save the answers with dates.
3. **Log names and links** Record every domain Claude cites for each question. Mark when your domain appears, and on which queries.
4. **Segment analytics by referrer** Create segments for claude.ai and any Anthropic chat host. Compare against chatgpt.com, perplexity.ai, copilot.microsoft.com and gemini.google.com.
5. **Use a tracker if you need scale** Tools like Profound, Peec AI, Otterly, Scrunch, and the AI toolkits in Semrush and Ahrefs, or DataForSEO’s LLM mentions, automate the asking.

A fixed process looks like a sheet with 20 rows for queries and weekly columns for “cited domains”. Your domain is green when named. You add a note with the page cited and any change you made that week.

## Diagnose missing citations and fix pages

If Claude skips your site, start with access and structure. Then fix content. Then earn mentions on third party pages Claude already trusts for the query.

- Check robots.txt for Claude-User and Claude-SearchBot allows.
- Fetch the target URL as Claude-User with curl to confirm access.
- Read the top cited pages. Note the headings, lists, date and author blocks they use.
- Match the question in your h1 and first paragraph. Answer in plain terms in the first 2 paragraphs.
- Add a short list or a table with the exact figures the answer needs.
- Add internal links from related pages so the page is easy to find and crawl.
- Pitch the page to roundups and comparison posts that already rank for the query.

For example, if Claude cites “/compare/tool-a-vs-tool-b” pages and you only have a blog post, publish “/compare/yourproduct-vs-tool-b” with a table of plan limits and a dated summary. Link it from “/alternatives/tool-b”.

## Common edge cases and how to handle them

- Single page apps: assistants cannot read content that only appears after scripts run. Render critical content server side.
- Stale dates: if your page shows 2023, assistants may prefer a 2026 page with the same facts. Update and date visibly.
- Inconsistent naming: if your “Starter” plan is called “Basic” on one page, assistants may misquote. Standardise names first.
- Thin comparison pages: if your compare page only says “we are better”, assistants will skip it. Add sourced specs and limits.
- Country variants: if pricing varies by country, state the region on the page and show the currency in the table.

## Questions

### Does Claude support web search?

Yes. Claude searches the web and shows inline citations with links. It uses a search index built from its own crawl and partners, and it fetches pages on demand when a user asks about a topic or a URL.

### How do I allow Claude but block training?

In robots.txt allow Claude-SearchBot and Claude-User, and disallow ClaudeBot. That keeps your site visible in Claude’s answers while opting out of Anthropic’s training crawl. Vendors honour robots.txt voluntarily and cannot retroactively remove pages already collected.

### Is Claude good at searching the web?

It retrieves from a web index and cites sources in line. It tends to pick pages that answer the question early with clear structure, consistent names, a date and an author. Your job is to publish pages in that shape and to be listed on the review and comparison pages it already cites.

### How does Claude Code web search work?

Claude Code can call web tools to search and fetch when a task needs external context. The same crawl and fetch rules apply. Allow Claude-SearchBot and Claude-User in robots.txt so those tools can read your pages when users ask about your product or category.

### How do I see clicks from Claude in analytics?

Look at referrers. Visits from assistants arrive with domains like claude.ai, chatgpt.com, perplexity.ai, copilot.microsoft.com and gemini.google.com. Google’s AI surfaces arrive as google.com with the rest of Search.

### How do I get Claude to cite my site?

Match the buyer question in your h1 and first paragraph, add a quotable list or table, and keep names consistent. Get listed on third party lists and comparisons that Claude already cites for the query. Make sure Claude-User and Claude-SearchBot are allowed in robots.txt.

## Read next

- [robots.txt for AI crawlers: GPTBot, ClaudeBot, PerplexityBot and what to allow](https://porteur.ai/guides/robots-txt-for-ai-crawlers): Decide which AI crawlers to allow in robots.txt, why it matters, and copy‑paste examples for GPTBot, ClaudeBot, PerplexityBot, Google‑Extended and more.
- [Generative engine optimization (GEO): how to be named by AI answers](https://porteur.ai/guides/generative-engine-optimization): How to get your site named in AI answers: sources, crawlers, lists, structure, llms.txt, comparisons, and how to measure by asking assistants.
- [How ChatGPT search picks and cites sources](https://porteur.ai/guides/how-chatgpt-search-cites-sources): See how ChatGPT search finds pages and cites them, what your page needs to be picked, and how to check if your site is named.
- [How Perplexity picks its sources, and how to become one](https://porteur.ai/guides/how-perplexity-picks-sources): How Perplexity finds, indexes and cites pages, what it favours, the crawler tokens to allow, how to see your mentions, and a monthly GEO routine.
- [Google AI Mode: what it is and how a small site gets cited](https://porteur.ai/guides/google-ai-mode): What AI Mode is, how it picks citations, how clicks appear in Search Console, and the page changes that make a small site quotable.
- [How to track your AI visibility without a subscription](https://porteur.ai/guides/ai-visibility-tracking): Track AI visibility without tools: list 20 buyer questions, ask each assistant, log names and links, segment referrers, and know when automation is worth it.
- [Gemini and your site: grounding, Google-Extended, and what to allow](https://porteur.ai/guides/gemini-and-your-site): What AI Mode and Gemini are, how grounding and Google-Extended work, what to allow in robots.txt, and the simplest way to be cited.

Paste your homepage URL and get a free check of how assistants see your niche, which rivals they cite, and three fixes to make your next Claude citation more likely. Free check: https://porteur.ai/
