X-Robots-Tag

X-Robots-Tag is an HTTP response header that tells crawlers how to index a URL. Use it when you need robots directives on files without HTML heads, or when control is easier at the server or edge.

By , founder of Porteur · Updated 14 September 2026 · Markdown

What X-Robots-Tag is

X-Robots-Tag is an HTTP response header carrying the same directives as the meta robots tag. It works for any file type, including PDFs and images.

Typical directives: noindex, nofollow, nosnippet, max-snippet, and unavailable_after. You can scope it to a crawler name too, for example googlebot.

HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex, nofollow

HTTP/1.1 200 OK
Content-Type: text/html; charset=utf-8
X-Robots-Tag: googlebot: noindex

When a small site should use it

Use X-Robots-Tag when you must control indexing for non-HTML files. Think /legal/terms.pdf, /docs/guide.pdf, or image variants you do not want in search.

Use it when templating meta tags is hard. For example, you serve files from a CDN bucket or a static export where adding an HTML head is not possible.

Keep meta robots for normal pages. It is simpler to read in templates like /pricing or /guides/getting-started. Use the header where HTML is not practical.

Where and how to set it

Set the header per path on your server or at the edge. Most stacks let you add response headers by route, file extension or folder prefix.

# Example rules you might apply
# All PDFs under /docs are noindex
X-Robots-Tag: noindex  for path starts_with /docs/ and content-type = application/pdf

# Block only Googlebot for a test image
X-Robots-Tag: googlebot: noindex  for path = /images/test.jpg

If you use a CDN, set it in the CDN behaviour for that path or content type. If you render from an app, add the header in the controller or middleware for those routes.

For static hosts, many allow custom headers via a config file at build. Keep rules narrow so you do not hide your whole site by mistake.

How to read and measure it

First verify the header on the wire. Use curl or your browser’s network panel. You are checking the response of the final URL, after redirects.

curl -I https://yourproduct.com/docs/guide.pdf
# Look for a line like:
# X-Robots-Tag: noindex, nofollow

Then confirm how Google sees it. Use Search Console’s URL Inspection for the exact URL. It shows whether indexing is allowed by a robots meta or header.

Watch coverage for unintended changes. If many pages move to Excluded by ‘noindex’, review recent header deploys by path. Roll back the rule and recrawl a sample.

Common traps to avoid

  • Blocking crawl with robots.txt and expecting noindex to work. The header is only read when the URL is crawlable.
  • Setting the header on a 301 or 302 and assuming it applies after the hop. Check the target URL’s own response.
  • Adding noindex globally on */*.pdf. Audit which PDFs you do want indexed, like product sheets or whitepapers.
  • Forgetting other directives. If you only need to hide a snippet, use nosnippet or a max-snippet value, not noindex.
  • Sending the header with the wrong casing or duplicated values. Check the single, final header Googlebot receives.

Questions

Sources

Check my site, free

Paste one URL and get a free check that reads your site and rivals on your searches, in about thirty seconds, and shows three findings whole.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next