Blocked by robots.txt: what Google can still do with the URL

You saw “Blocked by robots.txt” or “Indexed, though blocked by robots.txt” in Search Console. Here is what each status means, why the second exists, and how to fix your file without hurting real pages.

By , founder of Porteur · Updated 14 September 2026 · Markdown

The two statuses, in plain words

Search Console’s Page indexing report splits known URLs into Indexed and Not indexed, each with a reason. Two reasons here matter.

  • Blocked by robots.txt: your robots.txt file disallows Googlebot from fetching the URL. Google will not crawl the content.
  • Indexed, though blocked by robots.txt: the URL is disallowed for crawling, but Google still added the URL to the index from links or a sitemap. It indexed the URL without fetching the content.

You can get both on the same site. A disallow stops crawling. It does not guarantee that the URL will not enter the index if Google finds strong signals to it.

Why Google can index a URL you block from crawling

Indexing and crawling are separate. Crawling fetches the content. Indexing stores a URL and what Google knows about it. Links alone can be enough to index a URL stub.

If /guides/getting-started is disallowed but many pages link to it, Google may index the URL. It will show minimal data and may rank poorly. The Page indexing report will list it as Indexed, though blocked by robots.txt.

Check your robots.txt and test a single URL

  1. Fetch the file

    Open yourproduct.com/robots.txt in a browser. If you get 404, there is no file. If you get 200, read the rules in plain text.

  2. Find the rule that applies

    Look for a User-agent line that matches Googlebot. Rules under it apply. If none, the default is allow. A Disallow: / blocks everything for that agent.

  3. Test the URL in Search Console

    Use URL Inspection on the exact URL. Check Crawling allowed and Indexing allowed. Click Test live URL to fetch now. Review View tested page for response and blocked resources.

  4. Confirm the report reason

    Open the Page indexing report. Filter to the reason. Review the example URLs. Use Validate fix when you have changed the file.

If you run a site-wide change, retest one public page like /pricing and one private page like /admin/settings. That catches broad misrules fast.

Common accidents that trigger these statuses

  • A staging Disallow: / copied to production. All pages become Blocked by robots.txt.
  • Blocking CSS or JS folders, for example Disallow: /assets/. Google cannot render layout or text. Pages can index, but render poorly.
  • Blocking sitemaps with Disallow: /sitemap.xml or placing the sitemap under a disallowed path. Discovery weakens.
  • Over-broad wildcard rules, like Disallow: /*? blocking all parameters including pagination or filters you want crawled.
  • Blocking media or fonts needed for page render, like Disallow: /static/ or Disallow: /fonts/.

Each can leave indexed stubs or tank how your pages render in Test live URL. Both slow down fixes because Google cannot fetch what you blocked.

What to do on a small site

Keep robots.txt simple. Allow Google to crawl pages and assets. Only disallow true system paths that should never be crawled, like /admin/ or /cart/checkout if they have no public content.

  • For pages you want out of Google, allow crawling and add noindex via a meta robots tag or an X-Robots-Tag header.
  • For render assets, allow crawling. That includes CSS, JS, images, fonts and inline API endpoints needed for HTML render.
  • For staging, add HTTP auth, add noindex, and, if you like, disallow crawling too. Do not ship the same file to production.
  • For sitemaps, reference them in robots.txt with a Sitemap: line or submit in Search Console. Do not disallow them.

A fixed /robots.txt on yourproduct.com is short, names the sitemap, allows assets, and only disallows true dead ends. Your public pages crawl and render in Test live URL without missing resources.

Clean patterns you can copy

User-agent: *
# Allow core assets
Allow: /assets/
Allow: /static/

# Keep admin out of crawl, but add noindex in HTML too
Disallow: /admin/
Disallow: /checkout/

# List your sitemap(s)
Sitemap: https://yourproduct.com/sitemap.xml

Then implement noindex on the pages that must stay out of search. For HTML pages, add a meta robots noindex. For PDFs and other files, set an X-Robots-Tag: noindex header. Both require that Google can fetch the URL at least once.

Use Search Console to triage and recheck

Start in the Page indexing report. It lists example URLs for each reason, up to 1,000. Work through the sample to spot a pattern, for example all /static/ blocked or everything under /blog/ disallowed.

  • Fix the robots.txt rule and deploy.
  • Use URL Inspection on one affected URL. Test live URL to confirm crawling allowed and that key resources load.
  • If a page should be indexed, request indexing. This queues a crawl. There is a small daily quota per property and no guarantee of indexing.
  • Back in the Page indexing report, click Validate fix on the reason. The re-check can take about two weeks.

If a URL truly should not exist, let it return 404. That status is correct. Fix internal links and sitemaps that still point to it or redirect to the nearest equivalent if it had traffic and links before.

Edge cases: parameters, duplicates and redirects

Blocking all parameters with a wildcard often prevents crawl of useful pages. If /docs?page=2 is blocked by Disallow: /*?, Google may index a stub from links and you will not get proper pagination signals. Scope the pattern or allow the needed paths.

If you see Duplicate without user-selected canonical together with robots disallows on alternates, remove the disallow. Let Google crawl duplicates so it can respect rel=canonical. Otherwise it may pick a different canonical than you want.

If a URL redirects, Page with redirect is normal. Make sure your internal links and sitemap use the final URL to avoid crawl waste. Do not add a disallow on the old URL. Let Googlebot follow the redirect to update its index.

How to keep a page out of Google the right way

  • Public but not to be indexed, like /search-results or /thank-you: allow crawling and add noindex. Keep it linked only where needed.
  • Private or sensitive, like staging or internal dashboards: put behind authentication. Optionally add noindex. You can also return 401 or 403. Do not rely on robots.txt.
  • Gone pages: return 404 or 410. Remove links. If it had value, redirect to a close match and update links and sitemaps.

If you must remove a previously indexed URL, allow one crawl with noindex, or serve an HTTP header X-Robots-Tag: noindex for non-HTML. After Google confirms Excluded by noindex tag, you can disallow future crawls if you want to save crawl budget.

Questions

Sources

Check my site, free

Check your robots.txt and indexation from a URL in about thirty seconds, free, and see three concrete findings you can ship today with Porteur.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next