Orphan pages: finding the pages nothing links to

You have pages no one can reach from your site. Search bots struggle. Users never see them. Here is how to find every orphan page, and what to do next on a small site you ship yourself.

By , founder of Porteur · Updated 13 September 2026 · Markdown

What an orphan page is and why it costs you

An orphan page is a URL on your site with no internal links pointing to it. It can sit in your sitemap or even get traffic from a bookmark. But your own pages do not link to it.

  • Search bots may not find it, or crawl it rarely. Internal links guide crawl and pass context.
  • Even if Google finds it from your sitemap or a backlink, it lacks signals from your site. That can cap its reach.
  • Users cannot click to it. So it cannot help journeys from / to conversion pages.
  • It is a risk for duplicate or stale content. You forget it exists. It rots.

On a small site, each page should pull its weight. You either connect it, combine it, or clear it out.

Why orphan pages happen on small sites

  • Old campaign landers like /spring-offer that you stopped linking after the push.
  • Generated pages from tags, filters or calendars that create thin URLs you never linked in menus.
  • Docs and changelogs moved off the header. The old URLs still live but no nav item points to them.
  • Programmatic pages at scale, like /guides/how-to-x, where some never got linked from hub pages.
  • Design refreshes that drop footer links to policies, regions, or legacy posts.
  • Drafts or test URLs you published and forgot, like /test-variant-b.
  • Product variants or PDPs created by a feed, later removed from category pages.

None of this is weird. It is normal site drift. The fix is a repeatable inventory and triage.

Get your full URL list before you hunt

You need three sources. Your crawl, your XML sitemaps, and Google Search Console. Each sees a different slice. Together they show gaps.

  1. Export your XML sitemap URLs

    Open your main sitemap at /sitemap.xml or the path in robots.txt. Save every listed URL. Include nested sitemaps. Keep a single column of canonical URLs.

  2. Crawl your site from the home page

    Run a full crawl that follows links on your HTML pages. Save the final list of HTML URLs that responded 200. This is your linked set.

  3. Export indexed and discovered URLs from Search Console

    In the Page indexing report, export URLs in valid, valid with warnings, and discovered but not indexed, as of now. Also export any URL list from the Links report, internal links section, if you use it.

  4. Normalise the data

    Lowercase where your server treats URLs case-insensitive. Strip URL parameters you know are tracking only, like utm_source. Keep a clean, deduped list per source.

How to find orphan pages

An orphan is a URL that exists in your sitemap or Search Console, or both, but is missing from your crawl list. That is the core compare.

  1. Compare crawl vs sitemap

    Mark any URL that is in the sitemap but not in your crawl as orphan-candidate. Prioritise clean paths over parameter URLs.

  2. Compare crawl vs Search Console

    Mark any URL Google knows but your crawl missed as orphan-candidate. Include valid and discovered but not indexed. These are often old or thin pages.

  3. Remove false positives

    Exclude URLs behind forms or auth. Exclude intentionally noindexed pages like /cart or /account. Exclude non-HTML assets.

  4. Spot duplicates and canonicals

    If A and B are variants and A is canonical, then B is not an orphan to keep. B is a redirect or noindex job, not a link job.

You now have a list of true orphan pages. Sort by value next: conversions, links, and search potential.

  • Pages with backlinks or saved traffic first. Check if any orphan has external links. Keep and connect those.
  • High intent pages second, like /pricing, /contact, or a feature page. If they are orphaned, link them today.
  • Search matches third. Pages that answer clear queries, like /guides/getting-started, deserve links if they still fit your product.

Watch for noindex, canonicals and blocked paths

Many orphan pages are not link problems. They are index control problems. Check these flags before you add links.

  • Meta robots noindex. If set, the page will not index even if you link it. Only keep noindex for pages like /login or /cart.
  • Canonical points elsewhere. If the orphan has rel=canonical to another URL, decide if that is right. If wrong, fix it before linking.
  • robots.txt blocks crawling. If blocked, Google may not fetch the content. Remove the block for public pages and let your sitemap list them.
  • X-robots-tag at the header level. It can mark whole paths noindex. Check if a CDN rule is catching more than it should.

A monthly sweep in under 30 minutes

You do not need to run a full audit each week. A light routine keeps drift under control.

  1. Run a quick crawl

    Crawl from your home page to catch new internal links and live pages. Save the URL list.

  2. Pull sitemap URLs

    Export the current sitemap URLs. If you manage sitemaps by folder, scan each.

  3. Check Search Console deltas

    Open the Page indexing report. Export valid, and discovered but not indexed. Compare to last month’s export to spot new entries.

  4. Diff and triage

    Find URLs that are in sitemap or Search Console but missing from the crawl. Classify them with the four outcomes above.

  5. Fix the top five

    Add two or three solid links for keepers. Redirect or remove junk. Update the sitemap. Re-crawl to verify.

The fixed site is simple. No dead-end URLs. Hubs list all live pieces. Every key page is two clicks or fewer from the home page, and the sitemap mirrors reality.

Worked example: five orphans on a SaaS site

Say your crawl finds fewer URLs than your sitemap, and Search Console shows more valid plus some discovered but not indexed. That gap surfaces five notable orphans to fix now.

  • /pricing-old. Redirect to /pricing. Update any blog posts that linked the old slug.
  • /guides/getting-started-v1. Merge unique steps into /guides/getting-started. 301 the v1 URL.
  • /blog/launch-2023. Keep. Add links from /changelog and /blog. Add a related links block to point to /pricing and /guides/migration.
  • /integrations/slack. Keep. Link it from /integrations, /features/notifications, and one high-traffic blog post about alerts.
  • /blog/tag/product-updates. Remove tag pages. 410 or noindex, remove from sitemap, and prune any templates that generate them.

Re-crawl. All five now resolve. The two keepers appear with three internal links each. Search Console will reflect the redirects and removals over the next 28 days.

Prevent new orphans as you ship

  • Ship with a hub. Every new guide or feature page must be linked from its hub, like /guides or /features.
  • Set a nav rule. If a page is core, add it to header or footer. Review menus with each redesign.
  • Gate drafts. Use a staging password or noindex until the page is linked. Then remove the block.
  • Automate sitemaps. Generate from your canonical URL list, not from the file system.
  • Use templates. Add related links sections to posts and docs so new pages auto-surface.
  • Redirect on rename. When you change a slug, 301 the old one at deploy time and update internal links in the codebase.

One simple check before you publish helps. Ask: where will a user click to reach this page in two steps or fewer? If you cannot answer, add links now.

Questions

Sources

Check my site, free

Paste your URL to get a free orphan page check that reads your site and the searches around it in about thirty seconds and shows three findings whole.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next