Internal link checker
Crawlers find pages by following links, and they weigh a page partly by who links to it. A page listed in the sitemap and linked from nowhere is found late and ranked last. This tool takes a sample of your site, up to sixty pages from the sitemap or from the links themselves, fetches them, reads every internal link, and says who links to whom: the orphans, the pages more than three clicks from home, the least linked, the most linked.
By Théophile Louvart, founder of Porteur · Updated 14 September 2026 · Markdown
What it checks
The tool fetches the home page, then looks for the sitemap where robots.txt says it is, or at /sitemap.xml. When there is one, it takes up to sixty of its addresses spread evenly across the file, so the sample spans the site rather than its first folder. When there is none, it follows the home page's own links, then the links of those pages, until it has sixty. Each page is fetched with a plain crawler's user agent, six at a time, and every internal link on it is read.
From those links it builds the graph of the sample and computes, for each page: how many other pages of the sample link to it, how many distinct internal addresses it links to, and how many clicks it sits from the home page along the sampled links. Then it reports:
- Orphans in the sample: pages no other sampled page links to.
- Pages more than three clicks from the home page, or with no path from it at all.
- The ten least linked and the ten most linked pages.
- The average number of internal links a page carries.
- Pages that did not answer 200: redirects, 404s, timeouts.
Why internal links decide what gets found
A sitemap tells a crawler a page exists. A link tells it the page matters, and what it is about, and hands it some of the standing the linking page has. A page reachable only from the sitemap is crawled less often, indexed later, and holds no standing of its own; Search Console files it under Discovered, currently not indexed, and the fix is never the sitemap.
Depth works the same way. Pages one click from the home page are crawled first and most; pages four clicks down, found through a category, then a listing, then a pagination page, are crawled rarely. On a small site nothing should be more than three clicks from home, and the pages that earn, the pricing page, the best guides, the comparisons, should be one or two.
How to read the result
The verdict names the worst finding: orphans first, depth second, then the good case. Under it, the size of the sample and where it came from, the average links per page, then the tables. Each row gives the page's path, how many sampled pages link to it, how many internal links it carries, and its distance from home in clicks, or no path when the sampled links never reach it.
| Reading | What it means | What to do |
|---|---|---|
| Linked from: 0 | Only the sitemap knows the page | Link it from the pages on its subject, and from a hub |
| Linked from: 1 to 2 | Found through one door | Add a link from the body of two related pages |
| From home: no path | The sampled links never reach it | Same as an orphan: it may be reachable through pages outside the sample, but the sample is a fair read |
| From home: 4 clicks or more | Crawled rarely, weighed lightly | Link it from a hub or from a page one click from home |
| Links out: 0 on a 200 page | A dead end, or a page whose links are drawn by scripts | Add read-next links; check the links are real anchors in the HTML |
| Did not answer 200 | The sitemap lists a redirect or a dead page | Fix the sitemap, then the links pointing there |
A worked example. yourproduct.com has a sitemap of 140 addresses; the tool samples sixty. Four guides under /guides/ are linked from nothing but the guides hub, which itself is three clicks from home because the navigation links to /blog and the hub is reachable only from a footer link on the blog. Verdict: no orphans, six pages more than three clicks from home. The fix is one navigation link to /guides, which pulls the whole family two clicks closer; then two body links to each of the four guides from the guides on the same subject.
What to do with it
Give every orphan two links from its own subject
Not from the footer: from the body of the two pages closest to it, with an anchor that says what the page is. A hub page for the family gives a third.
Bring the earning pages within two clicks
Pricing, the best guides, the comparisons: a link from the navigation, the home page or a hub. Depth is a decision about what matters.
Add a read-next block to templates that dead-end
Guides, changelog entries and docs pages that link nowhere are the usual dead ends. Three related links at the foot of each, chosen by subject, and every page becomes a door.
Clean the sitemap of what did not answer 200
Redirects and dead pages in the sitemap spend the crawl for nothing. List the final address of every page that exists, and only those.
Run the check again in a month
New pages arrive as orphans by default. Make the check part of publishing: a page is not done until two others link to it.
Limits
- A sample of sixty pages, not the whole site. An orphan in the sample may be linked from a page outside it; a page that looks well linked may be linked only by pages inside it. On a site of a few hundred pages the sample is a fair read; on a site of ten thousand it is a sketch.
- Links are read from the HTML as sent. Links drawn by JavaScript after load are not seen, which is also true of the AI crawlers and of Googlebot's first pass.
- Fetching stops after forty seconds, six pages at a time; a slow site gets a smaller sample, and the result says so.
- Five checks an hour per visitor and a shared daily budget, because a check fetches up to sixty pages of someone's site. Nothing is stored.
Questions
A page no other page on the site links to. Crawlers can find it through the sitemap, but it holds no standing of its own and is indexed late or not at all. Search Console usually files orphans under Discovered, currently not indexed.
There is no right number. A page needs enough links in to be found and weighed, two from its own subject at least, and enough links out to lead somewhere. Hundreds of links on every page, from a mega menu, give each one almost nothing.
Google has said that pages closer to the home page are generally considered more important, and they are crawled more often. Keep the pages that earn within two clicks of home and nothing beyond three.
Because the path goes through a page outside the sixty sampled, or through a link drawn by JavaScript. Check the page's inbound count in the table; if it is zero as well, the page is an orphan in practice.
Yes, they are read like any other link, which is why every page shows a handful of inbound links from the start. The pages that stand out are the ones that also get links from the body of related pages.
Check my site, free
This reads who links to whom. The free check reads what the pages are worth: which searches each could take, which two of yours fight over one, and where a link from a stronger page would move a ranking.
- Free check, no card
- Read-only, your own accounts
- Readable by your agent
Read next
- GuideInternal linking for a small site: which pages link to which
- GuideOrphan pages: finding the pages nothing links to
- GuideDiscovered, currently not indexed: why Google has not crawled the page
- GlossaryOrphan page
- GlossaryInternal link
- GlossaryCrawl budget
- Free toolSitemap checker
- Free toolBroken link checker