Sitemap checker

Read a sitemap the way a crawler does: what it lists, how many, how dated, whether robots.txt declares it, and whether thirty of its addresses answer. A sitemap that lists dead or redirecting pages spends the crawl on them, and a small site has no crawl to spare.

By , founder of Porteur · Updated 14 September 2026 · Markdown

What it checks

The tool reads robots.txt first and takes the sitemap it declares, or /sitemap.xml, or the address you paste. If it is an index of sitemaps, it reads the first four and checks that the others answer. It counts the addresses, notes how many carry a lastmod and how many distinct dates there are, lists the folders they fall in, and looks for the mistakes: addresses on another host, addresses listed twice, http addresses on an https site, addresses with a query string, dates that are not dates, dates in the future, and every date the same (the build time, which Google learns to ignore). Then it fetches thirty of the addresses, spread through the file, with HEAD requests, and reports each one's status: 200, a redirect and where to, or an error.

A checker, not a generator

Most searches here want a sitemap generator. A site built on any modern framework or platform already generates one, and generating a second by hand creates the problem the check finds: two files, one stale. What a site with a sitemap needs is to know whether the file is true. That is this tool.

If the site has no sitemap at all, the fix is in the framework, not in a generator: every one of them writes the file from the routes, and keeps it current on each deploy. Generate once by hand and it is wrong by the second week.

How to read the result

  • The verdict says whether the sampled addresses answer, and how many things there are to tidy. Dead addresses first, redirects second, then the dates and the declaration.
  • The notes are the mistakes in sentences: the sitemap not declared in robots.txt, addresses on the wrong host or scheme, duplicates, query strings, unreadable or future dates, one date for every page.
  • Dated counts the addresses with a lastmod and the distinct dates among them. Two hundred pages and one date means the file is stamped at build; Google reads lastmod only when it is consistently true.
  • Parts is the count by first folder, which is the shape of the site: the pages you stand behind, by kind.
  • Sampled address lists the ones that do not answer 200 directly, with the redirect destination when there is one; the button shows all thirty.

What to fix

  1. Remove what should not be indexed

    Sign-in, account, search, filters, thin tag pages. A sitemap is a list of pages you stand behind.

  2. List final addresses

    If a page redirects, list where it lands. If it is gone, remove it.

  3. Make lastmod true

    Set it from the page's real change date, or leave it out. Never "now" on everything.

  4. Declare it in robots.txt

    A Sitemap: line with the full address, so every crawler finds it without being told.

  5. Split past fifty thousand

    The limit is fifty thousand addresses or fifty megabytes a file; past that, an index of sitemaps.

Limits

The tool reads plain XML. Gzipped sitemaps, RSS feeds used as sitemaps and text sitemaps are not read. It samples thirty addresses, not all; a sitemap of a thousand is judged on thirty spread through it, which finds a broken folder and misses one broken page. An index is read four children deep. Whether Google has indexed the addresses is a different question, answered in Search Console's Sitemaps and Page indexing reports.

Questions

Check my site, free

The sitemap says which pages you stand behind. The free check reads which of them Google shows, which nobody clicks, and which pages other sites link to that no longer answer.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next