# Sitemap checker

Read a sitemap the way a crawler does: what it lists, how many, how dated, whether robots.txt declares it, and whether thirty of its addresses answer. A sitemap that lists dead or redirecting pages spends the crawl on them, and a small site has no crawl to spare.

Updated 2026-09-14 · Source: https://porteur.ai/tools/sitemap-checker

## What it checks

The tool reads robots.txt first and takes the sitemap it declares, or /sitemap.xml, or the address you paste. If it is an index of sitemaps, it reads the first four and checks that the others answer. It counts the addresses, notes how many carry a lastmod and how many distinct dates there are, lists the folders they fall in, and looks for the mistakes: addresses on another host, addresses listed twice, http addresses on an https site, addresses with a query string, dates that are not dates, dates in the future, and every date the same (the build time, which Google learns to ignore). Then it fetches thirty of the addresses, spread through the file, with HEAD requests, and reports each one's status: 200, a redirect and where to, or an error.

## A checker, not a generator

Most searches here want a sitemap generator. A site built on any modern framework or platform already generates one, and generating a second by hand creates the problem the check finds: two files, one stale. What a site with a sitemap needs is to know whether the file is true. That is this tool.

If the site has no sitemap at all, the fix is in the framework, not in a generator: every one of them writes the file from the routes, and keeps it current on each deploy. Generate once by hand and it is wrong by the second week.

## How to read the result

- **The verdict** says whether the sampled addresses answer, and how many things there are to tidy. Dead addresses first, redirects second, then the dates and the declaration.
- **The notes** are the mistakes in sentences: the sitemap not declared in robots.txt, addresses on the wrong host or scheme, duplicates, query strings, unreadable or future dates, one date for every page.
- **Dated** counts the addresses with a lastmod and the distinct dates among them. Two hundred pages and one date means the file is stamped at build; Google reads lastmod only when it is consistently true.
- **Parts** is the count by first folder, which is the shape of the site: the pages you stand behind, by kind.
- **Sampled address** lists the ones that do not answer 200 directly, with the redirect destination when there is one; the button shows all thirty.

## What to fix

1. **Remove what should not be indexed** Sign-in, account, search, filters, thin tag pages. A sitemap is a list of pages you stand behind.
2. **List final addresses** If a page redirects, list where it lands. If it is gone, remove it.
3. **Make lastmod true** Set it from the page's real change date, or leave it out. Never "now" on everything.
4. **Declare it in robots.txt** A Sitemap: line with the full address, so every crawler finds it without being told.
5. **Split past fifty thousand** The limit is fifty thousand addresses or fifty megabytes a file; past that, an index of sitemaps.

## Limits

The tool reads plain XML. Gzipped sitemaps, RSS feeds used as sitemaps and text sitemaps are not read. It samples thirty addresses, not all; a sitemap of a thousand is judged on thirty spread through it, which finds a broken folder and misses one broken page. An index is read four children deep. Whether Google has indexed the addresses is a different question, answered in Search Console's Sitemaps and Page indexing reports.

## Questions

### Where should a sitemap be?

At /sitemap.xml at the root, and declared in robots.txt with a Sitemap: line. Any address works if it is declared; the root is where crawlers look without being told.

### Should every page be in the sitemap?

Every page you want indexed, and no other. Sign-in, account, search and filter pages do not belong, nor do pages marked noindex.

### Does lastmod matter?

Yes, when it is true. Google uses a truthful lastmod to decide what to recrawl and ignores one that changes on every deploy.

### Do I need a sitemap generator?

Rarely. Every modern framework writes the sitemap from the routes on each build; a generated file goes stale within days. The question worth asking is whether the file the site has is true.

### Why does the tool care about identical lastmod dates?

Because Google uses lastmod only when it finds it consistently accurate. A file where every address carries the deploy time says that nothing and everything changed at once; after a few crawls Google stops reading the field on that site, and a real change no longer gets the early crawl it could have had. Emit the date the page changed, or no date at all.

## Read next

- [Technical SEO checklist for a small site](https://porteur.ai/guides/technical-seo-checklist): Run these 20 technical SEO checks, in order. Each shows how to check it free and what fixed looks like for a site under 1,000 pages.
- [Orphan pages: finding the pages nothing links to](https://porteur.ai/guides/orphan-pages): Find orphan pages on your site, why they happen, how to compare crawl, sitemap and Search Console, and what to do: link, merge, redirect or remove.
- [Canonical tags: what they do and the mistakes that cost rankings](https://porteur.ai/guides/canonical-tag): What a canonical tag does, when to use one, the mistakes that cost rankings, and how to check and fix canonicals on a small site.
- [robots.txt tester](https://porteur.ai/tools/robots-txt-tester): Paste a site, paths and a crawler: what its robots.txt allows by the rules Google reads with, the rule that decided each, why it won, and the file mended.
- [Broken link checker](https://porteur.ai/tools/broken-link-checker): Every link on one page, up to 150, fetched eight at a time: the ones that 404 with the words they carry, the ones that redirect and where to, as a list.
- [XML sitemap](https://porteur.ai/glossary/xml-sitemap): What to put in sitemap.xml, how lastmod works, how to create, validate and submit your XML sitemap, and the traps to avoid.
- [Indexation](https://porteur.ai/glossary/indexation): Indexation means a page is stored in Google’s index and can appear in search. Here is how to check status, read Page indexing, and fix what blocks it.

The sitemap says which pages you stand behind. The free check reads which of them Google shows, which nobody clicks, and which pages other sites link to that no longer answer. Free check: https://porteur.ai/
