# XML sitemap

An XML sitemap is a machine readable list of URLs you want crawled and indexed. It helps search engines find your important pages and their latest change dates. For a small site, it speeds discovery and keeps stale URLs out of the crawl queue.

Updated 2026-09-13 · Source: https://porteur.ai/glossary/xml-sitemap

## Why XML sitemaps matter for a small site

Google can find pages from links. Your sitemap removes guesswork: it lists the canonical URLs that matter, and when each changed. On a small site you often add or rename pages by hand. A tidy sitemap means fewer missed pages and faster re crawls after edits.

## What belongs in sitemap.xml, and what does not

- Include only canonical, indexable URLs that return 200: for example https://yourproduct.com/pricing, not /pricing?ref=twitter.
- Include unique content pages: home, features, pricing, docs like /guides/getting-started, blog posts, critical category pages.
- Use lastmod to signal updates. Update the date when the primary content changes, not for minor CSS or typo fixes.
- Exclude noindex, 404, 410, redirecting, duplicate, staging and parameter URLs.
- Do not list pages blocked in robots.txt if you want them indexed later. Remove the block or leave them out.
- Keep it small and clean. If your sitemap grows, split it into multiple files and reference them from a sitemap index as per sitemaps.org. This is the standard approach, typical, not measured on your site.

```xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://yourproduct.com/pricing</loc>
    <lastmod>2026-10-10</lastmod>
  </url>
  <url>
    <loc>https://yourproduct.com/guides/getting-started</loc>
    <lastmod>2026-11-15</lastmod>
  </url>
</urlset>
```

> The schema version in the namespace is the standard one for sitemaps, typical, not measured from your site.

## How to create, host and validate your sitemap

1. **Create the file** Most CMSs generate sitemap.xml for you. If not, script it from your URL list, or export from your router. Keep absolute URLs and ISO 8601 dates.
2. **Host it at a stable URL** Serve it at https://yourproduct.com/sitemap.xml. If you use multiple sitemaps, create https://yourproduct.com/sitemap_index.xml referencing them.
3. **Declare it in robots.txt** Add a line: Sitemap: https://yourproduct.com/sitemap.xml. This helps crawlers discover it without submission.
4. **Validate the XML** Open it in a browser to catch obvious errors. Then use a sitemap validation tool or your own XML linter to confirm it follows the sitemaps.org schema.
5. **Keep it fresh** Automate updates when you publish, unpublish or change canonical URLs. Remove dead URLs within a day of deprecating a page.

## Submit and read it in Search Console

In Google Search Console, open Sitemaps. Enter the sitemap URL and submit. Google will fetch it and report discovery and parsing. Use the report to spot invalid URLs, last fetch time and any crawl issues tied to sitemap URLs.

- Check coverage for sitemap URLs in the Indexing reports. Confirm the pages you listed are indexed.
- If a URL shows Alternate page with proper canonical tag, remove the non canonical variant from the sitemap.
- If you change structure, resubmit. You do not need to resubmit after routine updates, Google refetches.
- Track key pages like /pricing and /blog in the report after content updates. A correct setup shows them discovered quickly and indexed soon after.

## Common errors and quick fixes

- Including noindex or redirected URLs. Fix by removing them and updating internal links.
- Wrong canonical URLs. Ensure each sitemap URL has a canonical tag pointing to itself and returns 200.
- Stale lastmod on active pages. Wire lastmod to your publish or update timestamp for primary content.
- Mixing http and https or www and non www. Match your canonical scheme and host.
- Large, slow files. Compress with gzip and split into logical sitemaps, for example /sitemap-pages.xml and /sitemap-posts.xml.
- Blocked by robots.txt. If the sitemap URL or its targets are disallowed, lift the block or the sitemap will be ignored.

> A fixed sitemap lists only live, canonical URLs you actually want indexed, with accurate lastmod dates and no duplicates.

## Questions

### What does an XML sitemap do?

It lists the URLs you want search engines to crawl and index, plus optional metadata like lastmod. It improves discovery, especially for new, deep or orphaned pages. It does not force indexing, but it gives clear crawl hints.

### Is sitemap.xml still relevant?

Yes, as of 2026 it is part of standard crawling. Google can find pages without it, but a clean sitemap speeds discovery and reduces crawl waste. It is low effort and high signal for small sites.

### How do I create an XML sitemap?

If your CMS has one, enable it and check the output. Otherwise generate it from your routing table or database, write XML that follows sitemaps.org, host it at /sitemap.xml, declare it in robots.txt, then validate and submit in Search Console.

### How do I submit an XML sitemap to Google?

Open Google Search Console, choose your property, go to Sitemaps, enter the sitemap URL and submit. Google fetches it and shows status, errors and last read time. You can also rely on robots.txt discovery, but submission gives feedback.

### How often should lastmod change?

Update lastmod when the primary content changes, for example a rewritten /pricing page or a new section on a guide. Do not bump dates for minor style edits or tracking changes. Keep dates truthful, or crawlers learn to ignore them.

### How can I find a site's sitemap?

Try /sitemap.xml or /sitemap_index.xml on the root domain. Check robots.txt for one or more Sitemap lines. If neither exists, the site may not expose a public sitemap or uses a non standard path.

## Read next

- [Sitemap checker](https://porteur.ai/tools/sitemap-checker): Read a sitemap the way a crawler does: what it lists, how it is dated, whether robots.txt names it, and whether thirty of its addresses answer 200.
- [How to use Google Search Console in ten minutes a week](https://porteur.ai/guides/how-to-use-google-search-console): A quick weekly routine: set four filters, compare 28 days, check pages then queries, and fix three findings, without getting lost in noise.
- [Technical SEO checklist for a small site](https://porteur.ai/guides/technical-seo-checklist): Run these 20 technical SEO checks, in order. Each shows how to check it free and what fixed looks like for a site under 1,000 pages.
- [Orphan pages: finding the pages nothing links to](https://porteur.ai/guides/orphan-pages): Find orphan pages on your site, why they happen, how to compare crawl, sitemap and Search Console, and what to do: link, merge, redirect or remove.
- [Indexation](https://porteur.ai/glossary/indexation): Indexation means a page is stored in Google’s index and can appear in search. Here is how to check status, read Page indexing, and fix what blocks it.
- [Crawl budget](https://porteur.ai/glossary/crawl-budget): Crawl budget is how much Googlebot crawls your site. It matters on sites with thousands of URLs. Here is how to check it and avoid wasting it.

Want to see if your sitemap covers the right pages and rivals beat you on key searches, from a URL, in about thirty seconds and free? Run the check. Free check: https://porteur.ai/
