# Index bloat

Index bloat is when Google indexes many more URLs than pages worth indexing. It spreads crawl and ranking signals thin. If you run a small site, it wastes your chance to rank the pages that sell or sign up.

Updated 2026-09-14 · Source: https://porteur.ai/glossary/index-bloat

## What index bloat means for a small site

Index bloat is a site having many more URLs indexed than pages worth indexing. Typical sources: tag and author archives, filtered and sorted lists, parameters, on-site search results, paginated duplicates, thin generated pages.

On a small site this dilutes crawl and link signals across junk URLs. It slows updates on key pages like /pricing or /guides/getting-started. It also creates duplicate or near-duplicate entries that confuse which page should rank.

## Where it shows in Search Console

Open Performance. Filter by Search type Web. Compare Total impressions and Total clicks to the number of pages that actually drive your business. Then open Pages. You often see thousands of indexed URLs but only a few hundred that earn clicks.

- Performance, Pages tab: sort by Clicks, then by Impressions to see deadweight URLs.
- Pages report, filter by URL contains “?sort=” or “?page=” or “/tag/” to spot templates.
- URL Inspection: test a sample bloated URL to confirm Indexing allowed and the selected canonical.

If your sitemap lists thousands of tag and filter URLs, that is a sign. Check the Sitemaps report to see what you are feeding Google.

## Find the sources on your site

- Tags and authors: /blog/tag/javascript, /blog/author/jane.
- Filters and sorts: /shoes?colour=red, /list?sort=price.
- Pagination variants: /category/shoes?page=2, /category/shoes?p=2.
- On-site search: /search?q=crm.
- Tracking parameters: ?utm_source=, ?ref=.
- Programmatic or AI stubs: empty location pages, faceted combos with no stock.

1. **Crawl the site** Run a crawler and export all URLs. Group by path and by parameters. A quick pivot by “?” in the URL is revealing.
2. **Map URL to template** For each group, write the template that outputs it: category, tag, search, parameter, pagination, product, article.
3. **Decide intent** Ask one question per template: should this rank? If not, it should not be indexed. If yes, it needs a clean canonical and a place in the sitemap.

> Rule: make one decision per template, not per URL.

## Fix by template: noindex, canonical, robots, or removal

Pick one control per template, then keep the sitemap to only pages meant to rank. Examples below show a fixed page that keeps a single clean URL indexed.

- Tags and authors: add meta robots noindex, follow. Remove them from the sitemap. Example: /blog/tag/javascript now noindex, posts still crawlable.
- Filters and sorts: set a canonical to the unfiltered list. Keep them indexable only if they have unique copy and demand. Example: /shoes?colour=red canonical to /shoes.
- Pagination: page 1 indexable, later pages either noindex or canonical to page 1 if they are near-duplicates. Link to page 2, 3 for users.
- On-site search: disallow crawl in robots.txt if infinite, and add meta noindex to results. Never include /search in the sitemap.
- Tracking parameters: always canonical to the clean URL. Strip UTM parameters server-side where possible.
- Thin stubs: remove, or 410. If needed for users, keep, but add noindex until it has content.

After changes, resubmit the sitemap that lists only target pages. In URL Inspection, Request indexing on a few representative URLs to nudge the recrawl. Watch the Pages and Performance reports over 28 days.

## Monitor and prevent recurrence

- Lock rules in code: template-level meta robots and canonicals.
- Validate on deploy: check for “?” in new internal links.
- Keep one sitemap or a small sitemap index for rankable pages only.
- Review Performance quarterly: pages with zero clicks and zero impressions are candidates for noindex or removal.
- Avoid linking to facets you do not intend to index. Use buttons or forms without crawlable hrefs.

A healthy small site might have 200 pages in the sitemap and, typically, a similar number indexed, not 8,000. /pricing, /features, /guides/… should dominate clicks and impressions.

## Questions

### What is index bloat in plain terms?

It is Google indexing far more URLs than you want to rank. Most are duplicates, parameters, archives, or thin pages. They soak up crawl and scatter signals.

### What are the first symptoms of bloat?

Performance shows many impressions but few clicks spread across odd URLs. Pages shows thousands indexed, yet your sitemap lists far fewer. You also see parameters and tags in site: searches.

### Should I use robots.txt or noindex to fix it?

Use meta robots noindex when the page should be crawlable for links but not indexed. Use robots.txt to prevent crawling only when the content is infinite or pointless to fetch, like /search. Do not block pages you also want to canonicalise, because Google may not fetch the canonical.

### How often should I review indexing?

Quarterly suits most small sites. Also review after launches that add tags, filters or programmatic pages. Watch the Pages and Sitemaps reports after each change.

### Is pagination supposed to be indexed?

Index page 1. Later pages can be indexable if they list unique items users search for. If they repeat page 1 content, keep them crawlable for users but noindex or canonical to page 1.

### Will removing URLs hurt rankings?

Removing junk URLs does not hurt when you keep useful pages crawlable and linked. If a removed URL had links, redirect to the best match. If it has no value and no substitute, return 410.

## Read next

- [Indexation](https://porteur.ai/glossary/indexation): Indexation means a page is stored in Google’s index and can appear in search. Here is how to check status, read Page indexing, and fix what blocks it.
- [Crawl budget](https://porteur.ai/glossary/crawl-budget): Crawl budget is how much Googlebot crawls your site. It matters on sites with thousands of URLs. Here is how to check it and avoid wasting it.
- [The Sitemaps report: what Success, Has errors and Couldn’t fetch mean](https://porteur.ai/guides/search-console-sitemaps-report): How to submit your sitemap URL, read Success, Has errors and Couldn’t fetch, and fix invalid XML, 404s, cross‑host URLs and missing child sitemaps.
- [Crawled, currently not indexed: what it means and what to do](https://porteur.ai/guides/crawled-currently-not-indexed): What “Crawled, currently not indexed” means in Search Console, how to tell why it happened on your site, and what to change that gets pages indexed.
- [Discovered, currently not indexed: why Google has not crawled the page](https://porteur.ai/guides/discovered-currently-not-indexed): Google knows the URL but has not crawled it. Here is how to check why, cut junk URLs, add internal links, fix the sitemap, and set a date.
- [XML sitemap](https://porteur.ai/glossary/xml-sitemap): What to put in sitemap.xml, how lastmod works, how to create, validate and submit your XML sitemap, and the traps to avoid.
- [Google index](https://porteur.ai/glossary/google-index): The Google index is the store of pages Google keeps. See if your pages are indexed, how this differs from ranking, and what to fix in Search Console.

Want a second opinion on bloat templates and sitemaps, fast? Porteur reads your site, the searches around it and the rivals on them from a URL in about thirty seconds, free, and shows three findings whole. Free check: https://porteur.ai/
