How to fix duplicate content, case by case

There is no duplicate content penalty. There is something worse in practice: Google picks one URL out of a group and ignores the others, and it does not always pick the one you wanted. This page lists the eight places duplicates come from on a real site, and the fix for each.

By , founder of Porteur · Updated 15 September 2026 · Markdown

What Google actually does with duplicates

When Google finds several URLs with the same or nearly the same content, it groups them and chooses one as the canonical. That one can be indexed and ranked. The others are marked in Search Console as an alternate page with a proper canonical tag, or as a duplicate where Google chose a different canonical than you did.

The first of those two is normal and needs nothing. The second is the one to read, because it means Google disagreed with your declaration. The cost of duplicates is not a penalty. It is signals split across addresses, crawl spent on copies, and sometimes the wrong version of a page in the results.

The eight sources of duplicates on a small site

SourceWhat it looks likeThe fix
Protocol and hosthttp and https, www and non-www, all answering 200Pick one, 301 the others on every path, canonical to the chosen host
Trailing slash/pricing and /pricing/ both answerPick one form, redirect the other, keep internal links consistent
Tracking parameters/guides/getting-started?utm_source=newsletterSelf-referencing canonical to the clean URL, never link internally with UTMs
Sorting and filtering/tools?sort=new, /tools?sort=oldCanonical to the unfiltered page, or index the few filters with real demand
Pagination/guides?page=2 holding the same intro as page 1Each page self-canonical, unique title, no canonical to page one
Near-identical pages you wroteTwo guides answering the same questionMerge into the stronger URL and 301 the other
Syndication and republishingYour post on your site and on a partner'sAsk for a canonical to your URL, or accept that theirs may rank
BoilerplateTwenty pages whose only difference is a nameRewrite so each holds something of its own, or cut them

How to find your duplicates in twenty minutes

  1. Try the four addresses of your home page

    Request http and https, with and without www. Three of the four should redirect to the fourth in one hop. If two of them answer 200, that is your first fix.

  2. Read the Page indexing report

    In Search Console, open the Not indexed reasons. Alternate page with proper canonical tag is expected. Duplicate without user-selected canonical and Duplicate, Google chose different canonical than user are the two lists to work through.

  3. Inspect one URL from each list

    URL Inspection shows the canonical you declared and the one Google selected. When they differ, Google is telling you which page it considers the original.

  4. Search your own site for repeated titles

    A crawl, or a site: search on a phrase from a title, shows pages that say the same thing. Two pages with the same title usually answer the same query.

  5. Check your sitemap

    It should list only the canonical version of each page. A sitemap holding both slash and non-slash variants is telling Google you meant both.

Which fix to use, and when

Four tools, and each has one job. Using the wrong one is how a duplicate becomes an invisible page.

  • Redirect with a 301 when the duplicate has no reason to exist. It moves visitors and signals to the page you kept.
  • Canonical when the duplicate must stay reachable for people: a filtered list, a URL with tracking parameters, a print view.
  • Noindex when the page is useful to visitors and worthless in results, such as an internal search results page. Never combine it with a robots.txt block, because a blocked page cannot deliver its noindex.
  • Rewrite when the pages are near-identical because you wrote them that way. A page that only differs by a swapped noun is the case Google's scaled content guidance describes.
  • Leave it alone when Search Console says alternate page with proper canonical tag and the canonical points where you want. That is the system working.

A worked example on yourproduct.com

Search Console lists 140 pages as Not indexed. The reasons break down into three groups, and each group has one fix.

  • Sixty URLs with ?utm_source= from a newsletter, indexed as separate pages. They already carry a canonical to the clean URL, so Google folded them. Nothing to do except stop putting tracking parameters on internal links.
  • Forty URLs under /blog/ with a trailing slash while the site links to the version without. Both answered 200. One redirect rule at the edge and an update to the sitemap fixed the group in a week.
  • Two guides, /guides/getting-started and /guides/quick-start, both written for the same question, trading places on it. The longer one was kept, the better parts of the other folded in, and the second 301ed to the first.

A fixed site looks like this: one address per page, a self-referencing canonical on each, a sitemap that lists only those addresses, and internal links that all point at the same form.

Duplicates across sites, which you control less

Your text can appear elsewhere without any mistake on your side. The options are narrower, and panic is usually the wrong response.

  • Syndication: when a partner republishes your piece, ask for a canonical to your URL or a noindex on theirs. Without one, their version can be the one Google keeps.
  • Manufacturer descriptions: a store using the supplier's text shares it with every other store. The fix is to write your own, which is also the only way those pages earn anything.
  • Scrapers: sites that copy you automatically. Google is generally good at picking the original, and there is a removal request for the rare case it is not. It is not worth your week.
  • Your own second domain: a marketing site and an app site with the same landing pages compete with each other. Pick one home for a page.

What not to do

  • Do not block duplicates in robots.txt. Google cannot then read the canonical or the noindex, and can still index the URL from links.
  • Do not canonical every paginated page to page one. The items on page two then have no indexable home.
  • Do not rewrite a page just to beat a similarity score in a tool. Google groups pages by what they are about, not by a percentage.
  • Do not redirect a duplicate to the home page. That is read as a soft 404 and loses whatever the page had.
  • Do not disavow links to a scraped copy of your page. The disavow file is for a manual action or a known paid history.

Questions

Check my site, free

Paste your URL and the free check reads your site the way a crawler does: the addresses that answer twice, the pages that say the same thing, and the three findings worth your next hour.

  • Free check, no card
  • Read-only, your own accounts
  • Readable by your agent

Read next