• Technical SEO

Elimination of duplicate pages on the website: when the same product lives at a dozen addresses

The elimination of duplicate pages on the site does not start with edits, but with two numbers: how many entries are in your sitemap and how many of them are really different addresses. The difference between these numbers is the size of the problem — it is then covered by canonical links, redirects, and parameter handling rules. In the garden tool store we measured, these two numbers looked like 33,493 and 8,690. That is, the bot is going around the same thing almost four times, and new items are in the queue. How many such repetitions do you have - we count for free.
See how it works
Price
after a free audit
Guarantee
30 days after project sign-off
The worst case in the sample
33,493 entries in the site map for 8,690 unique addresses
after a free audit
Price
30
Guarantee

days after project sign-off

33,493
The worst case in the sample

entries in the site map for 8,690 unique addresses

271
The smallest catalog with takes

records for 116 addresses with 81 product pages

13,590
Standard of purity in the sample

addresses and exactly one duplicate

07
Date of measurements

/31/2026, external technical measurement

7-20
Term of works

working days

2-3
Free calculation

working days, the document remains with you

increase in traffic and the number of pages in the index
What we do not promise
Who it's for

Situations where this service delivers results

Scenario 1 of 4

The catalog is updated, and new products are not visible in the search for a long time

The bot comes with a limited crawl resource. If each card lies in the site map four times, most of this resource goes to rescanning the already known. The measured garden tool store had 33,493 listings for 8,690 real-world addresses—a new item in that queue is waiting longer than it should, and no amount of content budget is going to speed it up.

We'll review your situation in a free audit
Cases

Tasks and results in numbers — all metrics measured by us

official online store of garden tools, catalog of about 8,600 cards

Task
Calculate the extent of duplication before cleaning the index: before the works, it was only known that the sitemap weighed too much.
Solution
A full review of sitemap.xml with separate counting of records and unique addresses, layout of unique ones into category nodes and leaf pages, checking of language branches in the map.
Result
8.57MB map: 33,493 entries for 8,690 unique addresses — 3.85 entries per page. Out of 8,690 unique addresses, 81 turned out to be category nodes, and 8,609 were leaf pages. Separately, it turned out that only the Ukrainian-speaking branch was included in the map, although two were announced in the marking. Measured on 07/31/2026.

online store of the official seller of garden tools, catalog of about 1.2 thousand cards

Task
Find out why the site map weighs more than a megabyte with a catalog of a thousand and small items.
Solution
Traversing the map with a count of unique addresses and a breakdown by nesting levels; the number of cards was checked against the manufacturer's article number, which is in the address itself.
Result
1.18 MB and 4,236 entries for 1,300 unique addresses — 3.3 entries per address: the item is mapped under multiple category paths. The distribution of these 1,300: 1,231 first-level addresses, 64 subcategories, and 5 service paths from index.php, which the driver mapped himself. Out of 1,231 addresses, 1,120 contain the manufacturer's article number - that's why we were able to calculate the catalog. Measured on 07/31/2026.

online store of natural cosmetics on beekeeping products, a small catalog

Task
Check the indexing status after measuring the speed: the site was the fastest of the lot, and it was necessary to understand whether it was just as clean with addresses.
Solution
Counting of records and unique addresses in the site map, reconciliation with the actual number of product pages, inventory of analytics counters.
Result
271 records for 116 unique addresses: individual goods are duplicated two to four times. Out of 116 addresses, 91 are letter pages, 10 of them are informational, i.e. 81 are commercial. The first server response of 126.3 ms is the best indicator of the lot. Separately, it was found that the site counts with a dead counter, and there is no GA4 identifier in the markup: the consequences of cleaning here simply would be nothing to measure. Measured on 07/31/2026.
What's included

Complete list of work and what you get as a result

  • Full sitemap crawl with separate count of records and unique addresses — you get the size of the problem in numbers, not by eye evaluation
  • Classification of repetitions by mechanics: category paths, filter parameters, language prefixes, pagination, session IDs - each type is further treated by its own tool
  • Choose one primary address for each product with you - the business needs the page to promote, not the one that the search engine chooses on its own
  • Canonical links on every type of page, self-referencing where appropriate: master version signal is no longer random
  • 301-redirects from old product paths to the main address - redundant addresses cease to exist, not just stop being recommended
  • Rules for handling filter and sort parameters so that a click in the directory does not generate a new address for the index
  • Separate work with language branches: hreflang where pages are not duplicates, but variants for different languages
  • Regeneration of the site map without repetitions and repeated traversal with the same method - the difference can be seen in the two numbers before and after
  • Updating robots.txt and indexing directives for the new address structure
  • Observations in Search Console 30 days after the cleanup explaining why the number of pages initially drops
When this service isn't right

What's not included — so there are no surprises at delivery

  • Writing unique texts for pages that remain after cleaning
  • Dealing with external sites copying your content
  • Moving to another CMS
  • Promises about the number of pages in the index - this is at the disposal of the search engine
Process steps

Transparent stages with approval at every step

Total duration:7–20 days

  1. Counting and classification

    2-3 working days

    Site map walkthrough, two numbers, duplication factor. Next, there is a breakdown of repetitions by mechanics and a list of addresses that are repeated most often. The output is a document that shows exactly what needs to be treated.

  2. Decisions on the main addresses

    1-3 working days

    We analyze together with you which path is considered the main one for each section. This is not a technical issue: it depends on the category in which the product will be promoted further. We fix it in writing, so as not to redo it later.

  3. Canonical links and redirects

    2-6 working days

    Self-referencing canonical on all page types, 301 from old product paths to main address, separate hreflang for language versions. Each change goes out in batches with a check so as not to collect a chain of redirects.

  4. Options, pagination, site map

    1-4 working days

    Rules for filters and sorting, processing of pagination pages, regeneration of the site map without repetitions, update of robots.txt to the new structure.

  5. Re-circulation and observation

    1-4 working days

    The same counting method, the same two numbers — a before-and-after comparison. Next, we look at the indexing report in Search Console for 30 days and explain the dynamics, including the expected drop in the number of addresses.

Free double counting in your sitemap

Dubs are not visible from the administrative panel - they are visible from the search engine. We take your sitemap, count the records and unique addresses separately, and the difference between the two numbers tells you the size of the problem. This calculation costs nothing.

What we measure

  • Entries against unique addressesHow many lines are in your sitemap and how many of them are actually different pages. Measured sampling benchmarks range from one duplicate per 13,590 addresses to 3.85 records per address.
  • Mechanics of generationWhat causes repetitions in you: the path of the category in the address, filter parameters, language prefixes, pagination, session identifiers.
  • The worst repeatersSpecific addresses that are found on the map most often. Usually, these are several dozen cards, arranged in the maximum number of categories.
  • Status canonicalIs there a canonical link on each type of page, is it self-referencing or leads somewhere else, and is it consistent with what is advertised in the sitemap.
  • Dubs between language versionsAre both versions with the same meta tags included in the index, and is it not the other way around? In the measured garden tool store, only one language branch out of the two announced entered the map.
  • Product in several categoriesHow many cards are available from more than one path. In a measured directory of 1,300 addresses, this resulted in 4,236 entries in the map.

What you get

  • Two numbers: entries in the site map and unique addresses, plus the duplication factor.
  • Analysis of where your repeats come from - with examples of specific addresses.
  • Cleaning plan: what is done by CMS settings, what by redirects, what by canonical links.
  • Conversation for 30 minutes on the document.

Timeline: 2-3 working days

Why is it free

Because the calculation is automated and takes us an hour or two. And without it, it is impossible to say whether it is work for three days or three weeks: the mechanics of duplication in each CMS is its own.

What's next

After the calculation, we name the list of works and the deadline. We warn right away: after cleaning, the number of pages in the index first drops - this is the expected course of events, and it is better to know about it before the start, and not after.

Short form: your contact and site URL

  • Contract, act and 30-day warranty

    Every project gets a written contract: scope, deadlines, amount, acceptance procedure. After delivery — act and invoice, then 30 calendar days of warranty.

  • Sole proprietor & bank transfer

    The contractor is a registered sole proprietor. Payment by invoice with closing documents.

  • Rights & access — yours

    Code, design and materials transfer to you after full payment. Domain, hosting, repository and analytics are registered to you.

  • Client portal instead of email chains

    During the project you get access to a portal: contracts, invoices, acts and project status in one place.

  • European clients

    Among our work — projects for Norway, Bulgaria, Moldova and Spain.

  • Verifiable numbers

    Every case in the portfolio comes with a link to a live site and a technical measurement.

  • Audit first, then pricing

    There is no price list on the site intentionally: the scope of the same work differs multiples between clients.

  • We say "no" when unsure

    If the task isn't ours or the deadline is unrealistic — we tell you upfront.

Did not find your case?

Describe how it works on your side — we will tell you whether “Fix duplicate URLs” fits and what it means in your situation. No brief and no call: one question, one answer.

What affects the price

Why two seemingly identical tasks are priced differently

  • How many takes actuallyOne duplicate per 13,590 addresses and 3.85 records per address is several weeks of work. We have seen the first number in a wholesale underwear store, the second in a garden tool store, and between them fits the entire volume fork.
  • How many mechanics work at the same timeOne site breeds addresses only by way of categories. The other is by paths, filter parameters, language prefixes and pagination together. Each mechanism is closed separately and checked separately by re-circulation.
  • Does the CMS know how not to generate a redundant addressWhen the driver knows how to deliver the goods to one address, the issue is closed with an hourly adjustment. When it does not know how, you have to write rules at the server level and then maintain them with each update.
  • The status of canonical links nowEmpty - we put from scratch according to one rule, it is predictable. The worst case is when there is a canonical, but it leads to the wrong place: first you need to understand the logic behind it, and only then rewrite it.
  • The number of language branchesEach language version doubles the number of checks and adds hreflang handling. Doubles between languages ​​are treated differently than doubles within the same language, so it's not the same volume multiplied by two.
  • Who decides what the main address isThe technical part is rarely a bottleneck. The longest is the answer of the business in which section the product is considered the main one - and if it is given for weeks, the calendar stretches without a single line of code.
Technologies & integrations

What we build on and what it connects to

Stack

  • Traversal of sitemap.xml with counting of unique addresses - shows only what you yourself have declared; he does not see doubles outside the map
  • Screaming Frog - crawling through internal links. It goes on for hours on tens of thousands of addresses and rests in the memory
  • Search Console is the only place where you can see which duplicates are actually in the index. Delayed and sampled data
  • rel=canonical is a hint, not a command: the search engine has the right to ignore it
  • 301-redirect — the decision is final. After a few cleanups, redirect chains become a separate issue
  • URL parameter handling rules - remove filters and sorting from the index, but not where the parameter fundamentally changes the content
  • nginx / .htaccess — the rule works outside the CMS; the site puts the error here, so the changes leave separately and with a rollback
  • OpenCart SEO URL, WordPress permalinks — the cleanest, when an extra address is simply not generated; the driver is not able to do it everywhere

Integrations

  • Google Search Console
  • Bing Webmaster Tools
  • GA4
  • Cloudflare
Cleaning by count vs. closing doubles blindly

How this option differs from the alternative

Where does the work begin?two numbers: entries in the sitemap and unique addresses among them
What is done with the extra address?one main, the rest - canonical or 301 depending on the type
Where does the page weight go?is gathered at one main address
Will reruns return?Spawning mechanics are closed in the engine and on the server
What is visible in Search Consolethe drop in the number of addresses, explained before the start of work
What we need from you

We can't start without this — best to prepare in advance

  1. The address of the site map — counting starts from it. If there is no map at all, this is the first line of work, and not a reason to refuse.
  2. Access to Search Console. Without it, duplicates are visible in the map, but it is not visible which of them actually got into the index.
  3. Access to the CMS admin — part of the rules are set in the engine settings, not in a file on the server.
  4. Access to the server configuration or the contact of the developer who has the right to change it.
  5. The decision, which address of the product is considered the main one, when it lies in several categories. This is a business decision: it determines in which section the product is promoted further.
  6. One person on your side who can answer such questions in a day, not two weeks.

If something is missing — let us know, we'll help you gather it or do it as a separate task.

FAQ

Most frequently asked questions — with concrete answers

How do duplicates really hurt - I don't lose pages, do I?

You don't lose pages, you lose two other things. The first one is the bypass resource: the bot comes with a limited limit and spends it rescanning the same thing. In the garden tool store measured, there were 33,493 entries in the map for 8,690 real addresses, that is, most of the walk was idle, and new goods were waiting for their turn. The second is weight: links and behavioral signals are split between copies, and instead of one strong page, you have three weak ones competing for the same query. In such a situation, the search engine itself chooses what to show, and not always the address you promoted.

How to understand if I have duplicates without hiring anyone?

Open your sitemap and count two numbers separately: how many entries there are in total and how many different addresses there are among them. If the first is noticeably more than the second, there are duplicates, and you have just calculated the duplication ratio. Landmarks from our measurements: 13,590 addresses and exactly one duplicate in a wholesale underwear store — that's clean; 4,236 records for 1,300 addresses in a garden tool store is already a mechanic that works every day. We do the same calculation for free and add to it an analysis of where the repetitions come from.

Why does the number of pages in the index drop after cleaning?

Because you remove the extra addresses, and they really disappear: instead of ten addresses of one product, only one remains. This is an expected course of events, not a glitch. We talk about it before starting work precisely because the graph in Search Console shows the same drop both when everything is done correctly and when the site is hacked. The difference is that after cleaning, addresses are reduced, not goods, and this is visible in order - we show exactly which addresses have disappeared.

Is Canonical enough, or are redirects necessary?

It depends on whether the redundant addresses should remain accessible to the person. Canonical - hint: the search engine has the right to ignore it, especially if the content of the pages is different. 301 is a decision after which the address simply does not exist. In practice, filters and sorting are closed with a canonical link, and old product paths are closed with a redirect. And a separate tip: don't close duplicates via robots.txt. Prohibition of traversal does not remove the address from the index, it only deprives the bot of the opportunity to see your canonical - the duplicate remains forever after that.

And if the product really falls into three categories, is it convenient for the buyer?

It is convenient, and there is no need to break this convenience. It is necessary to break the fact that each path gives a separate address. There are two options: either the paths do not generate separate URLs at all — this is a driver setting and the cleanest way — or each additional path sends a canonical to the main address. Which of the addresses is the main one is decided by the business: it depends on which section the goods are promoted. We show the consequences of both options, the choice is yours.

How long does it last and what is extending the deadline?

7-20 working days. Counting and classification 2-3 days, decisions on main addresses 1-3, canonicals and redirects 2-6, parameters and sitemap 1-4, re-traversal and observation 1-4. The technical part is rarely a bottleneck. The answer to the question of which address to consider as the main one for each section is longer: if it is given in a day, we fit in the lower limit, if a meeting is being held - in the upper limit.

Is it possible to just disable the sitemap so that duplicates don't get there?

It is a mirror treatment instead of a face. Duplicates exist not because they are on the map, but because the site delivers the same product to different addresses — the bot will reach them through internal links and without a map. A map only makes the problem visible, and that's all it's good at this point. By the way, it also shows what the CMS considers a page: in one of the measured stores, the driver himself put five service paths with index.php on the map next to 1,231 addresses of the first level.

Will traffic increase after cleaning?

There will be no such promise. The mechanism works differently: cleaning does not add new visitors, it stops spraying the existing ones and frees the crawling resource for pages that the search engine has not yet seen. Google describes canonicalization precisely as a way to indicate the main version among duplicates and separately warns that otherwise it will choose it on its own - the forecast of traffic growth does not follow from this documentation. On a small site, the effect is generally close to zero, and we are talking about this before the works, not after.

Submit your sitemap address and we'll count the duplicates and show you where they're coming from.

In response: two numbers, a duplication factor, a breakdown of the mechanics with examples of specific addresses, and a cleaning plan. If you have almost no repetitions, you will hear it in the first email.

From measured cases8.57MB map: 33,493 entries for 8,690 unique addresses

View cases
  • Reply within 2 hours
  • No commitment
  • We work under a contract

There is no price list on the site on purpose: the same work differs several times over between two clients, and a “from” figure explains nothing in that case. First a free audit — we count your pages, duplicates and speed — then we name the sum and the deadline and fix both in the contract.