Duplicate Content Audit
Crawls your top pages live and finds the title tags and meta descriptions that are identical or near-identical, so your own pages stop competing with each other.
Last updated 2026-08-06
Summary#
Duplicate Content Audit crawls up to twelve of your most important pages and cross-compares every title tag and meta description. It shows which pages are identical, which are near-identical, which have no tag at all, and what proportion of the crawled set is clean.
Purpose#
The duplicate content that hurts most is not scraped copies on other domains. It is your own pages carrying the same title.
When two of your pages say Products | Acme in the title, you have asked Google to choose between them for the same query. It will choose, and it may not choose the one you would. Meanwhile both pages compete, neither ranks as well as one strong page would, and the click-through rate on whichever wins is set by a title that describes neither page precisely.
This tool exists to find those collisions before they cost you anything, on the pages that actually matter — not across a whole site crawl you will never read.
Overview#
Enter a domain. The tool builds a short, high-value page list, then crawls it.
Discovery works in priority order:
- The homepage.
- Your highest-traffic organic pages, ranked by estimated traffic.
- Whatever else is in your sitemap.
The list is filtered to your own domain and its subdomains only, deduplicated so that www and non-www variants of one page are not counted twice, and then capped at 12 pages. The cap is deliberate: the crawl runs live and in parallel, and twelve pages return in seconds.
Comparison then runs over the crawled titles and descriptions:
- Values are normalised before comparison — lowercased, whitespace collapsed, and trailing separators such as a dangling
|or-removed. SoShoes |andshoesare treated as the same title. - Exact groups are two or more pages sharing an identical normalised value.
- Near-duplicate groups are pages that are not identical but share at least 80% of their words, clustered together.
- Missing is any crawled page with no title, or no meta description, at all.
Every page in the report was fetched on this run. Nothing is estimated.
Benefits#
- Finds the collisions that matter. By crawling your top-traffic pages first, the report is about pages that earn something, not about pagination stubs.
- Near-duplicates, not just exact matches.
Blue Running Shoes | AcmeandRunning Shoes Blue | Acmenever match a string comparison, and they are exactly the pair you want to find. - A clean-page percentage that tells you the size of the problem in one number.
- Groups, not a flat list. Each collision shows every page sharing the value, so you fix a set rather than a symptom.
Use Cases#
- Post-migration check. A template change has propagated one title across a whole section, and nobody noticed.
- E-commerce variant sprawl. Color and size variants of one product each get their own indexable URL with the same title.
- Content audit prep. You need a defensible list of pages to rewrite before a quarterly content review.
- Ranking cannibalisation. Two pages keep swapping positions for one query; this shows whether their titles are telling Google they are the same page.
- CMS validation. You confirm that a new template is actually emitting unique titles and descriptions before it ships site-wide.
Requirements#
- A paid plan. The audit costs 1 credit, and the Free plan cannot run any tool that costs credits.
- A publicly crawlable domain. A site that blocks automated crawlers returns an error rather than a partial report.
- Ideally a working sitemap, which widens the page list beyond your top organic pages. Check it with Sitemap Validator.
- No integration and no verified domain.
Permissions#
Available to every team role on a paid plan. On the Free plan the run is refused with:
This tool needs a paid plan. Free includes the 10 technical SEO tools; upgrade to Pro to unlock the rest.
At 1 credit the tool sits below the premium threshold of 3, so on a paid plan it is never blocked by your monthly quota. Each run still adds 1 to your recorded usage.
Cost#
1 credit per run.
Important: The run button reads Run Duplicate Audit · 1 report, and the run charges 1 credit. The missing chip is a display gap, not a free tool.
The audit is not held in the shared data cache, so every run is a live crawl and every run is charged. Runs appear in the Library's Activity timeline as Ran a tool, but they are not stored as reopenable saved work, so there is no free re-open and no You already ran this prompt. Export the result if you need to keep it. See How credits work and Result caching and freshness.
Navigation Path#
Dashboard → Site Health → Duplicate Content Audit
Inputs#
| Field | Accepts | Required | Default | Validation | Notes |
|---|---|---|---|---|---|
Domain field (e.g. nike.com) | A bare domain or a full URL | Yes | Empty | An empty value returns Enter a value first.; https:// is added when the scheme is missing | Only the domain is used. A path is ignored — the tool always builds its own page list |
Run Duplicate Audit | — | — | — | — | Runs the audit. Pressing Enter in the field does the same |
Try sample | — | — | — | — | Fills nike.com when the field is empty, then runs |
Private, loopback and internal addresses are refused before any fetch happens.
Step-by-Step Guide#
- Open Site Health in the left rail, then select Duplicate Content Audit.
- Type your domain in the field marked
e.g. nike.com, or select Try sample to run againstnike.com. - Select Run Duplicate Audit, or press Enter.
- Wait while the panel reads
Analyzing… this can take a few seconds for live data.Twelve pages are fetched in parallel, so this is usually quick. - Read the tier badge in the header, then the four KPI tiles.
- Work through Colliding Title Tags first, then Colliding Meta Descriptions, then Missing Tags.
- Use All Crawled Pages to confirm which pages were actually inspected.
- Export with Export PDF, Export Excel, Export CSV, Export JSON or Share public link at the top of the card.
Reading the Results#
The header#
The domain, a Live crawl verified badge, N top pages compared, and one tier badge summarizing the whole audit:
| Badge | Means |
|---|---|
| ALL UNIQUE | No exact duplicates, no near-duplicates, nothing missing |
| NEEDS REVIEW | No exact duplicates, but near-duplicates or missing tags exist |
| DUPLICATES FOUND | At least one set of pages shares an identical title or description |
DUPLICATES FOUND is the one to act on today. NEEDS REVIEW is a backlog item.
The four KPI tiles#
| Tile | What it shows | Good | Bad | What to do |
|---|---|---|---|---|
| Pages Crawled | How many pages were fetched, sourced from sitemap + top pages | 12 | Fewer than 5 | A low count means discovery found little. Check the sitemap, and whether the site blocks crawlers |
| Duplicate Titles | The number of sets of pages sharing an identical title, with +N near-duplicate underneath when near-matches exist | 0 | 1 or more | Each set is one rewrite job covering several pages |
| Duplicate Descriptions | The same count for meta descriptions | 0 | 1 or more | Lower priority than titles, but it directly costs click-through rate |
| Clean Pages | The percentage of crawled pages involved in no collision and missing nothing | 100% | Under 70% | Below 70% usually points at a template, not at individual pages |
Note that the duplicate counts are sets, not pages. Duplicate Titles: 2 can mean six pages across two collisions.
What we found#
A short list of plain-language findings, one per problem type, with the page counts attached — for example "3 set(s) of pages share an identical title tag (7 pages). Search engines may pick the wrong page to rank, or treat them as competing."
When nothing is wrong the section says so explicitly: "No duplicate or near-duplicate titles or descriptions found across the N pages crawled. Every page has a distinct title and description."
Colliding Title Tags#
One card per collision set, labelled N set(s) in the section header.
- An EXACT card shows the shared title in quotes,
Shared by N pages, and the path of every page carrying it. - A NEAR card is headed
Near-identical across N pages, lists each page's own title so you can see how they differ, and then the paths.
Fix an EXACT set by giving each page a title that names what is unique about it. If the pages are genuinely the same content on different URLs, the fix is not a title rewrite — it is a canonical tag or a redirect.
Fix a NEAR set by asking whether the pages target different queries. If they do, make the titles reflect that difference. If they do not, you probably have two pages where one belongs.
Colliding Meta Descriptions#
The same card structure for meta descriptions. Duplicate descriptions rarely change rankings directly, but Google will often discard a description that does not match the query and write its own snippet, so you lose control of the result text. Treat these as a click-through problem.
Missing Tags#
Shown when any crawled page has no title or no meta description, split into two rows: N page(s) with no title tag and N page(s) with no meta description, each with the paths listed.
A missing title is more serious than a duplicate one. Google will invent something from the page content, usually the first heading, and you have no say in it.
All Crawled Pages#
A table with three columns — Page, Title and Description — and one row per crawled page. Each cell carries a status badge:
| Badge | Means |
|---|---|
| Unique | No collision, tag present |
| Near-dup | In a near-duplicate cluster |
| Duplicate | In an exact-duplicate set |
| Missing | The tag is absent |
Read this table for two reasons. First, it confirms which twelve pages were actually inspected, which matters when the report looks cleaner than you expected. Second, it is the fastest way to hand a list to whoever is doing the rewriting.
Recommendations#
When collisions are found, the findings are handed to the Get recommendations panel below the result, ranked with exact duplicates and missing titles as high-severity, near-duplicates as medium, and missing descriptions as low. Each item can route you to the screen where the fix belongs. The panel is a paid feature; see Get Recommendations.
Examples#
Example: You audit a 400-product store. The header reads DUPLICATES FOUND over 12 top pages compared. Pages Crawled is 12, Duplicate Titles is 1, Duplicate Descriptions is 3, and Clean Pages is 42%. Colliding Title Tags shows one EXACT set: five color variants of one shoe all titled Trail Runner | Acme. Colliding Meta Descriptions shows three sets, all carrying the store's default category boilerplate. Missing Tags lists two pages with no meta description. The read is that this is a template problem, not a content problem: the product template is not interpolating the variant name into either tag. One template change fixes the five titles and, by extension, the same collision across every other product in the catalogue.
Screenshots#
Tips#
- Run it after any template change, not before. The problems this finds are almost always template-shaped.
- Read Clean Pages first. A number under 70% means stop looking at individual pages and go and read the template.
- If Pages Crawled is low, fix the sitemap first and re-run. Discovery quality sets the value of the whole report.
- Export to CSV and hand the All Crawled Pages table straight to whoever writes the new titles.
Best Practices#
- Give every indexable page a title that would make sense on its own in a search result, without the rest of the site around it.
- Put what is unique about the page at the front of the title, and the brand at the end.
- Write a distinct meta description for pages that earn traffic; leave it off entirely rather than repeating boilerplate on pages that do not.
- Where two pages genuinely are the same content, use a canonical tag or a redirect. Rewriting the titles of duplicate pages hides the problem rather than solving it.
- Re-run after the rewrite to confirm the collisions have actually cleared.
Common Mistakes#
- Assuming the run is free despite the
· 1 reportchip. It charges 1 credit. - Reading the duplicate counts as page counts. They are sets. One set can hold many pages.
- Expecting a whole-site crawl. Twelve pages is the cap, chosen so the crawl finishes live.
- Fixing the twelve pages listed and stopping. They are a sample of your most important pages. A collision here almost always repeats across pages that were not crawled.
- Rewriting near-duplicates without asking whether both pages should exist. Sometimes the correct fix is a merge.
Limitations#
- Maximum 12 pages per run. The page list is capped so the parallel crawl completes within the run window.
- Only title tags and meta descriptions are compared. Body content, headings and image alt text are not.
- Near-duplicate detection uses word overlap at an 80% threshold. Two titles that share few words but mean the same thing are not flagged.
- Discovery covers your own domain and subdomains only. Pages that are neither in the sitemap nor among your top organic pages are not seen.
- Results are not cached and are not stored as reopenable saved work, so every run costs a credit.
- The tool does not check canonical tags, so it cannot tell you whether a duplicate set is already correctly canonicalised. Use On-Page SEO Checker on an individual page for that.
Troubleshooting#
| Symptom | Likely cause | Fix |
|---|---|---|
Could not crawl any pages for <domain>. The site may block automated crawlers, or have no indexable pages. | The site refuses automated requests, or discovery found nothing | Confirm the site loads for an anonymous visitor, check robots.txt, and confirm the sitemap resolves with Sitemap Validator |
| Pages Crawled is far below 12 | The sitemap is missing or small, and the domain has few ranking pages | Publish or repair the sitemap and re-run |
This tool needs a paid plan… | You are on the Free plan | Upgrade; see Plans and pricing |
Enter a value first. | The field was empty | Type a domain, or select Try sample |
| The report says ALL UNIQUE but you know there are duplicates | The duplicate pages were not among the twelve crawled | The cap is per run; check the All Crawled Pages list to see what was inspected |
Hourly fair-use limit reached… | More than 100 light-tool calls this hour | Wait for the next hourly reset; see Tool limits and performance |
| Titles look different on screen but are reported as duplicates | Comparison normalises case, whitespace and trailing separators | Expected. A title with a dangling separator and the same title without one are one value to a search engine |
FAQs#
Why only twelve pages? Because the crawl is live. Every page is fetched at the moment you press the button, in parallel, and the cap keeps the run inside its time budget. Twelve of your highest-traffic pages will surface a template-level problem; a thousand pages would mostly repeat it.
Does it find content duplicated on other websites? No. This tool compares your pages against each other. For text copied elsewhere on the web, use the Plagiarism check inside Pre-Publish Trust Checks.
Why does it charge a credit when the button does not say so? The credit chip is missing from the button, which is a display gap. The run costs 1 credit because it uses a paid data provider both to rank your top pages and to crawl them.
Can I re-open a previous audit for free? No. Unlike the premium research tools, this one is not stored as reopenable saved work, so there is no free re-open. Export the result if you need to keep it.
What counts as "near-identical"? Two values that share at least 80% of their words after normalisation. That threshold catches reordered and lightly edited titles while ignoring pairs that merely share a brand suffix.
See also
Was this article helpful?
Thanks — feedback noted for the docs team.