// SEO
Duplicate content is identical or substantially similar primary content available at more than one URL, either within one site or across sites. Common examples include tracking and filter parameters, print versions, HTTP and HTTPS variants, alternate hostnames, copied pages, and same-language regional variants. Search systems can group similar URLs and select a representative canonical rather than show every version.
Why it matters: Normal duplication is not automatically a spam violation or a “duplicate-content penalty.” The practical risks are confusing visitors, splitting reporting across URLs, consuming crawl resources on large sites, and having a search engine choose a different canonical from the one you prefer. First identify why each variant exists and whether users need it. Redirect retired or unnecessary variants; use consistent internal links, sitemaps, and `rel="canonical"` signals for versions that must remain; keep canonicals accessible and indexable; and make pages intended to serve distinct search needs meaningfully different. Check Google-selected canonicals in URL Inspection. Treat scraped, scaled, or deceptive duplication separately under relevant content and spam policies.
Explore related checks and guidance for duplicate content on your own site.
Open Duplicate Content AgentLooking for practical context? Start with the guidance behind these checks and definitions.
Read the SEO audit guide