Stop drowning in duplicates: how to deduplicate a literature search properly

Quick answer: search three databases for a literature review and the same paper shows up in each — formatted just differently enough that nothing matches automatically. The clean fix is to export everything, dedup with a purpose-built tool, and keep source tags so your PRISMA counts stay auditable. Here’s the workflow.

Anyone running a serious review knows the dirty secret: comprehensiveness means duplicates, and lots of them. A public-health team recently imported 4,212 records from four databases — and about a third were duplicates. Left unmanaged, they inflate your screening load and distort your counts.

The workflow

  1. Export in RIS from each database separately.
  2. One collection per source in your reference manager (Zotero works well) — source tagging is what makes your PRISMA numbers auditable later.
  3. Run the deduplicator. Zotero’s built-in *Duplicate Items* detector compares titles, DOIs, and ISBNs and is the most accessible option most researchers already have installed. For audit-grade accuracy, ASySD (free, from the University of Edinburgh) tested at 0.95–0.99 sensitivity across datasets up to 80,000 citations.
  4. Merge to the richest record — when you collapse a cluster, keep the master with the fullest metadata.
  5. Record the counts straight into your PRISMA flow diagram.

That public-health team’s 4,212 records collapsed to 2,790 unique in under an hour, source tags intact.

One pro move: maintain a master “Screened” library with decision tags. Each quarter, dedup new results against it — anything flagged as a duplicate is something you’ve already screened. Quarterly review time drops from days to hours.

Full profiles of Zotero, ASySD, and the systematic-review toolkit are in the directory.

Try it: run *Duplicate Items* on your current reference library right now. How many duplicates have been quietly inflating it?

Subscribe for one workflow like this every week.

[Subscribe CTA]

Leave a Comment