Literature search · Record management
Deduplicate literature search results without losing provenance
A repeated database record should not become a second screening task. A related paper should not disappear just because its title looks familiar.

In this article
Combining two literature searches often gives you several descriptions of the same paper. Removing those repeated records is straightforward until one export has no DOI, another abbreviates the title, and a third contains a companion paper. The safe output is a smaller screening set plus a record of every merge, not a list whose missing rows cannot be explained.
This workflow uses a spreadsheet and, optionally, Zotero. It produces a retained-record list and a merge ledger. It does not determine whether a paper is eligible for your review or whether several papers describe the same underlying study.
1. Preserve the inputs before editing
Save each original RIS, BibTeX, or CSV export unchanged. Record the database, search date, query version, filename, and imported record count in your literature search log. Assign a stable local ID to every imported row, including repeated rows. A database accession number can be an additional identifier; it should not replace your local ID when several databases are involved.
Create a working sheet with these columns: record_id, source_run, source_record_id, title_raw, doi_raw, doi_key, retained_id, decision, and reason. Keep the raw values beside any normalized values. Store the merge ledger separately from the screening decisions so a duplicate is never accidentally recorded as an excluded study.
2. Separate records, reports, and studies
A record is a search result describing a document. A report is the document itself. A study is the underlying investigation, which may have several reports. The Cochrane Handbook, Chapter 5 requires reviewers to link multiple reports of a study rather than treat each report as an independent study.
Use two different operations: deduplicate records that describe the same report, then link distinct reports when evidence indicates they concern the same study. A conference abstract, full journal article, protocol, and follow-up analysis may need to remain separately accessible. Shared authors or participants are reasons to investigate a relationship, not reasons to discard a document.
3. Generate candidates, then verify the match
For a DOI comparison key, remove surrounding whitespace and an explicitly recognized DOI resolver prefix such as https://doi.org/. Preserve the original field. Do not indiscriminately remove punctuation or truncate the suffix: an apparently untidy character may belong to the identifier. If the export contains a full URL with extra parameters or an ambiguous identifier, inspect it before normalizing.
Group exact keys first, then inspect the title, authors, year, and source page. Matching identifiers are strong candidates, but incorrectly assigned metadata can still produce a false match. When a DOI is missing, compare the full title and bibliographic details manually. A fuzzy title score can help order a review queue; it is not a deletion decision.
- Confirmed duplicate: retain one record, map every merged record ID to it, and preserve all source-run memberships.
- Related but distinct report: keep both records and describe the relationship separately.
- Uncertain: keep both pending review. Record what evidence is missing instead of guessing.
4. Work through a real DOI pair
The following three rows are an illustrative import fixture, not results from a database search. They use two real publications: the PRISMA 2020 statement paper, DOI 10.1136/bmj.n71, and its explanation and elaboration paper, DOI 10.1136/bmj.n160. The first paper appears twice to demonstrate a resolver-prefix difference. Titles in the visual are shortened for readability.
10.1136/bmj.n71https://doi.org/10.1136/bmj.n7110.1136/bmj.n160After normalization, A-01 and B-07 have the same DOI key. Verify that both describe the statement paper, retain A-01, and write B-07 → A-01 into the ledger with the reason “same DOI and verified publication.” A-01 now carries both source memberships; B-07 remains traceable in the unchanged export.
B-08 stays separate. Its DOI and full title identify a different paper, even though both publications concern PRISMA 2020. The PRISMA authors explicitly recommend reading the statement together with the explanation paper. Collapsing them would remove useful guidance. This example reconciles to three imported records minus one duplicate record equals two retained reports. It makes no claim about the number of studies in a review.
5. If using Zotero, merge rather than delete
In Zotero, open Duplicate Items within the relevant library and inspect each proposed group. Zotero detects candidates using fields including title, DOI, and ISBN; detection is limited to that library. Choose the master record, inspect conflicting fields, and use Merge Items once the records are confirmed duplicates. Zotero's duplicate-detection documentation explains that merging preserves collections and tags and is recognized by its word processor plugins.
Record the local ID mapping before the merge. The application's merged item is useful for reading and citation, while your external ledger explains which export rows contributed to it. Do not assume the library interface is a substitute for your search-run accounting. Before importing a cleaned library into another workspace, inspect its exported metadata and the retained-record count.
6. Reconcile before screening
For a batch with no other removals, check this identity: imported records = retained records + duplicate records merged. Count removed records, not duplicate groups: a group of four contributes three merged duplicates. Keep format errors and other pre-screening removals in separate categories if they occur. Do not force them into the duplicate count to make the arithmetic work.
Test the ledger in both directions. Every merged ID should resolve to exactly one retained ID, and every retained ID should lead back to its original source rows. Investigate broken mappings, a retained record pointing to another removed record, or a source run that disappears after a merge. Then export the screening set and retain a dated copy of the ledger with it.
Check likely false negatives too: records without identifiers, abbreviated titles, and publication-version pairs. Keep preprints, corrections, and companion reports available until their relationship has been reviewed. Deduplication cannot establish study independence, assess retractions, or decide which version supports a claim.
Once the records are stable, carry the retained IDs into your evidence matrix. That connects a screening decision to a source record and, later, a passage. The benefit is a review trail you can reconstruct when a colleague asks why a paper was retained or where another record went.