<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>errors caused by automated deduplication processes &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/errors-caused-by-automated-deduplication-processes/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 21:32:31 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>errors caused by automated deduplication processes &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Deduplication Software Can Silently Distort Systematic Reviews, Paediatric Neuroradiology Study Warns</title>
		<link>https://scienmag.com/deduplication-software-can-silently-distort-systematic-reviews-paediatric-neuroradiology-study-warns/</link>
		
		<dc:creator><![CDATA[Harold Sullivan]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 21:32:31 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[accuracy of bibliographic record merging tools]]></category>
		<category><![CDATA[best practices for managing duplicates in research]]></category>
		<category><![CDATA[bibliographic records]]></category>
		<category><![CDATA[challenges in pediatric radiology literature synthesis]]></category>
		<category><![CDATA[deduplication]]></category>
		<category><![CDATA[deduplication software in systematic reviews]]></category>
		<category><![CDATA[DOI]]></category>
		<category><![CDATA[errors caused by automated deduplication processes]]></category>
		<category><![CDATA[evaluation of software tools for literature deduplication]]></category>
		<category><![CDATA[evidence synthesis]]></category>
		<category><![CDATA[impact of database merging errors on research validity]]></category>
		<category><![CDATA[implications of hidden duplicates in systematic review outcomes]]></category>
		<category><![CDATA[neuroradiology]]></category>
		<category><![CDATA[paediatric radiology]]></category>
		<category><![CDATA[pediatric neuroradiology systematic review challenges]]></category>
		<category><![CDATA[PRISMA]]></category>
		<category><![CDATA[PRISMA flow diagram limitations]]></category>
		<category><![CDATA[Rayyan]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[research integrity and data management in]]></category>
		<category><![CDATA[risks of silent data deletion in literature reviews]]></category>
		<category><![CDATA[subacute sclerosing panencephalitis]]></category>
		<category><![CDATA[systematic review]]></category>
		<category><![CDATA[Zotero]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=207871</guid>

					<description><![CDATA[A study in Pediatric Radiology finds that widely used deduplication tools introduce both deterministic errors and non-reproducible record loss in systematic reviews, threatening the integrity of evidence synthesis before screening even begins.]]></description>
										<content:encoded><![CDATA[<p>Deduplication is the unglamorous housekeeping step of every systematic review. Before a single abstract is screened, reviewers must merge the overlapping records that PubMed, Embase and other databases return for the same paper. It is a task so routine that most reporting guidelines barely mention it, and the Preferred Reporting Items for Systematic reviews and Meta-Analyses, or PRISMA, flow diagram records only the tidy end result: how many records were found, how many were removed as duplicates, and how many moved on to screening. What the diagram never shows is whether those removals were correct. A new study published in Pediatric Radiology argues that this blind spot is far more dangerous than the research community has assumed, because the software tools trusted to do the merging can quietly delete unique studies or keep hidden duplicates, and neither error is visible anywhere in the final published review.</p>
<p>The study, led by Lakshmi Vennela Chowdary Kaza of the Indian Institute of Technology Bombay with colleagues from Myro Health, Basildon Hospital and the All India Institute of Medical Sciences Raipur, set out to measure exactly how wrong popular deduplication tools can be. The team assembled a corpus of 603 bibliographic records drawn from PubMed and Embase on a single, well-defined clinical question: neuroimaging in paediatric subacute sclerosing panencephalitis, a rare and devastating progressive brain disorder caused by a persistent measles virus infection. Because the corpus was modest in size and the clinical literature narrowly focused, the researchers could afford something most review teams never do: a complete, record-by-record manual reference standard against which every automated decision could be judged.</p>
<p>Building that reference standard was the methodological heart of the experiment. Two reviewers independently ran the same combined corpus through two of the most widely used tools in academic workflows, Zotero version 8.0.4, the open-source reference manager maintained by the Corporation for Digital Scholarship, and Rayyan, a web-based screening platform updated as recently as 21 January 2026. Whenever either tool flagged a pair of records as potential duplicates, the reviewers adjudicated the decision manually, comparing titles, publication years, author lists and, crucially, Digital Object Identifiers. Through this dual-reviewer process they established that the corpus contained exactly 120 true duplicate records, leaving 483 genuinely unique studies. That adjudicated set became the ground truth for scoring every subsequent automated output.</p>
<p>The results split into two distinct failure modes, each with its own implications. Zotero behaved with perfect determinism: both reviewers, running the tool independently on identical data, received identical outputs every time. Determinism is usually a virtue in software, and reproducibility is a cornerstone of systematic review methodology, so at first glance Zotero looks like the safer choice. But its accuracy was far from perfect. Against the manual reference standard, Zotero produced 16 false-positive duplicate decisions, wrongly merging 16 pairs of records that were in fact distinct studies, and 5 false negatives, failing to detect 5 true duplicates. The errors were not random: the authors traced the majority to DOI collisions, cases in which the same DOI string appeared on records that were not the same paper, or in which DOI-based matching overrode other evidence that the records differed.</p>
<p>That failure mode matters because DOI-based matching is widely regarded as the gold standard for identifying identical publications. A DOI is meant to be a unique, persistent identifier, and bibliometric databases have long been known to suffer from indexing errors in which identifiers are misassigned or duplicated. The new study provides a concrete, quantified demonstration of how those upstream indexing problems cascade into evidence synthesis: when a reference manager treats a shared or corrupted DOI as conclusive proof of duplication, two different papers can vanish from a review without any trace in the PRISMA flow diagram. In a small, specialised field such as paediatric neuroradiology, where the total evidence base for a rare disease may comprise only a few hundred studies, losing even a handful of unique records can materially change the conclusions of a meta-analysis.</p>
<p>Rayyan failed in the opposite and, in some ways, more unsettling way. The web platform detected more of the true duplicates than Zotero, committing only a single false negative per run, but it generated 16 to 17 false positives and, more importantly, did not produce identical outputs across independent runs. Two reviewers performing the same task on the same corpus at the same time received different deduplication results. The study does not identify the precise internal mechanism behind this non-determinism, but the consequences are easy to state: the record set entering screening depends on when and how the tool was run, on which server processed the request, and on the state of the underlying software at that moment. A review team could therefore publish a PRISMA diagram that no independent team, or even the same team a week later, could reproduce.</p>
<p>The reproducibility problem strikes at a foundational expectation of evidence synthesis. Systematic reviews derive their authority from transparency and auditability: any competent team following the published methods should be able to arrive at the same set of included studies. If the deduplication stage is non-deterministic, that chain of reproducibility is broken before screening even begins, and no amount of careful documentation downstream can restore it. The authors note that errors introduced at this preprocessing stage are invisible in standard reporting frameworks, which means the research community currently has no systematic way to detect them in published reviews. A review could be silently biased, either by missing studies or by double-counting records, and every conventional quality-assurance check would pass.</p>
<p>The scale of the errors in this corpus gives a sense of the stakes. Sixteen false positives and five false negatives out of 603 records may sound like a rounding error, but it represents roughly 3.5 percent of the unique record set being misclassified by a single tool in a single pass. Across the thousands of systematic reviews published each year in medicine, many of which inform clinical guidelines and treatment decisions for children, the cumulative effect of similar error rates could be substantial. The problem is compounded by the fact that deduplication is almost never validated in individual reviews. Teams routinely pilot their screening decisions and report inter-rater agreement for inclusion judgments, yet almost no one measures whether their deduplication software deleted the right records, and journals rarely ask.</p>
<p>The authors&#8217; recommendations are pragmatic rather than dramatic. They call for the deduplication software name and version to be explicitly reported in every systematic review, alongside the verification strategy used to check its decisions, so that readers can assess the reliability of the record set. They also recommend auditable secondary checks before any records are removed: for example, having a human reviewer adjudicate every flagged pair, or running a second, independent tool and reconciling the differences. The study&#8217;s supplementary materials include record-level error logs and the full adjudicated reference corpus, allowing other teams to replicate the evaluation on their own data. In an era when artificial intelligence tools are increasingly being layered onto every stage of the review pipeline, from search query generation to screening triage, the authors&#8217; findings serve as a reminder that even the oldest, simplest automation steps deserve the same scrutiny as the newest ones.</p>
<p>For the paediatric radiology community specifically, the study arrives at a moment of growing attention to the quality and equity of the evidence base. Subacute sclerosing panencephalitis remains a disease of profound public health importance in regions where measles vaccination coverage is incomplete, and the neuroimaging literature that guides its diagnosis is scattered, small and vulnerable to loss. If the studies that survive into a review depend on the quirks of a reference manager&#8217;s DOI parser or the non-deterministic behaviour of a web platform, then the synthesis built on top of them rests on sand. The message of this evaluation is deceptively simple: deduplication is not bookkeeping, it is a measurement step, and like any measurement it needs a validated instrument, a documented version, and a human check before its output is trusted.</p>
<p><strong>Subject of Research:</strong> Evaluation of the validity and reproducibility of deduplication tools used in systematic reviews on a paediatric neuroradiology corpus</p>
<p><strong>Article Title:</strong> Reproducibility and validity failures in systematic review deduplication tools: evaluation on a paediatric neuroradiology corpus</p>
<p><strong>Article References:</strong> Reproducibility and validity failures in systematic review deduplication tools: evaluation on a paediatric neuroradiology corpus. (n.d.). <a href="https://doi.org/10.1007/s00247-026-06775-z" rel="noopener noreferrer">https://doi.org/10.1007/s00247-026-06775-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00247-026-06775-z" rel="noopener noreferrer">10.1007/s00247-026-06775-z</a></p>
<p><strong>Keywords:</strong> systematic review, deduplication, reproducibility, Rayyan, Zotero, paediatric radiology, neuroradiology, evidence synthesis, PRISMA, DOI, subacute sclerosing panencephalitis, bibliographic records</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">207871</post-id>	</item>
	</channel>
</rss>
