<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>bibliographic data management and deduplication &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/bibliographic-data-management-and-deduplication/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 02:36:23 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>bibliographic data management and deduplication &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Open-Source R Workflow Aims to Make Systematic Literature Reviews Reproducible</title>
		<link>https://scienmag.com/new-open-source-r-workflow-aims-to-make-systematic-literature-reviews-reproducible/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 02:36:23 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automation of evidence extraction]]></category>
		<category><![CDATA[bibliographic data management and deduplication]]></category>
		<category><![CDATA[bibliographic data visualization and mapping]]></category>
		<category><![CDATA[bibliometrics]]></category>
		<category><![CDATA[bridging quantitative bibliometrics and qualitative synthesis]]></category>
		<category><![CDATA[integrated bibliometric analysis tools]]></category>
		<category><![CDATA[methodological transparency in literature reviews]]></category>
		<category><![CDATA[modular software for literature screening]]></category>
		<category><![CDATA[open-source R workflow for literature reviews]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[PRISMA]]></category>
		<category><![CDATA[R programming]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[reproducible research in systematic reviews]]></category>
		<category><![CDATA[research synthesis]]></category>
		<category><![CDATA[SLR Suite]]></category>
		<category><![CDATA[SoftwareX]]></category>
		<category><![CDATA[SPAR]]></category>
		<category><![CDATA[systematic literature review]]></category>
		<category><![CDATA[systematic literature review reproducibility]]></category>
		<category><![CDATA[TCCM framework]]></category>
		<category><![CDATA[text classification]]></category>
		<category><![CDATA[theory-oriented evidence synthesis]]></category>
		<category><![CDATA[version control in research workflows]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=260906</guid>

					<description><![CDATA[A new open-source R workflow called SLR Suite integrates bibliometric analysis, dictionary-based TCCM coding, and reproducibility controls into a single version-controlled pipeline for systematic literature reviews.]]></description>
										<content:encoded><![CDATA[<p>Systematic literature reviews are among the most labor-intensive exercises in modern science. Researchers must collect thousands of bibliographic records, deduplicate them, screen them against eligibility criteria, run bibliometric analyses, and then somehow convert citation networks and keyword maps into theory-oriented evidence. In practice, this usually means juggling a patchwork of separate tools, each with its own data formats and settings, and the methodological trail between them often breaks down. A newly published open-source project called SLR Suite, described in the journal SoftwareX by developer Metin Akbulut, proposes a different approach: a single, modular R workflow that carries a review from raw bibliographic data all the way to structured, theory-oriented evidence tables, with every step logged, versioned, and repeatable.</p>
<p>The core problem the software targets is reproducibility. When researchers rely on separate applications for data collection, screening, bibliometric analysis, and visualization, transferring data and methodological decisions between them makes analytical provenance and configuration management difficult to maintain. A second, subtler gap concerns the divide between quantitative bibliometrics and theory-oriented synthesis. Frameworks such as Theory–Context–Characteristics–Methodology (TCCM) and Structure–Process–Analysis–Research Agenda (SPAR) offer structured ways to organize evidence, but applying them has typically been a manual exercise conducted far from the computational pipeline. SLR Suite attempts to bridge that divide by creating a traceable connection between selected bibliometric outputs and configurable, inspectable TCCM evidence for researcher-led synthesis.</p>
<p>Architecturally, the suite is organized as a pipeline of nine ordered modules: bibliographic acquisition and deduplication, rule-based screening support, bibliometric analysis, VOSviewer export, dictionary-based TCCM matrix construction, thematic evolution analysis, citation-impact assessment, SPAR-oriented evidence aggregation, and generation of PRISMA flow-count artifacts. A shared launcher orchestrates execution but contains no analytical logic itself; each module reads its own declared configuration and input files and writes standardized outputs to designated data layers. Modules exchange data through documented file contracts rather than direct function calls, a design choice that reduces coupling and allows any individual module to be rerun independently when its upstream inputs are available.</p>
<p>The data flow is organized into three structured layers—raw, interim, and processed—so that every transformation is inspectable and intermediate datasets are preserved for auditing or debugging. Configuration lives in version-controlled YAML files: screening criteria in one file, and researcher-defined TCCM dictionaries of labels and case-insensitive regular-expression patterns in another. Because the classification procedure is rule-based rather than inferential, researchers can inspect exactly which patterns produced each label and revise the dictionaries for their particular domain. The direct input format is a Web of Science plain-text export; Scopus CSV records are handled upstream by a companion tool called BibexPy, which merges and harmonizes mixed-source data before SLR Suite consumes the combined output.</p>
<p>Reproducibility controls run through the entire project. The suite uses renv to lock all 130 package dependencies, ships with automated tests and quality gates, and runs continuous integration on Ubuntu, macOS, and Windows to verify installation and execution on bundled data. The current release, version 2.3.2, is permanently archived on Zenodo with a specific Git commit identifier, and the developers report that the analytical source matches the benchmark source exactly. Seven successive public releases have progressively added modular project structure, dependency locking, end-to-end tests, software–documentation traceability, and static Sankey diagram export, documenting steady improvements in maintainability and technical reproducibility.</p>
<p>The TCCM coding module illustrates the software&#8217;s philosophy in detail. It combines titles, abstracts, author keywords, and Keywords Plus fields, normalizes the text, and tests every record against every configured pattern. Multiple matched labels are retained as semicolon-separated values, while records with no match are reported as missing rather than assigned an inferred label. The method is fully deterministic: identical text and dictionary versions always produce identical matches, and no probability or confidence scores are generated. Dictionary additions are expected to come with domain justification, version control, and regression checks against known examples, with researchers reviewing the results for false positives, false negatives, and ambiguous assignments.</p>
<p>To verify that the machinery actually works, the developer ran the full nine-module pipeline on an illustrative corpus of 754 publications from 2015 to 2026 covering artificial intelligence, tourism, and customer experience—a dataset spanning 390 sources, 2,519 unique authors, and more than 2,700 author keywords. A separate technical benchmark then executed the pipeline five times on each of three collections containing 20, 208, and 500 records. All 15 runs completed successfully, with all 135 module executions logged as successful and required outputs, including thematic Sankey visualizations, verified. Median execution time rose from about 19 seconds for the smallest collection to roughly 53 seconds for the largest, while approximate peak memory increased from about 474 to 544 MiB on a Windows 11 test machine.</p>
<p>The developers are notably careful about what these results do and do not demonstrate. Software verification and scientific validation are treated as strictly separate concerns: automated tests confirm that the code installs and executes its declared operations, but they do not establish that a literature search is complete, that eligibility decisions are correct, or that the generated classifications are scientifically valid. A blinded dual-AI pilot on 60 records, used only to dry-run the coding protocol, showed perfect agreement between the two AI coders and 96.67 percent agreement with the software, but the authors explicitly decline to present this as evidence of classification accuracy, since AI agents are not independent human researchers and no independently human-coded reference set exists. Notably, an earlier benchmark of version 2.3.1 exposed a genuine bug—duplicate year boundaries caused the thematic-evolution calculation to fail silently while logs indicated success—which motivated corrections to interval construction and error propagation in the current release.</p>
<p>How does SLR Suite compare with the crowded ecosystem of existing review tools? The accompanying comparison table spans bibliometrix, VOSviewer, CiteSpace, Rayyan, ASReview, DistillerSR, EPPI-Reviewer, Parsifal, Colandr, Nested Knowledge, SWARM-SLR, BibexPy, Covidence, ARC, and INRA SLR Copilot, each excelling at particular stages such as screening, science mapping, or AI-assisted triage. What distinguishes SLR Suite, according to the assessment, is its combination of a fully reproducible versioned workflow with documented TCCM support and partial SPAR support—capabilities none of the compared tools document. The authors are careful to note that the comparison does not establish absolute uniqueness or superior performance, and that controlled comparative evaluation would be required for such claims.</p>
<p>The project&#8217;s limitations are stated with unusual candor. It does not formulate review protocols, execute database searches, replace independent eligibility judgments, assess risk of bias, or validate scientific interpretations. In domains without an established theory base, the workflow cannot discover or infer new theories; researchers must first build a provisional vocabulary from domain literature and refine it iteratively. Future evaluation priorities include two-researcher manual coding with agreement analysis, usability testing with real participants, controlled comparisons with alternative workflows, and cross-domain replication in health sciences, social sciences, and engineering. Positioned as a transparent support environment rather than a replacement for researcher judgment, SLR Suite represents a growing movement in metascience: treating the review process itself as an auditable, version-controlled computational pipeline whose every decision can be inspected, rerun, and improved.</p>
<p><strong>Subject of Research:</strong> A reproducible open-source R software workflow for conducting systematic literature reviews</p>
<p><strong>Article Title:</strong> SLR suite: A reproducible R workflow for systematic literature reviews</p>
<p><strong>Article References:</strong> Akbulut, M. (2026). SLR suite: A reproducible R workflow for systematic literature reviews. <em>SoftwareX, 36</em>, Article 103101. <a href="https://doi.org/10.1016/j.softx.2026.103101" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103101</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103101" rel="noopener noreferrer">10.1016/j.softx.2026.103101</a></p>
<p><strong>Keywords:</strong> systematic literature review, SLR Suite, R programming, reproducibility, bibliometrics, TCCM framework, SPAR, PRISMA, open-source software, text classification, research synthesis, SoftwareX</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">260906</post-id>	</item>
	</channel>
</rss>
