<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data quality &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/data-quality/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 00:41:54 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>data quality &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Rural Cancer Patients Get a Lifeline as Dresden Team Pilots Regional Tumour Network</title>
		<link>https://scienmag.com/rural-cancer-patients-get-a-lifeline-as-dresden-team-pilots-regional-tumour-network/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 00:41:54 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[cancer care networks]]></category>
		<category><![CDATA[cancer treatment access in rural areas]]></category>
		<category><![CDATA[clinical trial access for rural patients]]></category>
		<category><![CDATA[Clinical Trials]]></category>
		<category><![CDATA[collaboration between hospitals and specialists]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[Dresden cancer research]]></category>
		<category><![CDATA[Germany]]></category>
		<category><![CDATA[health equity]]></category>
		<category><![CDATA[health services research]]></category>
		<category><![CDATA[innovative cancer care models]]></category>
		<category><![CDATA[interdisciplinary oncology teams]]></category>
		<category><![CDATA[multicentre pilot study]]></category>
		<category><![CDATA[multidisciplinary tumour board]]></category>
		<category><![CDATA[patient pathways]]></category>
		<category><![CDATA[patient-reported outcomes]]></category>
		<category><![CDATA[pilot study]]></category>
		<category><![CDATA[regional tumour network]]></category>
		<category><![CDATA[rural cancer care]]></category>
		<category><![CDATA[rural healthcare access]]></category>
		<category><![CDATA[Surgical Oncology]]></category>
		<category><![CDATA[telemedicine in cancer treatment]]></category>
		<category><![CDATA[tumor board implementation]]></category>
		<category><![CDATA[underserved regions cancer management]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=220430</guid>

					<description><![CDATA[A multicentre pilot study in rural East Saxony is testing whether a central tumour board, standardised care pathways, and mobile data nurses can give countryside cancer patients the same access to specialised surgery and trials as city dwellers.]]></description>
										<content:encoded><![CDATA[<p>In the rural stretches of East Saxony, a cancer diagnosis often comes with a second, quieter burden: distance. Patients living far from Germany&#8217;s maximum-care hospitals can face long journeys, fragmented communication between local doctors and specialists, and limited access to clinical trials that might offer them the newest surgical and systemic therapies. A team of researchers and clinicians from TUD Dresden University of Technology, the University Hospital Carl Gustav Carus, and the National Center for Tumor Diseases (NCT) Dresden has now published the detailed study protocol for an ambitious attempt to close that gap. The project, known as MISSION4Sax, is described in the open-access Journal of Cancer Research and Clinical Oncology as a multicentre pilot study designed to test whether a tightly networked, cross-sectoral cancer care model can work in routine practice in a medically underserved region.</p>
<p>The core of the model is a regional surgical indication tumour board hosted at the NCT Dresden. In modern oncology, treatment decisions are ideally made not by a single physician but by an interdisciplinary panel of surgeons, medical and radiation oncologists, radiologists, and pathologists who review each case together. In practice, however, patients treated in smaller rural hospitals frequently never reach such a board, because the expertise is concentrated in urban maximum-care centres. MISSION4Sax brings the board to the region: cases from participating rural hospitals and oncology practices are presented to the central panel, which then issues evidence-based, standardised treatment recommendations. The goal is that every patient in the network, regardless of postcode, receives a guideline-based therapy recommendation and, where appropriate, access to innovative surgery and clinical trials.</p>
<p>To translate those recommendations into coordinated care, the study develops and pilots standardised patient care pathways across three regional hospitals and two oncology practices in East Saxony. These pathways define, step by step, what should happen after a cancer diagnosis: which diagnostics are needed, when the tumour board should be consulted, how surgery or systemic therapy is scheduled, and how follow-up is organised. Pathways of this kind are intended to reduce waiting times, prevent duplication of examinations, and make deviations from guideline care visible and correctable. The pilot study is explicitly exploratory and prospective in design, following patients longitudinally to assess how the model performs under real-world conditions rather than testing the causal effectiveness of any single intervention.</p>
<p>Because a care model is only as good as the data infrastructure beneath it, MISSION4Sax places unusual emphasis on study management and digital coordination. The project builds its database on the infrastructure of the NCT Core Unit Registry Trial Platform, a professional research data system already used for other oncology registry studies. Targeted training programmes prepare staff at the participating sites for the new processes, from documentation standards to communication routines. Perhaps the most distinctive element is the deployment of so-called Flying Data Nurses: mobile study personnel who travel between the rural hospitals and practices to support coordination, data flow, and communication on site. This travelling workforce is meant to solve a chronic problem in multicentre research, where small institutions often lack dedicated study staff and data quality suffers as a result.</p>
<p>The researchers also widen the lens beyond clinical measurements. Patient-reported outcomes, collected systematically, complement the clinical data and capture how patients themselves experience the care pathway: their symptoms, quality of life, and satisfaction with coordination and information. In addition, routinely collected health insurance data will be used strictly as external reference material for a descriptive comparison with the MISSION4Sax cohort. The protocol is careful to specify that no individual-level data linkage will be performed with these insurance records, a design choice that reflects both privacy considerations and the pilot character of the study. A mixed-methods process evaluation will run alongside the entire project, combining quantitative indicators with qualitative interviews and observations to understand not just whether the model functions, but why, for whom, and under what conditions.</p>
<p>What the study will actually measure follows directly from its feasibility focus. The team plans to assess whether the care model can be implemented in routine care, whether patients and professionals find it acceptable, and how well the study infrastructure supports data quality. Secondary exploratory questions concern access to specialised care, waiting times from diagnosis to treatment, adherence to the defined pathways, coordination between sectors, and clinical as well as patient-reported outcomes. This framing matters scientifically: many promising care innovations fail not because the underlying medicine is wrong but because logistics, communication, and workload make them unsustainable. By testing feasibility, acceptability, implementation, and data quality first, MISSION4Sax follows the recognised logic of pilot studies that prepare the ground for later, larger trials of effectiveness.</p>
<p>The regional context gives the project its urgency. East Saxony combines mid-sized towns and villages with long travel distances to tertiary centres, and like many rural areas in Germany and across Europe it faces a shortage of specialists and growing pressure on hospital infrastructure. For surgical oncology in particular, where the choice of operation, the timing of neoadjuvant therapy, and the availability of high-volume expertise can decisively influence outcomes, the distance between a patient and a maximum-care centre is not a trivial inconvenience. Strengthening cancer care infrastructure, interdisciplinary cooperation, and digital systems, the authors argue, may support equitable access to high-value care in rural regions, so that the quality of treatment no longer depends on where a patient happens to live.</p>
<p>The study is registered in the German Clinical Trials Register under identifier DRKS00036134, with registration dated 30 June 2025. Ethical approval was granted by the Ethics Committee of TUD Dresden University of Technology, and the study will be conducted in line with the principles of the Declaration of Helsinki, with informed consent obtained from all participants. The project is funded by the German Federal Ministry of Research, Technology and Space under grants 01KD2402A and 01KD2402B, with the funding institution explicitly not interfering in any part of the study. The participating institutions named in the acknowledgements include Krankenhaus Bautzen, Asklepios-ASB Krankenhaus Radeberg, Klinikum Oberlausitzer Bergland in Zittau, and two oncology care providers in the region, forming the practical backbone of the network.</p>
<p>For the wider field of health services research, the protocol offers a template worth watching. It combines a central specialist resource, standardised pathways, professional data management, mobile study staff, and systematic patient-reported measures into a single regional architecture, and it documents the evaluation plan in enough detail that other regions could replicate or adapt it. If the pilot demonstrates that the tumour board, the pathways, and the trained staff improve care quality and collaboration across sectors, the model could inform how rural cancer networks are built well beyond Saxony. Conversely, if the process evaluation reveals friction points, whether in documentation burden, communication delays, or staff capacity, those findings will be equally valuable, identifying precisely where investment is needed before the model is scaled.</p>
<p>The first results will show whether a Flying Data Nurse with a laptop and a central tumour board can do what decades of structural reform have struggled to achieve: make excellent cancer surgery and trial access a matter of course, rather than geography, for patients in the countryside. As the authors conclude, the proposed tumour board, patient pathways, and trained staff may improve care quality and collaboration across sectors, and the pilot study now underway in East Saxony will provide the evidence on whether that promise holds in the demanding reality of routine clinical care.</p>
<p><strong>Subject of Research:</strong> A multicentre pilot study protocol for cross-sectoral oncological care, data integration, and rural-urban networking in East Saxony</p>
<p><strong>Article Title:</strong> Optimized surgical treatment and study management for oncology patients in East Saxony (MISSION4Sax): study protocol of a multicentre pilot study</p>
<p><strong>Article References:</strong> Richter, C., Döring, A., Berg, F., Roost, M., Hasselberg, A., Martin, C., Schlieter, H., Piontek, D., Schoffer, O., Schmitt, J., Weitz, J., Krause-Jüttler, G., &amp; Kirchberg, J. (2026). Optimized surgical treatment and study management for oncology patients in East Saxony (MISSION4Sax): study protocol of a multicentre pilot study. <em>Journal of Cancer Research and Clinical Oncology, 152</em>(9), Article 173. <a href="https://doi.org/10.1007/s00432-026-06599-2" rel="noopener noreferrer">https://doi.org/10.1007/s00432-026-06599-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00432-026-06599-2" rel="noopener noreferrer">10.1007/s00432-026-06599-2</a></p>
<p><strong>Keywords:</strong> surgical oncology, multidisciplinary tumour board, patient pathways, health services research, rural healthcare access, cancer care networks, pilot study, patient-reported outcomes, data quality, Germany, clinical trials, health equity</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">220430</post-id>	</item>
		<item>
		<title>Scientists Devise a Three-Step Test to Check Whether ICU Data Annotations Can Be Trusted</title>
		<link>https://scienmag.com/scientists-devise-a-three-step-test-to-check-whether-icu-data-annotations-can-be-trusted/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 21:04:21 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI training data quality in neuro ICU]]></category>
		<category><![CDATA[annotation reliability in critical care]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[CENTER-TBI]]></category>
		<category><![CDATA[clinical data annotation]]></category>
		<category><![CDATA[clinical data plausibility assessment]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[high-resolution neurocritical care datasets]]></category>
		<category><![CDATA[ICU]]></category>
		<category><![CDATA[ICU data annotation verification]]></category>
		<category><![CDATA[impact of data quality on AI in healthcare]]></category>
		<category><![CDATA[intracranial pressure]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[medical note accuracy in intensive care units]]></category>
		<category><![CDATA[methods for validating ICU physiologic data]]></category>
		<category><![CDATA[neurocritical care]]></category>
		<category><![CDATA[neuromonitoring]]></category>
		<category><![CDATA[neurotrauma dataset analysis]]></category>
		<category><![CDATA[osmotherapy]]></category>
		<category><![CDATA[physiological signals]]></category>
		<category><![CDATA[systematic evaluation of clinical annotations]]></category>
		<category><![CDATA[traumatic brain injury]]></category>
		<category><![CDATA[traumatic brain injury monitoring]]></category>
		<category><![CDATA[traumatic brain injury treatment documentation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216339</guid>

					<description><![CDATA[Researchers have developed a three-step framework combining visual inspection and automated physiological classification to assess the reliability of time-stamped clinical annotations in the CENTER-TBI high-resolution neuromonitoring dataset.]]></description>
										<content:encoded><![CDATA[<p>In the intensive care units where traumatic brain injury patients fight for their lives, monitors stream torrents of physiological data every second: intracranial pressure, arterial blood pressure, heart rhythm, oxygen saturation. Woven into those waveforms are time-stamped notes made by bedside staff recording when a drug was given, a patient was turned, or an airway was suctioned. Those annotations are supposed to give the raw signals their clinical meaning. But a new study has exposed just how fragile that layer of human documentation can be, and offers the first systematic method for deciding which annotations deserve to be trusted.</p>
<p>Researchers analysing the high-resolution dataset from the Collaborative European Neuro Trauma Effectiveness Research in Traumatic Brain Injury study, known as CENTER-TBI, developed and tested a three-step framework for evaluating what they call annotation plausibility. Their work, published in the journal Neurocritical Care, examined more than 15,000 annotated interventions across 205 patients and found that a striking proportion of manually entered treatment records were physiologically implausible, a problem with serious consequences for the artificial intelligence tools increasingly built on such data.</p>
<p>The core difficulty is deceptively simple. High-frequency neuromonitoring has transformed neurocritical care research, enabling individualised treatment strategies and predictive models that detect rapid physiological changes. Yet the annotations that contextualise these signals are usually entered manually, often under pressure, and sometimes hours after the event. Staff at the 21 European centres participating in the CENTER-TBI high-resolution sub-study used a touch-screen interface in the ICM+ software to log interventions from nine predefined categories, ranging from osmotherapy and suctioning to physiotherapy and sedation changes. The fixed categories were designed to standardise documentation across centres, but the process remained vulnerable to human error, inconsistency and temporal imprecision.</p>
<p>Those weaknesses matter enormously for modern data science. Supervised machine learning models assume that their training labels are accurate, a condition rarely met in clinical settings. As previous research has highlighted, inconsistent human annotations can compromise model performance, producing incorrect predictions and undermining clinical decision-support systems. Most studies have simply treated annotation imprecision as a limitation to acknowledge rather than a problem to solve. The CENTER-TBI team set out to close that methodological gap with a pragmatic framework that does not require an impossible gold standard, since no definitive ground truth exists for retrospective bedside notes.</p>
<p>The first step of the framework relies on human eyes. Because suctioning and physiotherapy leave short, recognisable fingerprints in the physiological record, such as transient spikes in arterial blood pressure and intracranial pressure, they serve as reference events. A reviewer visually inspected each patient file, tolerating timing errors of up to twenty minutes, and classified the file as showing high, moderate or low evidence of a genuine relationship between annotations and signals. Files with fewer than ten total annotations were automatically deemed low evidence. A second independent reviewer then checked the classifications, achieving substantial agreement, with a Cohen&#8217;s kappa of 0.69, and only seven of the 205 recordings required reclassification after consensus review.</p>
<p>The results of this visual screen were broadly reassuring: 60 percent of files showed high evidence, 30.2 percent moderate evidence and 9.8 percent low evidence. Within the high-evidence files, an independent event-level review of 1,769 suctioning annotations found that 91.3 percent were physiologically plausible. In other words, where bedside teams documented consistently, their notes generally matched what the waveforms showed. The framework&#8217;s authors argue that this kind of file-level screening can inject a degree of trust into interventions, such as fluid boluses or sedation changes, for which algorithmic validation is not feasible.</p>
<p>The second and third steps turned to automation, using osmotherapy, the administration of mannitol or hypertonic saline to lower dangerous intracranial pressure, as a test case. The algorithm searched a forty-minute baseline window around each of the 388 osmotherapy annotations for sustained intracranial hypertension, defined as pressure above 20 mm Hg for at least five consecutive minutes. Events lacking such a baseline were rejected as implausible. Of the 248 events that passed this filter, the post-intervention window was then assessed: 67.7 percent were classified as effective, showing either a pressure drop of at least 10 mm Hg or normalisation, while 32.3 percent were ineffective. The rejected events, 36.1 percent of the total, typically showed intracranial pressure stably below the treatment threshold both before and after the recorded time stamp, strongly suggesting the annotations did not reflect genuine treatment responses.</p>
<p>The most compelling evidence for the framework came from the convergence of its two independent strategies. Among osmotherapy events from high-evidence files, only 18 percent were rejected by the algorithm, compared with 44.1 percent from moderate-evidence files and 57.3 percent from low-evidence files, a highly significant difference that persisted across alternative pressure thresholds of 15 and 25 mm Hg. Two methods, one qualitative and human, one quantitative and automated, arrived at consistent verdicts, supporting the credibility of both. The study also found that annotation counts bore no significant relationship to patients&#8217; six-month functional outcomes, and that documentation density peaked in the first 24 hours of intensive care before declining.</p>
<p>The researchers are careful about what their classifications mean. Treatment response is used here to separate plausible non-responders from annotations lacking a credible physiological context, not to judge the clinical efficacy of osmotherapy itself. Most ineffective events came from well-annotated files and showed prolonged pressure elevation, consistent with genuine treatment failure, whereas rejected events showed consistently low pressures. The team also acknowledges limitations: the framework cannot detect interventions performed but never documented, concurrent procedures can mimic reference signatures, and patients monitored exclusively with external ventricular drains were excluded because drain-open periods interrupt signal acquisition.</p>
<p>Beyond its immediate findings, the study carries a broader message for the era of clinical artificial intelligence. Until intensive care units systematically and automatically record medication types, doses and precise timings, manual annotations will remain the main bridge between physiological signals and clinical meaning, and that bridge needs inspection. The three-step framework, demonstrated for osmotherapy but applicable in principle to any intervention with definable physiological criteria, offers a practical audit tool. The authors call for external validation in other datasets and populations, but their conclusion is clear: in high-stakes neurocritical care, the data feeding tomorrow&#8217;s algorithms deserve the same scrutiny as the science built upon them.</p>
<p><strong>Subject of Research:</strong> A methodological framework for evaluating the physiological plausibility of time-stamped clinical annotations in high-frequency neuromonitoring data from traumatic brain injury patients</p>
<p><strong>Article Title:</strong> Time-Stamped Annotations in High-Frequency Physiological Data: Evaluating an Approach for Assessing Annotation Plausibility in the CENTER-TBI Dataset</p>
<p><strong>Article References:</strong> Turella, S., Beqiri, E., Bögli, S. Y., Ianosi, B., Olakorede, I., Zoerle, T., Tas, J., Helbok, R., Smielewski, P., the Smart Neuromonitoring to Support Precision Medicine in Acute Central Nervous Injury (SOPRANI) collaborators and the CENTER-TBI High-Resolution Sub-Study Participants and Investigators, Anke, A., Beer, R., Helbok, R., Bellander, B.-M., Nelson, D., Buki, A., Chevallard, G., Chieregato, A., Citerio, G., &#8230; Rehber, C. (2026). Time-Stamped Annotations in High-Frequency Physiological Data: Evaluating an Approach for Assessing Annotation Plausibility in the CENTER-TBI Dataset. <em>Neurocritical Care</em>. <a href="https://doi.org/10.1007/s12028-026-02637-6" rel="noopener noreferrer">https://doi.org/10.1007/s12028-026-02637-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12028-026-02637-6" rel="noopener noreferrer">10.1007/s12028-026-02637-6</a></p>
<p><strong>Keywords:</strong> traumatic brain injury, CENTER-TBI, neurocritical care, neuromonitoring, clinical data annotation, osmotherapy, intracranial pressure, machine learning, data quality, ICU, physiological signals, artificial intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216339</post-id>	</item>
		<item>
		<title>Bayesian Reasoning Problems Could Expose AI Bots Hiding in Online Surveys</title>
		<link>https://scienmag.com/bayesian-reasoning-problems-could-expose-ai-bots-hiding-in-online-surveys/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 21:18:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI detection in crowdsourced research]]></category>
		<category><![CDATA[Bayesian reasoning]]></category>
		<category><![CDATA[Bayesian reasoning in online survey validation]]></category>
		<category><![CDATA[Bayesian reasoning problems revealing AI bots]]></category>
		<category><![CDATA[behavioral research methods]]></category>
		<category><![CDATA[capability-gap test]]></category>
		<category><![CDATA[capability-gap testing for AI identification]]></category>
		<category><![CDATA[chatbot identification in research]]></category>
		<category><![CDATA[cognitive psychology methods for AI detection]]></category>
		<category><![CDATA[crowdsourcing]]></category>
		<category><![CDATA[data contamination]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[human versus AI performance in Bayesian tasks]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in behavioral studies]]></category>
		<category><![CDATA[large language models influencing survey responses]]></category>
		<category><![CDATA[natural frequencies]]></category>
		<category><![CDATA[online behavioral research integrity]]></category>
		<category><![CDATA[online research]]></category>
		<category><![CDATA[online survey data contamination]]></category>
		<category><![CDATA[positive predictive value]]></category>
		<category><![CDATA[predictive value calculation in survey validation]]></category>
		<category><![CDATA[Prolific]]></category>
		<category><![CDATA[signal detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212579</guid>

					<description><![CDATA[A new study proposes using Bayesian reasoning problems, with their well-established human performance ceilings, as calibrated detectors of large language model contamination in online research samples.]]></description>
										<content:encoded><![CDATA[<p>Online behavioral research is facing a quiet crisis. As large language models become woven into everyday life, researchers who recruit participants through crowdsourcing platforms increasingly suspect that some of their respondents are not human at all — or are humans outsourcing their answers to chatbots. A new study published in Behavior Research Methods proposes an elegant solution to this problem, and it comes from an unexpected corner of cognitive psychology: the humble Bayesian reasoning problem, a puzzle that humans have famously struggled with for fifty years.</p>
<p>The study, conducted by independent researcher Vera Wilde, introduces what she calls a capability-gap test. The logic is deceptively simple. Certain problems, such as calculating a positive predictive value from base-rate information, have been studied so extensively in humans that scientists know, with meta-analytic precision, exactly how well people can perform. When participants in an online study dramatically exceed those well-established human ceilings, the most plausible explanation is not superhuman cognition but machine assistance — a canary in the data coalmine, signaling that the sample may be contaminated by large language models.</p>
<p>The empirical basis for the proposal comes from two preregistered pilot studies of a Bayesian reasoning training tool, with a combined sample of 148 participants recruited through the Prolific platform. The tool was designed to teach people how to solve Bayesian inference problems, the kind of task exemplified by medical diagnosis questions: given a disease with a certain prevalence, a test with a certain sensitivity and false-positive rate, what is the probability that a person who tests positive actually has the disease? Decades of research, dating back to classic work by Daniel Kahneman and Amos Tversky and extended by Gerd Gigerenzer and Ulrich Hoffrage, have shown that most people fail such problems, even when the numbers are presented in natural frequency formats that make the underlying logic easier to grasp.</p>
<p>That failure is precisely what makes the task useful as a detector. A meta-analysis by McDowell and Jacobs found that only about 24 percent of people can solve a single Bayesian reasoning problem presented in natural frequency format — and that figure represents a ceiling, a level at which achieving a perfect score on a battery of five such problems is effectively unattainable for genuine human respondents. Yet in Wilde&#8217;s pilots, participants&#8217; accuracy on positive predictive value calculation problems reached roughly three times the established human performance ceiling. In the second pilot, 57 percent of participants achieved perfect 5-for-5 scores, a result that should be extraordinarily rare in an uncontaminated human sample.</p>
<p>The technical heart of the approach lies in distinguishing two outcome measures: accuracy and algorithm use. Accuracy refers simply to whether the participant produced the correct numerical answer. Algorithm use, by contrast, refers to evidence in the participant&#8217;s response that they actually followed the Bayesian reasoning process — for example, constructing a frequency tree, counting cases, or showing the intermediate steps of the calculation. This distinction matters because a training intervention designed to improve Bayesian reasoning should, if it works, change both measures in tandem. A large language model, however, can produce correct answers without any visible reasoning process, or with a reasoning process that does not respond to the training manipulation in the way human learning does.</p>
<p>By tracking both measures simultaneously, researchers can separate two rival explanations for suspiciously high performance. If accuracy spikes but algorithm use does not, contamination is the likelier culprit, because the model supplies answers without the participant acquiring the underlying skill. If both accuracy and algorithm use rise together, the pattern is consistent with authentic learning effects, and the treatment signal can be preserved rather than discarded. In this way, the capability-gap test does not merely flag bad data; it helps researchers decide which parts of their dataset reflect genuine psychological phenomena and which parts reflect machine-generated noise.</p>
<p>Wilde frames the detection problem itself as a signal detection problem, structurally analogous to the mass screenings for low-prevalence conditions — such as disease screening — around which the Bayesian reasoning literature was originally developed. Just as a medical test must balance hits against false alarms, a contamination detector must catch bot-driven responses without wrongly excluding honest participants who happen to be statistically savvy. The known reference distributions from the Bayesian reasoning literature make this calibration possible in a way that ad hoc attention checks cannot. Rather than relying on generic screening questions, researchers can compare observed performance against quantified human benchmarks and estimate the probability that a given response pattern arose from machine assistance.</p>
<p>The approach offers four practical advantages over existing data-quality tools. First, the human performance ceilings are grounded in meta-analyses rather than informal intuition, giving researchers a defensible threshold for suspicion. Second, the human–large language model performance gap on these problems is large, which increases the sensitivity of the test. Third, the known reference distributions allow nuanced assessment rather than crude pass–fail judgments. Fourth, Bayesian reasoning problems are easy to embed in existing surveys, requiring no special software or platform cooperation. Together, these properties make the method deployable at scale across the many fields — psychology, marketing, political science, epidemiology — that increasingly depend on online samples.</p>
<p>The stakes are considerable. Prior research on crowd work has documented substantial and growing use of large language models by online workers, and studies of data contamination in machine learning itself show how memorized content can masquerade as genuine capability. If a meaningful fraction of respondents in an online study are completing tasks with chatbot help, effect sizes may be distorted, replication attempts may fail for reasons that have nothing to do with the underlying science, and the credibility of entire literatures built on crowdsourced data could be undermined. The problem echoes an older statistical concern: John Tukey&#8217;s foundational work on sampling from contaminated distributions warned that even small amounts of contamination can seriously mislead inference drawn from nominally clean data.</p>
<p>Wilde is careful to note the provenance of the idea: neither pilot study was originally designed to validate a contamination detection method, and the proposal emerged from post hoc analysis of unexpectedly strong results. That origin makes the capability-gap test a promising hypothesis rather than a fully validated diagnostic, and the author provides practical recommendations for researchers who wish to use these problems as data-quality diagnostics while the validation literature matures. Both studies were preregistered on the Open Science Framework, and all data, materials, and analysis code are publicly available, allowing other teams to scrutinize and extend the approach. If the method holds up under broader testing, Bayesian reasoning problems — long a symbol of human statistical frailty — may find a second career as guardians of scientific integrity, ensuring that the data feeding behavioral science come from human minds rather than the machines trained on those minds&#8217; collective output.</p>
<p><strong>Subject of Research:</strong> Detecting large language model contamination in online behavioral research samples using Bayesian reasoning problems as capability-gap tests</p>
<p><strong>Article Title:</strong> An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples</p>
<p><strong>Article References:</strong> Wilde, V. (2026). An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples. <em>Behavior Research Methods, 58</em>(11), Article 301. <a href="https://doi.org/10.3758/s13428-026-03184-w" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03184-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03184-w" rel="noopener noreferrer">10.3758/s13428-026-03184-w</a></p>
<p><strong>Keywords:</strong> large language models, Bayesian reasoning, data contamination, online research, data quality, signal detection, natural frequencies, positive predictive value, crowdsourcing, behavioral research methods, capability-gap test, Prolific</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212579</post-id>	</item>
		<item>
		<title>New scoring framework reveals which textile data can survive a digital product passport</title>
		<link>https://scienmag.com/new-scoring-framework-reveals-which-textile-data-can-survive-a-digital-product-passport/</link>
		
		<dc:creator><![CDATA[Sloane Callahan]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 01:19:31 +0000</pubDate>
				<category><![CDATA[Climate]]></category>
		<category><![CDATA[automation]]></category>
		<category><![CDATA[blockchain readiness]]></category>
		<category><![CDATA[Challenges in digital product data verification]]></category>
		<category><![CDATA[Circular economy]]></category>
		<category><![CDATA[Circular economy and textile sustainability]]></category>
		<category><![CDATA[Consumer empowerment through product data]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[Data reliability in supply chain tracking]]></category>
		<category><![CDATA[Digital biography of textiles]]></category>
		<category><![CDATA[Digital Product Passport for textiles]]></category>
		<category><![CDATA[digital product passports]]></category>
		<category><![CDATA[Ecodesign for Sustainable Products Regulation]]></category>
		<category><![CDATA[ESPR]]></category>
		<category><![CDATA[EU regulation]]></category>
		<category><![CDATA[man-made cellulosic textiles]]></category>
		<category><![CDATA[Regulatory enforcement in textile sustainability]]></category>
		<category><![CDATA[standardisation]]></category>
		<category><![CDATA[Supply chain transparency in fashion industry]]></category>
		<category><![CDATA[sustainability indicators]]></category>
		<category><![CDATA[Sustainable fashion regulation in the EU]]></category>
		<category><![CDATA[Textile supply chain traceability]]></category>
		<category><![CDATA[textile supply chains]]></category>
		<category><![CDATA[traceability]]></category>
		<category><![CDATA[Traceability Degree Score for product data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211894</guid>

					<description><![CDATA[Researchers have created a Traceability Degree Score that grades 336 textile product data elements across ten lifecycle stages, revealing that only one in five is mature enough for the EU's incoming Digital Product Passports.]]></description>
										<content:encoded><![CDATA[<p>Europe is about to demand that every T-shirt, dress and bedsheet carry a digital biography. Under the Ecodesign for Sustainable Products Regulation, which entered into force in July 2024, textiles sold in the European Union will need a Digital Product Passport, a digital record that consolidates information on where a product&#8217;s fibres came from, how it was manufactured, what it is made of and what can happen to it at the end of its life. The ambition is sweeping: passports are meant to power informed consumer choice, regulatory enforcement and the circular economy all at once. But a new study argues that behind this policy momentum sits an uncomfortable question that regulators and companies alike have largely glossed over — is the underlying data actually traceable enough to make any of this work?</p>
<p>Researchers Prosper Dzidzienyo, Saeed Rahimpour and Luana Dessbesell, writing in Environmental and Sustainability Indicators, have developed what they call a Traceability Degree Score, a systematic framework for grading individual pieces of product data on how reliably they can be captured, verified and exchanged along a supply chain. Rather than asking whether data exists, the framework asks whether each data element can be systematically linked to its point of origin, checked for accuracy and transmitted across supply chain partners in a digitally structured format. That distinction matters, because the authors&#8217; analysis of a full textile lifecycle suggests that most of the data needed for passports currently sits in a grey zone: present in some form, but far from trustworthy or interoperable.</p>
<p>The framework rests on four indicators drawn from maturity-model theory, information quality research and supply chain traceability standards. The first is standardisation, which measures whether data are defined and reported in harmonised formats with persistent identifiers, so that a batch number means the same thing to a pulp mill in Finland and a retailer in Paris. The second is automation, assessing whether information is captured by sensors and enterprise systems rather than typed in by hand. The third is reliability, which captures whether measurements are accurate, consistent and independently verifiable. The fourth, blockchain readiness, is deliberately misnamed in the eyes of no one: it does not assume distributed ledgers are required, but instead evaluates whether data are structured, uniquely identifiable and linked to discrete lifecycle events, properties that any immutable, event-based traceability architecture would demand. Each data element is scored from one to five on every indicator, and the four scores are averaged into a single composite value.</p>
<p>To test the approach, the researchers constructed a detailed case study of a man-made cellulosic textile, a fabric made from 80 percent certified renewable wood and 20 percent recycled textile waste, both sourced in Finland. This hybrid of bio-based and circular material streams was chosen precisely because it mirrors the sourcing complexity that passport regulation is meant to address. Drawing on EU regulatory texts, the CIRPASS project&#8217;s passport prototypes, stakeholder consultations and life cycle assessment literature, the team identified 336 individual primary data elements spread across ten lifecycle nodes, from forestry and pulping through fibre, yarn, fabric and garment production to distribution, consumer use and end-of-life.</p>
<p>The results are striking for how unevenly traceability is distributed. Half of all data elements scored at the moderate level, meaning they exist in structured form but lack the full automation, verification or interoperability that seamless passport integration requires. Only 20 percent achieved the highest maturity, being digitally standardised, automatically captured, externally verified and ready for distributed ledger transmission. A quarter showed limited traceability, hampered by manual entry or inconsistent templates, while the remainder were essentially untraceable, concentrated in the consumer-use and end-of-life phases. In other words, the single most important stage for circular decisions — what happens to a garment after the sale — is exactly where the data infrastructure is weakest.</p>
<p>The node-by-node comparison sharpens that picture. Distribution and retail scored highest, with a mean traceability degree of 4.30, closely followed by dissolving pulp production at 4.28 and textile recycling at 4.13. The authors attribute the retail result to the quiet success of existing logistics standards: GS1 identifiers, batch numbering and trade barcodes have been standardised for decades, so the data foundations there are already mature. Manufacturing nodes such as fibre production, fabric making, yarn spinning and garment assembly landed in an intermediate band between 3.78 and 3.89, reflecting partial digitisation within factories but limited data exchange between companies. Raw material sourcing scored 3.55, while consumer use collapsed to 2.00 and end-of-life managed only 3.13. No single node achieved uniformly high scores across all four indicators, a finding the researchers say confirms that traceability maturity cannot be reduced to digitalisation alone.</p>
<p>The study also exposes a structural tension at the heart of the EU&#8217;s circular economy agenda. Digital Product Passports are explicitly intended to enable downstream value retention — reuse, repair and recycling — yet the phases where those decisions are made are the ones with the poorest data readiness. Consumer behaviour, device heterogeneity and the absence of standardised reporting at the end of a product&#8217;s life mean the data trail simply frays out. The researchers note that harmonised data definitions alone are insufficient: many nodes scored well on standardisation while lagging badly on automation and blockchain readiness, showing that agreeing on what to record is a very different problem from being able to record it reliably at scale.</p>
<p>Methodologically, the framework fills a gap that previous approaches left open. Earlier efforts to prioritise passport data, including ranking systems based on importance, availability and sensitivity, largely captured stakeholder perceptions of what data matters. Industry 4.0 maturity models, meanwhile, assess whole organisations rather than individual data points. By scoring data elements themselves, the new approach converts qualitative regulatory demands — data must be verifiable, data must be machine-readable — into a consistent numerical index that can pinpoint exactly where along a supply chain investment is most needed. The authors stress that the framework is technologically neutral: blockchain readiness measures architectural compatibility, not adoption, and European regulations remain agnostic about implementation technology.</p>
<p>The researchers acknowledge limitations. The case scores were derived from secondary sources reflecting average industrial realities in a Finnish context, and values would shift across regions, technologies and firms; the point of the exercise is to demonstrate a reproducible method rather than to deliver fixed numbers. The assessment is also a snapshot in time, whereas passport requirements and data systems will both evolve rapidly, calling for longitudinal studies and operational pilots. Inter-rater reliability and automated scoring are flagged as priorities for future work, as is validation across other product categories — the framework operates at the level of generic data elements, making it directly applicable to batteries, electronics, construction materials and other sectors that the passport regulation will eventually touch.</p>
<p>Still, the message for industry and policymakers is clear. As passport implementation accelerates, the readiness of product data is not a binary yes or no but a measurable, multi-dimensional capability that currently varies wildly across the lifecycle of even a well-documented textile. With only one in five data elements in the case study achieving full traceability maturity, and with the circular economy&#8217;s most critical decision points sitting at the bottom of the maturity curve, the gap between regulatory intent and operational data reality is wide. Diagnostic tools like the Traceability Degree Score offer a way to map that gap precisely — and, the authors argue, a way to close it with targeted investment rather than hopeful assumptions.</p>
<p><strong>Subject of Research:</strong> A traceability assessment framework for primary data elements in textile digital product passports</p>
<p><strong>Article Title:</strong> A framework for evaluating the traceability of primary data elements for textile digital product passports</p>
<p><strong>Article References:</strong> Dzidzienyo, P., Rahimpour, S., &amp; Dessbesell, L. (2026). A framework for evaluating the traceability of primary data elements for textile digital product passports. <em>Environmental and Sustainability Indicators, 32</em>, Article 101514. <a href="https://doi.org/10.1016/j.indic.2026.101514" rel="noopener noreferrer">https://doi.org/10.1016/j.indic.2026.101514</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.indic.2026.101514" rel="noopener noreferrer">10.1016/j.indic.2026.101514</a></p>
<p><strong>Keywords:</strong> digital product passports, textile supply chains, traceability, EU regulation, ESPR, circular economy, data quality, blockchain readiness, standardisation, automation, man-made cellulosic textiles, sustainability indicators</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211894</post-id>	</item>
		<item>
		<title>Pay Per Beep: Simple Incentive Tweaks Can Boost Experience Sampling Data Without Hurting Quality</title>
		<link>https://scienmag.com/pay-per-beep-simple-incentive-tweaks-can-boost-experience-sampling-data-without-hurting-quality/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 23:23:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[burden]]></category>
		<category><![CDATA[careless responding]]></category>
		<category><![CDATA[compliance]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[data quality in behavioral research]]></category>
		<category><![CDATA[ecological momentary assessment]]></category>
		<category><![CDATA[effect of monetary incentives on data completeness]]></category>
		<category><![CDATA[experience sampling]]></category>
		<category><![CDATA[experience sampling incentives]]></category>
		<category><![CDATA[experimental design in psychological research]]></category>
		<category><![CDATA[impact of payment methods on data quality]]></category>
		<category><![CDATA[incentives]]></category>
		<category><![CDATA[influence of incentive structures on survey response]]></category>
		<category><![CDATA[intensive longitudinal data]]></category>
		<category><![CDATA[KU Leuven study on experiment participation]]></category>
		<category><![CDATA[methodology for improving experience sampling compliance]]></category>
		<category><![CDATA[participant engagement in experience sampling]]></category>
		<category><![CDATA[participant payment]]></category>
		<category><![CDATA[personalized feedback]]></category>
		<category><![CDATA[personalized feedback in survey participation]]></category>
		<category><![CDATA[psychological methods]]></category>
		<category><![CDATA[real-time mood and behavior tracking]]></category>
		<category><![CDATA[smartphone survey response rates]]></category>
		<category><![CDATA[survey methodology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211190</guid>

					<description><![CDATA[A randomized experiment with 192 students found that paying participants per beep increased compliance in a 14-day experience sampling study without reducing data quality, while personalized feedback offered no measurable advantage over a flat payment.]]></description>
										<content:encoded><![CDATA[<p>Every beep of a smartphone survey is a small negotiation between science and daily life. Experience sampling — the method psychologists use to capture thoughts, moods, and behaviors in real time by pinging participants repeatedly throughout the day — lives or dies on whether people actually answer. A missed beep is not just an empty cell in a spreadsheet; it is a lost snapshot of someone&#8217;s lived experience, and enough lost snapshots can quietly erode the foundations of a study. Now a team of researchers at KU Leuven in Belgium has put one of the field&#8217;s most practical design questions under the experimental microscope: does the way you pay participants change how much data they give you, and how good that data is?</p>
<p>The study, published in the journal Behavior Research Methods, was led by Milla Pihlajamäki together with Ginette Lafit, Olivia J. Kirtley, Inez Myin-Germeys, and Gudrun Eisele. The team recruited 192 students and randomly assigned 64 of them to each of three incentive conditions. The first group received a fixed payment for taking part, the standard arrangement in most experience sampling research. The second group received the same fixed payment plus personalized feedback, a summary of the data they had contributed, on the theory that seeing one&#8217;s own psychological patterns might motivate more careful responding. The third group was paid per beep, earning money incrementally with every completed prompt, a structure that mirrors the piece-rate logic of behavioral economics rather than a flat salary model.</p>
<p>The protocol itself was demanding by design. For fourteen consecutive days, participants received nine prompts — or beeps — per day on their smartphones, producing up to 126 measurement moments per person and roughly 24,000 potential data points across the sample. This intensity is precisely what makes experience sampling so scientifically valuable and so operationally fragile. Unlike a one-off questionnaire, an intensive longitudinal design asks people to interrupt whatever they are doing — in a lecture, at dinner, with friends — to report on their inner states. Compliance, in this context, is not a given; it is an achievement that must be engineered through careful protocol design, and incentives are one of the most powerful levers available.</p>
<p>What makes the study methodologically notable is that the researchers did not stop at counting completed surveys. They systematically distinguished between data quantity and data quality, treating the two as separable outcomes that incentives might influence in different directions. Quantity was straightforward: compliance rates, or the proportion of beeps answered, and how that proportion changed over the two-week study period. Quality required more nuance. The team screened for careless responding — the phenomenon where participants click through items without genuine attention — and examined whether the temporal dynamics of such responses shifted across the study, alongside retrospective measures of participant burden and experience.</p>
<p>Careless responding matters because its consequences are surprisingly severe. As the methodological literature has repeatedly shown, even a modest proportion of inattentive respondents can distort correlations, inflate or deflate reliability estimates, and lead analysts astray in ways that are hard to detect after the fact. In experience sampling data, the problem is compounded by the sheer number of measurement moments and the fatigue that accumulates as the days wear on. A participant might respond attentively on day one and slide into patterned, mechanical answers by day ten. The Leuven team, drawing on their own prior work on the temporal dynamics of careless responding, built this dimension directly into the analysis, using analyses of variance and multilevel linear and logistic regression models to test whether incentive structure affected not only average behavior but its trajectory over time.</p>
<p>The headline finding is refreshingly clean: payment per beep increased compliance rates compared with the other two conditions, and that was essentially where the differences ended. The incremental payment scheme did not make participants answer more carelessly, did not change their reported burden, and did not alter their retrospective experience of the study in measurable ways. Personalized feedback, meanwhile — despite its intuitive appeal and growing popularity in mobile health and self-tracking applications — produced no measurable advantage over a plain fixed payment on any of the outcomes examined. The feedback condition neither boosted compliance nor improved the quality of responses, a null result that carries real practical weight for researchers weighing the added complexity of generating individualized reports against any assumed motivational payoff.</p>
<p>The finding that pay-per-beep worked without degrading quality is worth unpacking, because incentives in psychology have a complicated reputation. Classic motivation research has documented cases where external rewards crowd out intrinsic motivation or encourage gaming of the reward structure. In an experience sampling context, one might have worried that paying per beep would incentivize rapid, low-effort completion — quantity at the expense of quality. Recent related work on game-based rewards in experience sampling found exactly that pattern: gamification increased data quantity but reduced data quality. The Leuven results suggest that a straightforward monetary increment per completed beep does not carry the same risk, at least not in a motivated student population completing a two-week protocol.</p>
<p>The authors&#8217; recommendation follows directly from the evidence: for motivated samples such as students, use payment per beep to boost data quantity. This is not a trivial prescription. Compliance in intensive longitudinal research varies enormously across populations and study designs, with meta-analyses reporting widely different completion rates depending on the sample, the burden of the protocol, and the population under study. Clinical populations, adolescents, and people experiencing acute distress often show lower compliance, and every percentage point of missing data constrains the statistical models that researchers use to extract meaning from these datasets. Multilevel time-series analyses, network models of psychological dynamics, and idiographic prediction approaches all benefit directly from denser, more complete measurement streams.</p>
<p>There are, of course, boundaries to the conclusion. The study was conducted in a student sample, a population that is generally well-motivated, digitally fluent, and accustomed to participating in research. Whether payment per beep would deliver the same benefit — and remain free of quality costs — in harder-to-reach populations, longer protocols, or studies measuring stigmatized behaviors remains an open question. The reward magnitudes involved were also specific to this context; the psychology of incremental incentives plausibly depends on the size of each increment relative to the total payment. And burden, while measured, was assessed retrospectively and through momentary reports within a demanding 14-day design; incentive effects on the lived experience of participation might differ in studies with different rhythms and demands.</p>
<p>Still, the study exemplifies a broader and welcome shift in psychological methods research: treating study design choices — payment schemes, feedback provision, sampling frequency, questionnaire length — as empirical questions rather than conventions inherited from previous studies. The team post-registered their study, documented every deviation transparently, and made all materials and analysis code openly available on the Open Science Framework, with data accessible through a secure checkout system. In a field increasingly aware that methodological decisions ripple through every downstream finding, knowing that a simple per-beep payment can fill in more of the data grid without compromising what fills those cells is exactly the kind of actionable, evidence-based guidance that turns methodological debate into better science.</p>
<p><strong>Subject of Research:</strong> Effects of incentive type on data quantity and quality in experience sampling research</p>
<p><strong>Article Title:</strong> The effect of incentive type on data quality and quantity in an experience sampling study in a student population</p>
<p><strong>Article References:</strong> Pihlajamäki, M., Lafit, G., Kirtley, O. J., Myin-Germeys, I., &amp; Eisele, G. (2026). The effect of incentive type on data quality and quantity in an experience sampling study in a student population. <em>Behavior Research Methods, 58</em>(11), Article 300. <a href="https://doi.org/10.3758/s13428-026-03164-0" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03164-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03164-0" rel="noopener noreferrer">10.3758/s13428-026-03164-0</a></p>
<p><strong>Keywords:</strong> experience sampling, ecological momentary assessment, incentives, compliance, data quality, careless responding, survey methodology, participant payment, personalized feedback, intensive longitudinal data, burden, psychological methods</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211190</post-id>	</item>
		<item>
		<title>Tea Bags, Reminder Cards and Texts Boost Survey Responses in Breast Cancer Trial</title>
		<link>https://scienmag.com/tea-bags-reminder-cards-and-texts-boost-survey-responses-in-breast-cancer-trial/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 14:13:01 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[addressing questionnaire attrition in cancer survivorship studies]]></category>
		<category><![CDATA[breast cancer]]></category>
		<category><![CDATA[Breast cancer survivor survey response rates]]></category>
		<category><![CDATA[cancer survivors]]></category>
		<category><![CDATA[cancer survivorship]]></category>
		<category><![CDATA[challenges in follow-up data collection for breast cancer trials]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[enhancing data collection in longitudinal health research]]></category>
		<category><![CDATA[impact of reminder cards and text messages on survey participation]]></category>
		<category><![CDATA[improving questionnaire return in medical studies]]></category>
		<category><![CDATA[improving trustworthiness of patient-reported]]></category>
		<category><![CDATA[innovative methods to boost survey completion]]></category>
		<category><![CDATA[long-term clinical trial participant engagement]]></category>
		<category><![CDATA[mixed methods study]]></category>
		<category><![CDATA[mixed-methods approach to survey nonresponse]]></category>
		<category><![CDATA[participatory research]]></category>
		<category><![CDATA[patient retention in clinical trials]]></category>
		<category><![CDATA[questionnaires]]></category>
		<category><![CDATA[Randomized Controlled Trial]]></category>
		<category><![CDATA[reminder strategies]]></category>
		<category><![CDATA[response rates]]></category>
		<category><![CDATA[return to work]]></category>
		<category><![CDATA[strategies to reduce nonresponse in health research]]></category>
		<category><![CDATA[survey nonresponse]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205687</guid>

					<description><![CDATA[A mixed-methods study embedded in the FASTRACS randomized controlled trial found that personalized reminders, printed questionnaires and phone-based completion, co-designed with breast cancer survivors, raised questionnaire response rates dramatically during long-term follow-up.]]></description>
										<content:encoded><![CDATA[<p>Every researcher who has ever run a long-term clinical trial knows the quiet dread of watching response rates slide as the months pass. Questionnaires that once flowed back from enthusiastic participants begin to trickle, then stall, and the study that was designed to capture the lived experience of patients over years gradually loses the very data that give it scientific power. For the teams behind FASTRACS, a French randomized controlled trial evaluating a return-to-work intervention for breast cancer survivors, that dread became reality: as follow-up wore on, questionnaire completion declined markedly, threatening the integrity of an study built around 431 participants across multiple sites. Rather than accept the attrition as inevitable, the researchers launched a second investigation, FASTRACS-INSIGHT, to understand precisely why survivors stopped responding and what could realistically be done to bring them back.</p>
<p>The INSIGHT study, whose name spells out its mission as Identify Nonresponse and Strategies to Improve Questionnaires&#8217; Trustworthiness, adopted a mixed-methods design that treated nonresponse not as a statistical nuisance but as a phenomenon with identifiable, and potentially correctable, causes. The first phase was a needs assessment with two complementary strands. On the quantitative side, the team carried out a statistical analysis of the FASTRACS trial data to pinpoint the subgroups of participants at higher risk of falling silent, comparing the characteristics of respondents and nonrespondents at successive follow-up points. On the qualitative side, they conducted individual interviews with survivors to hear, in their own words, the barriers that had pushed them away from questionnaires and the facilitators that might draw them back.</p>
<p>What emerged from that analysis was strikingly consistent with what survey methodologists have long suspected but rarely tested inside an ongoing oncology trial. Nonrespondents were younger than their responding counterparts, and they reported more systemic side effects from their breast cancer treatment, the fatigue, pain and cognitive fog that follow chemotherapy and endocrine therapy. Three broad domains of influence were identified: the breast cancer experience itself, the design of the questionnaires, and the context in which completion took place, whether at the hospital or at home. A younger woman juggling a return to work, childcare and lingering treatment toxicity, the findings suggest, experiences a mailed or online questionnaire very differently from an older, retired participant with fewer competing demands.</p>
<p>Armed with that diagnosis, the researchers turned to a participatory development process. A brainstorming session brought together breast cancer survivors drawn from the FASTRACS Intersectoral Strategic Participatory Committee, a standing body created to embed patient and stakeholder voices in the project, and additional phone interviews gathered participant perspectives. This was not a token consultation exercise. The ideas that survived into the final intervention were those the survivors themselves judged acceptable, practical and likely to matter to their peers. The researchers then telephoned trial participants to introduce the strategies and to gauge their willingness to complete questionnaires under the new conditions, effectively testing acceptability in real time before formal evaluation.</p>
<p>Four concrete components made it into the FASTRACS-INSIGHT intervention, and their simplicity is part of their charm. Participants received personalized reminder cards, each accompanied by a tea bag, a small gesture of warmth designed to make the request feel human rather than bureaucratic. Printed paper questionnaires were sent to every participant, removing the digital barrier for those who preferred or needed pen and paper. A medical appointments booklet, bundled with a FASTRACS-branded pen, offered a practical tool that kept the study visible in daily life. Finally, the team reinforced telephone-based questionnaire completion, allowing women to answer surveys with a researcher over the phone, and deployed text message reminders timed to prompt action without overwhelming it.</p>
<p>The evaluation phase compared response rates before and after the intervention was rolled out, and the results were unambiguous. At the third follow-up time point, designated T3, questionnaire response rates rose from 54.4 percent to 74.4 percent, a difference that was statistically significant at p less than 0.001. At the fourth time point, T4, the improvement was even more dramatic: completion climbed from 46.2 percent to 76.7 percent, again with p less than 0.001. In plain terms, roughly half of the survivors who had been dropping out of the study&#8217;s data collection were brought back into the fold. For a longitudinal randomized controlled trial, in which every lost questionnaire erodes statistical power and invites attrition bias, a gain of twenty to thirty percentage points at successive waves is not a cosmetic improvement but a meaningful rescue of the study&#8217;s evidentiary foundation.</p>
<p>The findings also speak to a deeper methodological problem in health research. Survey nonresponse is rarely random, and the INSIGHT study confirms that in the context of cancer survivorship it is patterned by age and by treatment burden. Younger survivors face what the literature describes as distinct work-related, clinical and psychological pressures during the return-to-work period, and women experiencing more systemic side effects may find a lengthy questionnaire an intolerable additional demand rather than a welcome chance to contribute. If researchers ignore these patterns, their longitudinal datasets become progressively skewed toward older, healthier, more engaged participants, quietly distorting the very outcomes, quality of life, fatigue, functioning and employment, that survivorship research aims to measure. By identifying who was at risk and why, INSIGHT offered a template for protecting data quality prospectively rather than statistically patching it after the fact.</p>
<p>Equally instructive is the participatory character of the solution. None of the four components, the reminder cards with tea bags, the universal printed questionnaires, the appointments booklet with pen, or reinforced phone completion and text reminders, involved sophisticated technology or prohibitive cost. What distinguished them was that they were co-designed with the people expected to use them. The authors argue that involving breast cancer survivors in developing the intervention was essential to tailoring it to their needs, and the trial&#8217;s before-and-after results give empirical weight to that claim. The approach echoes the intervention mapping framework the FASTRACS team has used throughout the parent trial, in which theory, evidence and stakeholder input are systematically combined to design health behavior interventions. Response behavior, INSIGHT demonstrates, is itself a behavior that can be mapped, understood and changed.</p>
<p>The implications extend well beyond this single French trial. Randomized controlled trials in oncology routinely lose participants to questionnaire fatigue, and systematic reviews of survey methodology have catalogued dozens of strategies, from monetary incentives to shortened instruments, with mixed and context-dependent results. What INSIGHT adds is a replicable, multi-step playbook: analyze who is not responding, interview a sample to understand their reasons, co-develop low-burden interventions with a patient committee, and evaluate the package with straightforward before-and-after comparisons. The registered protocols underlying the work, the FASTRACS-RCT under NCT04846972 and the RECOVA-FASTRACS qualitative protocol under NCT05919498, with an amendment covering INSIGHT, give the effort a transparency that methodologists will welcome. Informed consent was obtained from all participants, and the authors report no competing interests, with funding from the League against cancer in the Auvergne Rhône-Alpes region, the National Agency for the Improvement of Working Conditions, the National Cancer Institute, and joint support from the Pension and Welfare Fund for Notaries&#8217; Clerks and Employees.</p>
<p>There is a quietly radical message here for the research enterprise. Behind every unreturned questionnaire is a person navigating one of the hardest chapters of her life, and the difference between a study that retains its participants and one that loses them may come down to whether the researchers asked those participants what would help. A tea bag and a handwritten-feeling reminder card will not cure cancer or eliminate treatment side effects, but they signal respect, and respect, as the FASTRACS-INSIGHT team has shown with response rates rising from the mid-forties to the mid-seventies, turns out to be a remarkably effective instrument of science. For the growing community of survivorship researchers studying return to work after cancer, an outcome with well-documented personal and economic stakes, the lesson is clear: sustainable data collection is not a logistics problem to be solved in a back office, but a relationship to be built with the survivors themselves, one thoughtfully designed touchpoint at a time.</p>
<p><strong>Subject of Research:</strong> Nonresponse causes and response-rate improvement strategies in a randomized controlled trial of return-to-work support for breast cancer survivors</p>
<p><strong>Article Title:</strong> Understanding nonresponse and strategies to improve response rates in a randomized controlled trial among breast cancer survivors: the FASTRACS-INSIGHT mixed-methods study</p>
<p><strong>Article References:</strong> Soulier, A., Blazer, A., Préau, M., Guittard, L., Maurin, A., Fervers, B., Letrilliart, L., Carretier, J., Broc, G., Fassier, J.-B., Péron, J., Rouat, S., &amp; Lamort-Bouché, M. (2026). Understanding nonresponse and strategies to improve response rates in a randomized controlled trial among breast cancer survivors: the FASTRACS-INSIGHT mixed-methods study. <em>Journal of Cancer Survivorship</em>. <a href="https://doi.org/10.1007/s11764-026-02132-z" rel="noopener noreferrer">https://doi.org/10.1007/s11764-026-02132-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11764-026-02132-z" rel="noopener noreferrer">10.1007/s11764-026-02132-z</a></p>
<p><strong>Keywords:</strong> breast cancer, cancer survivors, randomized controlled trial, survey nonresponse, questionnaires, response rates, return to work, cancer survivorship, mixed-methods study, participatory research, reminder strategies, data quality</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205687</post-id>	</item>
		<item>
		<title>Attitude, Training and Trust in Data Drive Use of Ethiopia&#8217;s Digital Health Records</title>
		<link>https://scienmag.com/attitude-training-and-trust-in-data-drive-use-of-ethiopias-digital-health-records/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:05:54 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[barriers to health data adoption]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[data utilization]]></category>
		<category><![CDATA[DHIS2]]></category>
		<category><![CDATA[DHIS2 health information system]]></category>
		<category><![CDATA[digital health transformation in Ethiopia]]></category>
		<category><![CDATA[electronic health record implementation]]></category>
		<category><![CDATA[Ethiopia]]></category>
		<category><![CDATA[Ethiopia digital health records]]></category>
		<category><![CDATA[evidence-based decision making in public health]]></category>
		<category><![CDATA[evidence-based decision-making]]></category>
		<category><![CDATA[Haramaya University]]></category>
		<category><![CDATA[health data utilization in Ethiopia]]></category>
		<category><![CDATA[health informatics]]></category>
		<category><![CDATA[health information system challenges]]></category>
		<category><![CDATA[health information systems]]></category>
		<category><![CDATA[health worker attitudes towards digital health]]></category>
		<category><![CDATA[healthcare data management in Ethiopia]]></category>
		<category><![CDATA[performance monitoring]]></category>
		<category><![CDATA[public health facilities]]></category>
		<category><![CDATA[routine health information systems]]></category>
		<category><![CDATA[supportive supervision]]></category>
		<category><![CDATA[trust and training in health data systems]]></category>
		<category><![CDATA[workplace culture and data use]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202484</guid>

					<description><![CDATA[A study of 220 health workers in eastern Ethiopia found that only about 45 percent use DHIS2 data, with attitude, self-competence, perceived data quality and supportive supervision as the key determinants.]]></description>
										<content:encoded><![CDATA[<p>In the public health facilities of eastern Ethiopia, a digital revolution has quietly been underway for years. The District Health Information System 2, known widely as DHIS2, is an open-source software platform designed to collect, store, analyze and distribute health data from the most remote clinics to the highest levels of health administration. On paper, it promises a transformation: health managers should be able to monitor disease trends, track service performance and make evidence-based decisions in near real time. In practice, however, a new study from Haramaya University reveals that the promise is only half fulfilled. Nearly half of the health workers surveyed in the Harari Regional State and the Dire Dawa City Administration are not actually using the data their own facilities generate, and the reasons why turn out to be as much about human psychology and workplace culture as about technology.</p>
<p>The research, published in BMC Health Services Research, was led by Daniel Gudina, Gudeta Ayele, Behailu Hawulte and Ibsa Mussa from the School of Public Health at Haramaya University&#8217;s College of Health and Medical Sciences. The team set out to answer a deceptively simple question: what determines whether health workers in public facilities actually use the data flowing through DHIS2? This question matters because data utilization is one of the key functions of any health information system. When performance monitoring teams can read and interpret their own data, they can identify gaps in immunization coverage, spot outbreaks earlier, allocate staff and supplies more rationally and ultimately improve the quality of health service delivery. When they cannot, or do not, the entire investment in digital infrastructure risks becoming an expensive exercise in data entry without insight.</p>
<p>To investigate, the researchers conducted an institutional-based cross-sectional study between June 1, 2020 and July 31, 2020. They used a stratified sampling technique to select 220 participants working in public health facilities across the two administrative areas of eastern Ethiopia. Data collection relied on a structured questionnaire complemented by an observational checklist, allowing the team to capture both what health workers reported about their practices and what could be verified on the ground. The completed questionnaires were thoroughly checked, coded and entered into Epi-data 3.1 before being transferred to Stata 14 for statistical analysis. To identify the factors associated with data utilization, the researchers applied a binary logistic regression model, treating results as statistically significant when the p value fell below 0.05 with 95 percent confidence intervals.</p>
<p>The headline finding is sobering. Overall utilization of data from DHIS2 stood at about 45 percent, with a 95 percent confidence interval ranging from 39 to 50 percent. That figure falls well below the national recommended threshold, which calls for more than two-thirds of health facilities to be actively using their routine health information data for decision making. In other words, even in a system where the software has been deployed and staff have been trained to feed it, a majority of facilities are not translating the numbers into action. The study also notes that prior research across Ethiopia has reported significant variability in DHIS2 data utilization, suggesting that this is not a local anomaly but a systemic challenge with roots that vary from place to place.</p>
<p>What, then, separates the facilities that use their data from those that do not? The regression analysis identified four independent determinants, and each one is revealing. The first is attitude. Health workers who held a favorable attitude toward data use were roughly three times more likely to utilize DHIS2 data, with an adjusted odds ratio of 3.0 and a 95 percent confidence interval of 1.90 to 8.62. Attitude, in this context, reflects whether staff believe that reviewing data is worth their time, whether they see it as relevant to their daily work and whether they feel that acting on evidence is part of their professional identity. A health worker who views monthly data review as a bureaucratic chore will behave very differently from one who sees it as a diagnostic tool for improving services.</p>
<p>The second determinant is perceived self-competence. Workers who felt confident in their ability to interpret and apply the data were nearly three times more likely to use it, with an adjusted odds ratio of 2.9 and a confidence interval of 1.14 to 7.38. This finding speaks to a well-documented problem in health information systems across low- and middle-income countries: staff may be trained to enter data but not to analyze it. DHIS2 offers dashboards, pivot tables and visualization tools, but these features are only useful to people who understand what the outputs mean and how they connect to programmatic decisions. Self-doubt, in this environment, becomes a silent barrier. A nurse or health officer who feels incompetent with data will avoid opening the very reports that could guide their work, and the avoidance reinforces the incompetence in a self-perpetuating cycle.</p>
<p>The third and strongest single factor was perceived data quality. Health workers who believed the data in the system were accurate, complete and timely were more than four times as likely to use them, with an adjusted odds ratio of 4.4 and a confidence interval of 1.76 to 10.9. This is perhaps the most intuitive of the findings, and also the most troubling. If staff suspect that the numbers in DHIS2 are riddled with errors, duplicates or gaps, they will reasonably distrust any conclusion drawn from them. Data quality and data use are locked in a feedback loop: poor quality suppresses use, and without use there is little incentive or feedback mechanism to correct quality problems. Breaking that loop requires deliberate investment in data verification, feedback to data enterers and a culture in which accuracy is valued and rewarded rather than assumed.</p>
<p>The fourth determinant was supportive supervision. Facilities whose staff received supportive supervision were more than four times as likely to use DHIS2 data, with an adjusted odds ratio of 4.3 and a confidence interval of 1.48 to 12.45. Supportive supervision, in the language of the Performance of Routine Health Information System framework, known as PRISM, means supervisors who do more than inspect forms. They review data with frontline staff, help troubleshoot technical problems, encourage discussion of trends and connect the numbers to concrete service improvements. The PRISM framework, which underpins much of the conceptual thinking in this field, holds that technical, behavioral and organizational determinants together shape whether routine health information systems deliver value. The Ethiopian findings map neatly onto that framework: attitude and self-competence are behavioral determinants, perceived data quality is a technical one, and supportive supervision is organizational.</p>
<p>The implications for policy are direct. Ethiopia&#8217;s Federal Ministry of Health has invested substantially in DHIS2 as the backbone of its routine health information system, and the platform is expected to increase the utilization of health data nationwide. But the study&#8217;s results suggest that software deployment alone does not close the gap between data availability and data use. Interventions should target the four determinants identified: building favorable attitudes toward evidence-based practice, strengthening the analytical self-competence of health workers through practical mentorship rather than one-off training, assuring and communicating data quality so that staff trust what they see, and institutionalizing supportive supervision so that every facility benefits from regular, constructive engagement around its own numbers. The performance monitoring teams that exist in Ethiopian facilities could become the natural vehicle for this work, provided they are equipped and encouraged to function as intended.</p>
<p>There is also a broader lesson for the global health informatics community. DHIS2 is now used in dozens of countries, and the dream of a digital health information backbone is closer to reality than ever. Yet the eastern Ethiopia study is a reminder that the last mile of any information system is human. A dashboard nobody opens is indistinguishable from a filing cabinet nobody opens. The researchers, whose work was financially supported by the Doris Duke Charitable Foundation as part of the Capacity Building and Mentorship Program project, with no funder role in study design or interpretation, obtained ethical clearance from Haramaya University&#8217;s Institutional Health Research Ethics Review Committee in accordance with the Helsinki II declaration. Their message to health systems everywhere is clear: to unlock the value of digital health data, invest as much in confidence, trust and supervision as in servers and software. Until the people closest to the data believe in its quality and in their own ability to act on it, the numbers will keep flowing, and the decisions will keep waiting.</p>
<p><strong>Subject of Research:</strong> Determinants of District Health Information System 2 data utilization among health workers in public health facilities in eastern Ethiopia</p>
<p><strong>Article Title:</strong> Determinants of district health information system 2 data utilization in public health facilities in Harari Regional States and Dire Dawa City Administration, Eastern Ethiopia</p>
<p><strong>Article References:</strong> Determinants of district health information system 2 data utilization in public health facilities in Harari Regional States and Dire Dawa City Administration, Eastern Ethiopia. (n.d.). <a href="https://doi.org/10.1186/s12913-026-15620-w" rel="noopener noreferrer">https://doi.org/10.1186/s12913-026-15620-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12913-026-15620-w" rel="noopener noreferrer">10.1186/s12913-026-15620-w</a></p>
<p><strong>Keywords:</strong> DHIS2, health information systems, data utilization, Ethiopia, public health facilities, health informatics, supportive supervision, data quality, evidence-based decision making, routine health information systems, Haramaya University, performance monitoring</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202484</post-id>	</item>
		<item>
		<title>New Open-Source Platform Puts Data Maturity Self-Assessment in Every Organization&#8217;s Hands</title>
		<link>https://scienmag.com/new-open-source-platform-puts-data-maturity-self-assessment-in-every-organizations-hands/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:26:31 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-driven data improvement roadmap]]></category>
		<category><![CDATA[automated data governance scoring]]></category>
		<category><![CDATA[cost-effective data process capability assessment]]></category>
		<category><![CDATA[data governance]]></category>
		<category><![CDATA[data management]]></category>
		<category><![CDATA[data management maturity model]]></category>
		<category><![CDATA[Data maturity assessment]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[international data standards ISO 8000 and IEC 33000]]></category>
		<category><![CDATA[ISO 8000]]></category>
		<category><![CDATA[ISO/IEC 33000]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[maturity models]]></category>
		<category><![CDATA[open-access data assessment software]]></category>
		<category><![CDATA[open-source data governance platform]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[process capability]]></category>
		<category><![CDATA[scalable data quality evaluation]]></category>
		<category><![CDATA[self-assessment]]></category>
		<category><![CDATA[self-assessment for organizational data capability]]></category>
		<category><![CDATA[SoftwareX]]></category>
		<category><![CDATA[standards-based data quality management]]></category>
		<category><![CDATA[UNE 0080]]></category>
		<category><![CDATA[web-based data management tool]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196295</guid>

					<description><![CDATA[Researchers have released DQPA, an open-source platform that automates ISO/IEC 33000-based self-assessment of data governance, management, and quality maturity, matching expert assessors' results while generating AI-curated improvement roadmaps.]]></description>
										<content:encoded><![CDATA[<p>Data has become the defining asset of the modern organization, yet most institutions still have no reliable way of knowing how well they actually govern, manage, and safeguard the quality of that data. Formal maturity assessments exist, anchored in international standards, but they are expensive, slow, and dependent on scarce expert assessors. A team of Spanish researchers now believes it has cracked the problem. In a paper published in the open-access journal SoftwareX, Fernando Gualo, Yolanda Ayuso, Ismael Caballero, and Mario Piattini of the University of Castilla-La Mancha and the Alarcos Research Group introduce DQPA, a web-based software platform that allows any organization to run a rigorous, standards-compliant self-assessment of its data governance, data management, and data quality maturity—complete with automated scoring and artificial intelligence–generated improvement roadmaps—as an open-source tool released under the GNU AGPL v3.0 license.</p>
<p>The scientific foundation of DQPA rests on two pillars of international standardization. ISO 8000 establishes the principles of data quality management, while the ISO/IEC 33000 family provides the general mechanism for process capability assessment: a process reference model, process attributes rated on an ordinal scale, and a maturity model that aggregates those attributes into organizational levels. Although this mechanism has a long track record in domains such as software development, Green IT, and data quality certification, no international instantiation had ever covered data governance, data management, and data quality management jointly as integrated disciplines. The only ISO instantiation for the data domain, the ISO 8000-6x series, is confined to data quality management alone. The researchers built on a direct precedent, the MAMD model, and on the UNE 0077 through 0080 specifications—what they describe as the first standardization initiative to close that gap with normative status—enriched with the governance principles of ISO/IEC 38505 and the DAMA-DMBOK body of knowledge.</p>
<p>What distinguishes self-assessment from formal certification is its purpose. Under ISO/IEC 33000, assessment by an independent team enables certification with validity toward third parties, while self-assessment, performed by the organization itself, is an equally recognized application of the same method aimed at understanding one&#8217;s own situation recurrently and affordably as a basis for continuous improvement. Until now, no adequate instrument has existed for this second application, for two reasons. Independent assessment is too costly to repeat frequently, and its most expensive phase—evidence collection—depends heavily on tacit human knowledge that is difficult to automate. Moreover, translating normative processes into language that business profiles can act upon, one of the very purposes of data governance, rarely occurs in manual practice. Existing frameworks fall short in different ways: COBIT 2019 lacks an integrated data-domain maturity model; DCAM and CMMI-DMM support self-assessment but are not grounded in ISO/IEC 33000; and DAMA-DMBOK systematizes the disciplines without defining a maturity model of its own. None combines integrated coverage, an ISO/IEC 33000-based mechanism, automated scoring, automated recommendations, and open-source availability.</p>
<p>DQPA&#8217;s architecture is deliberately engineered around a strict separation between deterministic computation and generative artificial intelligence. The platform is a multilayer web application: a React single-page client, a Node.js and Express server exposing a REST API and hosting the deterministic assessment engine, a MongoDB document store accessed through Mongoose, and an external large language model service invoked only after results have been computed. Authentication is token-based with role-based access control, and both tiers deploy as independent containers. Crucially, the normative model itself—processes, questions, weightings, and improvement tasks—is maintained as configurable data through an administration module restricted to expert users, meaning the platform can adapt to revisions of the specifications or even to equivalent frameworks without touching the source code. This data-driven design is what makes the tool reusable and future-proof in a way that hard-coded assessment instruments cannot be.</p>
<p>The heart of the platform is its assessment engine, which algorithmically reproduces the measurement framework of ISO/IEC 33020. Question answers are first converted into weighted percentage scores at two granularities: one per process, from process-specific capability-level-1 questions, and one per capability level from 2 to 5, from cross-cutting questions shared across the scope. The four-category achievement scale maps onto these scores: Not implemented for 0 to 15 percent, Partially implemented for above 15 to 50 percent, Largely implemented for above 50 to 85 percent, and Fully implemented for above 85 to 100 percent. The staged aggregation rule then applies: a process reaches a given capability level when its process-specific score and all lower cross-cutting scores are Fully achieved and the score at that level is at least Largely achieved. The organizational maturity level, on a six-level scale from 0 to 5, is derived from the consolidated capability of the assessed processes in a staged manner. The entire computation is deterministic, traceable, and executed without any intervention from humans or the AI service—a design choice the authors argue is a property rather than a limitation, because transparency and auditability are explicit requirements of the ISO/IEC 33000 method itself.</p>
<p>Only after the numbers are settled does artificial intelligence enter the picture—and the boundaries are strict. The recommendation service neither trains nor fine-tunes any model. Instead, a pre-trained language model, Gemini 2.0 Flash Lite, selected after a structured comparison against alternatives including GPT-4o and GPT-4.1 nano for its large context window and high throughput, is conditioned at inference time by a purpose-built structured prompt. The model&#8217;s grounding is entirely deterministic: its input consists solely of the computed as-is state—levels, ratings, and gaps—and a catalogue of predefined improvement tasks curated by domain experts from an anonymized corpus of real projects, assessment reports, standards, and technical documentation. The prompt explicitly forbids the re-computation of levels and imposes prioritization criteria including impact on maturity, criticality of the gap, dependencies, feasibility, urgency, and normative alignment. If the model call fails, the service degrades gracefully to a template-based ordering of the curated catalogue, so a usable plan is always produced without the generative component.</p>
<p>The platform was validated in striking fashion against a real organization: a Spanish public river basin management body responsible for hydrological data acquisition and exploitation. Questionnaire responses were recorded in parallel through DQPA and through the manual procedure of an external expert assessment team, with each business process owner completing the instrument in roughly two and a half hours with support from assessors. The engine computed a largely achieved process-specific score, a partially achieved level-2 dimension, and an unattained level-3 dimension, yielding maturity level 1—precisely the level determined independently by the human experts. Verification went further: the engine&#8217;s logic was exhaustively checked against a reference spreadsheet used by the consulting team in professional practice, across all 1,024 possible rating combinations, with full agreement in every case, including boundary conditions between capability levels.</p>
<p>The quality of the AI-generated recommendations was then scrutinized by four expert evaluators—three of them external to the author team and blind to the study—who rated thirty recommendations on a five-point rubric covering consistency with the computed state, alignment with the specifications, actionability, and clarity for non-technical profiles. The overall mean was 4.35 out of 5, with no two evaluators differing by more than one point on any of the 480 ratings, and the restricted external-only mean of 4.23 confirmed the result does not depend on the internal rater. The platform was also applied to three further organizations—a local public administration, a port authority operating critical infrastructure, and a large private technology corporation—producing consistent operation across markedly different sectors, with resulting maturity levels ranging from 0 to 1. Perceived usability, measured with the System Usability Scale across four participants, averaged 80.6, comfortably above the scale&#8217;s commonly cited average of about 68. Performance testing showed the deterministic engine computing results in a median of 105 milliseconds and sustaining 150 concurrent users, while end-to-end report generation took a median of three seconds, dominated by the external model call.</p>
<p>The implications reach well beyond convenience. For public administrations and resource-constrained organizations facing obligations under the European Data Governance Regulation, DQPA substantially lowers the barrier to understanding and improving their data practices without depending on scarce certified assessors. For researchers, the platform&#8217;s elimination of inter-assessor variability in the computation phase opens the door to genuinely reproducible, longitudinal, and sector-level empirical study of data maturity—questions such as which processes systematically act as bottlenecks, or how maturity evolves after improvement plans are applied, that have been nearly impossible to address empirically until now. The authors are careful to delimit their claims: the platform does not verify declared evidence, cannot prevent deliberate misstatement, does not replace third-party certification, and its recommendations must be contextualized by each organization since the AI knows nothing of internal budgets or politics. Future work includes ablation studies contrasting catalogue-anchored with unconstrained generation, a self-hosted AI deployment for stricter data-residency control, sector-level benchmarking, and what-if simulation of improvement scenarios under explicit resource constraints. But the core message is already clear: the machinery once reserved for expensive consulting engagements has been reproduced as transparent, inspectable, open-source software that any organization can pick up and run.</p>
<p><strong>Subject of Research:</strong> An open-source software platform for ISO/IEC 33000-based self-assessment of data governance, data management, and data quality maturity</p>
<p><strong>Article Title:</strong> DQPA: A software platform for ISO/IEC 33000-based self-assessment of data governance, data management, and data quality maturity</p>
<p><strong>Article References:</strong> Gualo, F., Ayuso, Y., Caballero, I., &amp; Piattini, M. (2026). DQPA: A software platform for ISO/IEC 33000-based self-assessment of data governance, data management, and data quality maturity. <em>SoftwareX, 35</em>, Article 103012. <a href="https://doi.org/10.1016/j.softx.2026.103012" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103012</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103012" rel="noopener noreferrer">10.1016/j.softx.2026.103012</a></p>
<p><strong>Keywords:</strong> data governance, data quality, data management, maturity models, ISO/IEC 33000, ISO 8000, self-assessment, process capability, large language models, open-source software, SoftwareX, UNE 0080</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196295</post-id>	</item>
	</channel>
</rss>
