In the intensive care units where traumatic brain injury patients fight for their lives, monitors stream torrents of physiological data every second: intracranial pressure, arterial blood pressure, heart rhythm, oxygen saturation. Woven into those waveforms are time-stamped notes made by bedside staff recording when a drug was given, a patient was turned, or an airway was suctioned. Those annotations are supposed to give the raw signals their clinical meaning. But a new study has exposed just how fragile that layer of human documentation can be, and offers the first systematic method for deciding which annotations deserve to be trusted.
Researchers analysing the high-resolution dataset from the Collaborative European Neuro Trauma Effectiveness Research in Traumatic Brain Injury study, known as CENTER-TBI, developed and tested a three-step framework for evaluating what they call annotation plausibility. Their work, published in the journal Neurocritical Care, examined more than 15,000 annotated interventions across 205 patients and found that a striking proportion of manually entered treatment records were physiologically implausible, a problem with serious consequences for the artificial intelligence tools increasingly built on such data.
The core difficulty is deceptively simple. High-frequency neuromonitoring has transformed neurocritical care research, enabling individualised treatment strategies and predictive models that detect rapid physiological changes. Yet the annotations that contextualise these signals are usually entered manually, often under pressure, and sometimes hours after the event. Staff at the 21 European centres participating in the CENTER-TBI high-resolution sub-study used a touch-screen interface in the ICM+ software to log interventions from nine predefined categories, ranging from osmotherapy and suctioning to physiotherapy and sedation changes. The fixed categories were designed to standardise documentation across centres, but the process remained vulnerable to human error, inconsistency and temporal imprecision.
Those weaknesses matter enormously for modern data science. Supervised machine learning models assume that their training labels are accurate, a condition rarely met in clinical settings. As previous research has highlighted, inconsistent human annotations can compromise model performance, producing incorrect predictions and undermining clinical decision-support systems. Most studies have simply treated annotation imprecision as a limitation to acknowledge rather than a problem to solve. The CENTER-TBI team set out to close that methodological gap with a pragmatic framework that does not require an impossible gold standard, since no definitive ground truth exists for retrospective bedside notes.
The first step of the framework relies on human eyes. Because suctioning and physiotherapy leave short, recognisable fingerprints in the physiological record, such as transient spikes in arterial blood pressure and intracranial pressure, they serve as reference events. A reviewer visually inspected each patient file, tolerating timing errors of up to twenty minutes, and classified the file as showing high, moderate or low evidence of a genuine relationship between annotations and signals. Files with fewer than ten total annotations were automatically deemed low evidence. A second independent reviewer then checked the classifications, achieving substantial agreement, with a Cohen’s kappa of 0.69, and only seven of the 205 recordings required reclassification after consensus review.
The results of this visual screen were broadly reassuring: 60 percent of files showed high evidence, 30.2 percent moderate evidence and 9.8 percent low evidence. Within the high-evidence files, an independent event-level review of 1,769 suctioning annotations found that 91.3 percent were physiologically plausible. In other words, where bedside teams documented consistently, their notes generally matched what the waveforms showed. The framework’s authors argue that this kind of file-level screening can inject a degree of trust into interventions, such as fluid boluses or sedation changes, for which algorithmic validation is not feasible.
The second and third steps turned to automation, using osmotherapy, the administration of mannitol or hypertonic saline to lower dangerous intracranial pressure, as a test case. The algorithm searched a forty-minute baseline window around each of the 388 osmotherapy annotations for sustained intracranial hypertension, defined as pressure above 20 mm Hg for at least five consecutive minutes. Events lacking such a baseline were rejected as implausible. Of the 248 events that passed this filter, the post-intervention window was then assessed: 67.7 percent were classified as effective, showing either a pressure drop of at least 10 mm Hg or normalisation, while 32.3 percent were ineffective. The rejected events, 36.1 percent of the total, typically showed intracranial pressure stably below the treatment threshold both before and after the recorded time stamp, strongly suggesting the annotations did not reflect genuine treatment responses.
The most compelling evidence for the framework came from the convergence of its two independent strategies. Among osmotherapy events from high-evidence files, only 18 percent were rejected by the algorithm, compared with 44.1 percent from moderate-evidence files and 57.3 percent from low-evidence files, a highly significant difference that persisted across alternative pressure thresholds of 15 and 25 mm Hg. Two methods, one qualitative and human, one quantitative and automated, arrived at consistent verdicts, supporting the credibility of both. The study also found that annotation counts bore no significant relationship to patients’ six-month functional outcomes, and that documentation density peaked in the first 24 hours of intensive care before declining.
The researchers are careful about what their classifications mean. Treatment response is used here to separate plausible non-responders from annotations lacking a credible physiological context, not to judge the clinical efficacy of osmotherapy itself. Most ineffective events came from well-annotated files and showed prolonged pressure elevation, consistent with genuine treatment failure, whereas rejected events showed consistently low pressures. The team also acknowledges limitations: the framework cannot detect interventions performed but never documented, concurrent procedures can mimic reference signatures, and patients monitored exclusively with external ventricular drains were excluded because drain-open periods interrupt signal acquisition.
Beyond its immediate findings, the study carries a broader message for the era of clinical artificial intelligence. Until intensive care units systematically and automatically record medication types, doses and precise timings, manual annotations will remain the main bridge between physiological signals and clinical meaning, and that bridge needs inspection. The three-step framework, demonstrated for osmotherapy but applicable in principle to any intervention with definable physiological criteria, offers a practical audit tool. The authors call for external validation in other datasets and populations, but their conclusion is clear: in high-stakes neurocritical care, the data feeding tomorrow’s algorithms deserve the same scrutiny as the science built upon them.
Subject of Research: A methodological framework for evaluating the physiological plausibility of time-stamped clinical annotations in high-frequency neuromonitoring data from traumatic brain injury patients
Article Title: Time-Stamped Annotations in High-Frequency Physiological Data: Evaluating an Approach for Assessing Annotation Plausibility in the CENTER-TBI Dataset
Article References: Turella, S., Beqiri, E., Bögli, S. Y., Ianosi, B., Olakorede, I., Zoerle, T., Tas, J., Helbok, R., Smielewski, P., the Smart Neuromonitoring to Support Precision Medicine in Acute Central Nervous Injury (SOPRANI) collaborators and the CENTER-TBI High-Resolution Sub-Study Participants and Investigators, Anke, A., Beer, R., Helbok, R., Bellander, B.-M., Nelson, D., Buki, A., Chevallard, G., Chieregato, A., Citerio, G., … Rehber, C. (2026). Time-Stamped Annotations in High-Frequency Physiological Data: Evaluating an Approach for Assessing Annotation Plausibility in the CENTER-TBI Dataset. Neurocritical Care. https://doi.org/10.1007/s12028-026-02637-6
Image Credits: AI Generated
DOI: 10.1007/s12028-026-02637-6
Keywords: traumatic brain injury, CENTER-TBI, neurocritical care, neuromonitoring, clinical data annotation, osmotherapy, intracranial pressure, machine learning, data quality, ICU, physiological signals, artificial intelligence
Cite Scienmag News
Cassandra Pierce. (September 26, 2026). Scientists Devise a Three-Step Test to Check Whether ICU Data Annotations Can Be Trusted. Scienmag. https://scienmag.com/scientists-devise-a-three-step-test-to-check-whether-icu-data-annotations-can-be-trusted/
Cassandra Pierce. "Scientists Devise a Three-Step Test to Check Whether ICU Data Annotations Can Be Trusted." Scienmag, 26 September 2026, https://scienmag.com/scientists-devise-a-three-step-test-to-check-whether-icu-data-annotations-can-be-trusted/. Accessed 26 September 2026.
Cassandra Pierce. "Scientists Devise a Three-Step Test to Check Whether ICU Data Annotations Can Be Trusted." Scienmag. September 26, 2026. https://scienmag.com/scientists-devise-a-three-step-test-to-check-whether-icu-data-annotations-can-be-trusted/

