Few moments in medicine carry the weight of a prognosis delivered after cardiac arrest. When a patient survives the initial collapse but remains comatose, families must often decide within days whether to continue life-sustaining treatment, and those decisions rest on an imperfect synthesis of neurologic examinations, electroencephalograms, brain imaging, blood biomarkers, and the accumulated bedside judgment of clinicians. Much of that reasoning never appears in any structured field of the electronic health record. It lives in the free-text clinical notes: the neurologist’s concern, the intensivist’s uncertainty, the nurse’s overnight observations, the record of a difficult family meeting. A new study published in Neurocritical Care asks whether machine learning can read between those lines, and whether the words clinicians write down might reveal not only a patient’s likely outcome but also the hidden logic, and possible biases, of prognostic decision-making itself.
The research, led by Furlow and colleagues, tackles one of the most consequential prediction problems in critical care. Up to 80 percent of patients who regain spontaneous circulation after cardiac arrest are comatose on admission, and their recovery is frequently uncertain. Because brain injury is the most common cause of death and disability in this population, accurate neuroprognostication is essential, yet reliable predictors remain limited and many apply only to narrow subsets of patients. Current guidelines emphasize a multimodal approach, combining examination findings, EEG patterns, imaging, and biomarkers rather than relying on any single signal. Even so, the process remains as much art as science, and the reasoning behind a withdrawal-of-care decision is often scattered across hundreds of narrative notes.
To build their model, the researchers extracted unstructured text from notes spanning the entire hospitalization of comatose cardiac arrest survivors treated at three hospitals. Natural language processing algorithms processed this text and identified potentially informative keywords, which were then used to train a logistic regression model to predict outcomes on a stratified Cerebral Performance Category scale, the standard framework for grading neurologic recovery after cardiac arrest. The baseline model performed well, achieving an area under the receiver operating characteristic curve of 0.9 and an area under the precision-recall curve of 0.89. Those figures place the note-based model in the range of respectable clinical prediction tools, though the authors caution that performance fell when the model was restricted to notes from only the first 24 hours of admission, with the decline especially pronounced for predicting good outcomes.
That last detail matters enormously, and it is where the accompanying commentary by Sarah Nutman and Christopher Horvat of the University of Pittsburgh sharpens the discussion. Post-cardiac arrest guidelines from the American Heart Association, the European Resuscitation Council and European Society of Intensive Care Medicine, and the Neurocritical Care Society recommend waiting at least 72 hours before attempting to prognosticate, because sedation, targeted temperature management, and the natural evolution of injury can cloud early assessment. A model trained on notes from the first day of admission is therefore being asked to make a judgment that guidelines explicitly say should not yet be made. Nutman and Horvat suggest that the drop in performance on early notes is unsurprising, and that a 72-hour window might have been more instructive for evaluating the model’s true clinical utility.
The commentary also raises a subtler statistical concern: temporal leakage. When a model is trained on notes drawn from the entire hospitalization, the patient’s eventual outcome can become progressively clearer as time passes, so the text may simply echo a conclusion that clinicians have already reached rather than genuinely predicting it. The authors of the original study worried that their models’ weaker performance on good outcomes creates a risk of self-fulfilling prophecies, in which a pessimistic algorithmic prediction influences the decision to withdraw care, thereby making the prediction come true. Nutman and Horvat propose an alternative interpretation: the asymmetry may reflect clinicians incorporating information that never makes it into the written record, leaving the model blind to the very signals that would allow it to recognize patients on a path to recovery.
Perhaps the most provocative contribution of the work is not the AUROC at all, but the window it opens into the hidden architecture of prognostic decision-making. By identifying language patterns associated with withdrawal of life-sustaining treatment, the researchers demonstrated that natural language processing can do more than forecast outcomes; it can interrogate the clinical reasoning, and the possible biases, embedded in routine documentation. Two findings stand out. Keywords related to intimate partner violence were significantly more common in the notes of patients on whom life-sustaining treatment was withdrawn. And although several guidelines recommend against using myoclonus as a prognostic marker, myoclonus appeared significantly more often in the records of patients whose care was withdrawn, raising the possibility that the finding is still quietly influencing bedside decisions despite the guideline warnings.
These observations point toward a use case that extends well beyond prediction. Even without a predictive model, the approach suggests that natural language processing could be deployed to systematically audit differences in clinical outcomes using variables that structured data never captures. In an era when health systems are increasingly attentive to equity, the ability to mine the narrative record for evidence of inconsistent decision-making represents a genuinely novel form of quality control. The same technique that flags a guideline-discordant reliance on myoclonus could, in principle, reveal whether prognostic pessimism is distributed unevenly across patient populations, institutions, or clinical teams.
The study also highlights an underappreciated dimension of prognostication: whose observations count. Neuroprognostication has historically been driven by physician perspectives, yet nurses, therapists, and other health professionals spend far more continuous time at the bedside and often notice changes, in spontaneous movements, responsiveness to family, or sleep-wake cycling, that physicians document less frequently. Incorporating these complementary observations through interdisciplinary evaluation may improve predictions, and the note-mining approach provides a structured mechanism to capture and integrate perspectives from across the clinical team. In that sense, the model is not merely reading the chart; it is aggregating the collective perception of everyone who wrote in it.
Still, significant questions remain before such an algorithm could approach the bedside. Post-arrest guidelines insist on a multimodal approach rather than reliance on any single predictor or algorithm, so the critical unknown is the additive benefit of this tool when used alongside established predictors, including the clinical examination, EEG and neuroimaging findings, and serum biomarkers. The original study compared its model’s performance to other natural language processing models, but not to known neuroprognostic predictors, which limits how much can be said about its incremental value. It also remains unknown whether the algorithm captures anything genuinely new, or whether it simply rediscovers, in narrative form, the same information clinicians already extract from structured testing. A model trained on clinical notes may be learning patient biology, clinician interpretation, family preferences, institutional practice patterns, or an inseparable mixture of all four, and in post-arrest care that distinction is not academic.
The stakes of getting this right could hardly be higher. A falsely pessimistic prediction does not merely misclassify a patient; it can trigger decisions that make the prediction come true, converting an algorithmic error into an irreversible outcome. As Nutman and Horvat conclude, the next step is not to build better note-based models but to determine what these models are actually learning, when their predictions become reliable, and how they should be used alongside established multimodal prognostic tools. Used carefully, natural language processing may help clinicians extract real signals from the clinical record and strengthen one of the most difficult judgments in medicine. Used carelessly, it risks transforming documentation artifacts into clinical authority. The enduring promise of this work lies in making the invisible parts of prognostication visible, and then deciding, with appropriate humility, which of those signals deserve to guide care.
Subject of Research: Natural language processing of clinical notes for neurologic outcome prediction after cardiac arrest
Article Title: Reading Between the Lines After Cardiac Arrest
Article References: Nutman, S. K., & Horvat, C. M. (2026). Reading Between the Lines After Cardiac Arrest. Neurocritical Care. https://doi.org/10.1007/s12028-026-02649-2
Image Credits: AI Generated
DOI: 10.1007/s12028-026-02649-2
Keywords: cardiac arrest, neuroprognostication, natural language processing, machine learning, electronic health records, clinical notes, coma, withdrawal of life-sustaining treatment, EEG, critical care, predictive models, healthcare bias
Cite Scienmag News
Ophelia Keating. (October 2, 2026). AI Reads Clinical Notes to Predict Recovery After Cardiac Arrest. Scienmag. https://scienmag.com/ai-reads-clinical-notes-to-predict-recovery-after-cardiac-arrest-2/
Ophelia Keating. "AI Reads Clinical Notes to Predict Recovery After Cardiac Arrest." Scienmag, 2 October 2026, https://scienmag.com/ai-reads-clinical-notes-to-predict-recovery-after-cardiac-arrest-2/. Accessed 2 October 2026.
Ophelia Keating. "AI Reads Clinical Notes to Predict Recovery After Cardiac Arrest." Scienmag. October 2, 2026. https://scienmag.com/ai-reads-clinical-notes-to-predict-recovery-after-cardiac-arrest-2/

