<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>clinical notes &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/clinical-notes/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 20:19:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>clinical notes &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Doctors Know Stigmatizing Words Harm Patients, Yet Most Cannot Spot Them in Charts</title>
		<link>https://scienmag.com/doctors-know-stigmatizing-words-harm-patients-yet-most-cannot-spot-them-in-charts/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 20:19:11 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[clinical documentation]]></category>
		<category><![CDATA[clinical notes]]></category>
		<category><![CDATA[clinician attitudes towards stigmatizing language]]></category>
		<category><![CDATA[clinician awareness of stigmatizing words]]></category>
		<category><![CDATA[effect of language on healthcare disparities]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[faculty development]]></category>
		<category><![CDATA[healthcare disparities]]></category>
		<category><![CDATA[healthcare provider training on stigmatization]]></category>
		<category><![CDATA[hospital documentation practices]]></category>
		<category><![CDATA[identifying bias in medical charts]]></category>
		<category><![CDATA[impact of language on patient care]]></category>
		<category><![CDATA[internal medicine]]></category>
		<category><![CDATA[medical bias]]></category>
		<category><![CDATA[Medical documentation bias]]></category>
		<category><![CDATA[Medical Education]]></category>
		<category><![CDATA[medical education on bias recognition]]></category>
		<category><![CDATA[Open Notes]]></category>
		<category><![CDATA[patient-centered communication in hospitals]]></category>
		<category><![CDATA[patient-centered language]]></category>
		<category><![CDATA[stigmatizing language]]></category>
		<category><![CDATA[stigmatizing language in clinical notes]]></category>
		<category><![CDATA[survey research]]></category>
		<category><![CDATA[tools for detecting bias in medical records]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=245481</guid>

					<description><![CDATA[A multicenter survey of internal medicine clinicians finds that while nearly all recognize the importance of avoiding stigmatizing language in medical records, fewer than one in twenty could identify all stigmatizing terms on a knowledge test and most lack confidence in their own documentation.]]></description>
										<content:encoded><![CDATA[<p>Every day, clinicians at hospitals across the United States write thousands of notes describing their patients. Some of those notes contain words that quietly shape how the next clinician will treat the person on the other end of the stethoscope: descriptions of patients as manipulative, non-compliant, difficult, or not credible. A new multicenter survey published in the Journal of General Internal Medicine reveals a striking paradox at the heart of modern medical documentation. Nearly every physician and trainee surveyed believed that word choice matters and could bias care, yet almost none could reliably identify stigmatizing language when it appeared in front of them, and only a small fraction felt confident that their own notes were free of it.</p>
<p>The study, led by Julia Caton of Northwell Health and colleagues at Stanford School of Medicine, surveyed internal medicine residents, hospitalist faculty, and advanced practice providers at two large academic medical centers between April and September 2025. The researchers combined a traditional attitude survey with a knowledge test, asking participants to spot stigmatizing terms embedded in realistic snippets of clinical documentation. In total, 82 faculty and advanced practice providers responded, a 52 percent response rate, along with 63 residents, a 21 percent response rate. Because no validated instrument existed for this emerging field, the team built their survey from scratch, refining it through cognitive interviews designed to ensure that respondents at both institutions interpreted each question the same way.</p>
<p>The results expose a profound gap between intention and ability. Ninety-one percent of faculty and advanced practice providers and 92 percent of residents reported noticing stigmatizing language when reading clinical documentation. Roughly nine in ten said they considered how their own word choices might bias other clinicians toward or against a patient, and large majorities also considered the impact on patients and families who might read the notes. Yet when asked how confident they were that their documentation consistently avoided stigmatizing language, only 10 percent of faculty and advanced practice providers and 24 percent of residents reported being very or extremely confident. The clinicians knew the problem existed and believed it mattered, but they did not trust themselves to solve it.</p>
<p>The knowledge test made the scale of that distrust concrete. Participants were asked to identify 11 stigmatizing terms hidden across five documentation excerpts, with the terms selected through a literature review to capture a range of stigmatizing sentiments. Only 4 percent of faculty and advanced practice providers and 3 percent of residents correctly identified all 11. Mean scores were nearly identical between the two groups, with faculty and providers averaging 61 percent correct and residents averaging 59 percent, a difference that was not statistically significant. Experience, in other words, conferred no advantage. Senior attending physicians performed no better than first-year trainees at recognizing language that their own profession&#8217;s literature has repeatedly linked to worse care.</p>
<p>Recognition varied dramatically from term to term, and that variation may be the study&#8217;s most instructive finding. Terms such as non-compliant and claims were widely recognized as stigmatizing, while denies and the use of scare quotes around patient statements were far less frequently flagged. This pattern suggests that many clinicians have absorbed a short list of forbidden buzzwords without internalizing the underlying principles of patient-centered writing that would generalize to unfamiliar contexts. Education that simply teaches clinicians which words to avoid, the authors argue, will fail; effective training must convey the reasoning behind patient-centered language so that clinicians can navigate novel situations on their own.</p>
<p>The downstream consequences of stigmatizing documentation are not hypothetical. A 2018 study found that clinicians exposed to biased clinical vignettes developed more negative attitudes toward patients and became less willing to prescribe adequate pain management. More recently, researchers demonstrated that biased language used during verbal handoffs impaired clinical recall among medical students and residents and was associated with decreased expressions of empathy. Multiple studies have shown that stigmatizing language appears more frequently in the charts of Black patients, women, patients with substance use disorders, and patients with public insurance, raising the possibility that documentation practices actively perpetuate existing healthcare disparities. The medical record, once written, becomes a durable transmission channel for bias, read by every subsequent clinician who touches the case.</p>
<p>Open notes have raised the stakes considerably. Under the 21st Century Cures Act, patients in the United States can read their clinical notes immediately, and a recent study found that 10.5 percent of patients reported feeling offended or judged by something they read in a clinician&#8217;s note. A note is no longer a private communication between professionals; it is a document the patient may read within hours, one that can erode trust precisely when trust is most needed for treatment to succeed. The researchers note that this shift transforms documentation word choice from a stylistic quibble into a core clinical competency, one that the Accreditation Council for Graduate Medical Education implicitly recognizes through its Patient- and Family-Centered Communication milestone.</p>
<p>Faculty behavior adds another layer to the problem. Seventy-one percent of faculty said that avoiding stigmatizing language was very or extremely important for trainees, yet only 17 percent reported often or always providing feedback when they noticed stigmatizing language in resident notes. The most commonly cited barrier was lack of time, selected by 66 percent of faculty, followed by lack of prior training in giving feedback on biased language, cited by 29 percent, and lack of prioritization, cited by 26 percent. Documentation practices have traditionally been learned implicitly, through observation and the internalization of unspoken norms, rather than through explicit instruction. If supervisors rarely correct biased language due to time pressure and their own uncertainty, harmful habits pass unchallenged from one generation of clinicians to the next.</p>
<p>The study&#8217;s free-text responses revealed genuine tensions that any intervention must confront. Respondents worried about over-policing language and suppressing clinically meaningful information; a blanket prohibition on documenting that a patient refused a treatment, for example, could obscure safety-relevant events. Others pointed to conflicts with standardized terminology: terms such as obesity carry specific ICD-10 codes, and substituting alternative phrasing without guidance could affect reimbursement, risk adjustment, and disease capture in administrative datasets. Documentation simultaneously serves as a communication tool, a legal record, a billing instrument, and a quality measurement input, and optimizing language for one function can create friction in another. The authors argue that education must therefore frame patient-centered writing not as word prohibition but as a skill for achieving accuracy, respect, and clinical utility simultaneously, delivered with psychological safety and acknowledgment of legitimate gray zones such as direct quotations.</p>
<p>The findings arrive at a moment of technological inflection. Ambient listening technologies and large language model-assisted note generation are rapidly entering clinical workflows, and if these tools are trained on existing clinical notes, they risk encoding and perpetuating the very stigmatizing patterns that educators are trying to eliminate. The researchers have begun piloting microlearning modules for faculty and residents, developed an infographic for broader dissemination, and are pursuing electronic health record integration initiatives, including revised templates and automated flags that suggest alternatives without adding to faculty workload. Their data suggest the cultural groundwork is already laid: clinicians overwhelmingly agree that language matters. What remains is to build the skills, the feedback systems, and the technological safeguards that turn that agreement into notes that inform without harming, and that treat every patient reading their own record with the respect the profession claims to owe them.</p>
<p><strong>Subject of Research:</strong> Clinician knowledge and practices regarding stigmatizing language in clinical documentation</p>
<p><strong>Article Title:</strong> Hospital-Based Clinicians’ Knowledge and Practices Regarding Stigmatizing Language in Clinical Documentation: A Multicenter Survey</p>
<p><strong>Article References:</strong> Caton, J., Sun, B., Steele, N., Santiago, C., Antara, F., Hom, J., &amp; Dougherty, R. (2026). Hospital-Based Clinicians’ Knowledge and Practices Regarding Stigmatizing Language in Clinical Documentation: A Multicenter Survey. <em>Journal of General Internal Medicine</em>. <a href="https://doi.org/10.1007/s11606-026-10803-x" rel="noopener noreferrer">https://doi.org/10.1007/s11606-026-10803-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11606-026-10803-x" rel="noopener noreferrer">10.1007/s11606-026-10803-x</a></p>
<p><strong>Keywords:</strong> stigmatizing language, clinical documentation, electronic health records, medical education, patient-centered language, healthcare disparities, open notes, internal medicine, medical bias, clinical notes, faculty development, survey research</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">245481</post-id>	</item>
		<item>
		<title>AI Reads Clinical Notes to Predict Recovery After Cardiac Arrest</title>
		<link>https://scienmag.com/ai-reads-clinical-notes-to-predict-recovery-after-cardiac-arrest-2/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 01:48:29 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI in neurocritical care]]></category>
		<category><![CDATA[bedside clinical note interpretation]]></category>
		<category><![CDATA[brain injury outcome prediction]]></category>
		<category><![CDATA[cardiac arrest]]></category>
		<category><![CDATA[cardiac arrest prognosis]]></category>
		<category><![CDATA[clinical notes]]></category>
		<category><![CDATA[clinical notes analysis]]></category>
		<category><![CDATA[coma]]></category>
		<category><![CDATA[comatose patient prognosis]]></category>
		<category><![CDATA[critical care]]></category>
		<category><![CDATA[EEG]]></category>
		<category><![CDATA[electronic health record data]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[healthcare bias]]></category>
		<category><![CDATA[ICU decision-making biases]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in critical care]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing in medicine]]></category>
		<category><![CDATA[neurocritical care predictions]]></category>
		<category><![CDATA[neuroprognostication]]></category>
		<category><![CDATA[predictive models]]></category>
		<category><![CDATA[recovery prediction after cardiac arrest]]></category>
		<category><![CDATA[withdrawal of life-sustaining treatment]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224978</guid>

					<description><![CDATA[A natural language processing model trained on clinical notes predicts neurologic outcome in comatose cardiac arrest survivors while exposing hidden patterns and possible biases in prognostic decision-making.]]></description>
										<content:encoded><![CDATA[<p>Few moments in medicine carry the weight of a prognosis delivered after cardiac arrest. When a patient survives the initial collapse but remains comatose, families must often decide within days whether to continue life-sustaining treatment, and those decisions rest on an imperfect synthesis of neurologic examinations, electroencephalograms, brain imaging, blood biomarkers, and the accumulated bedside judgment of clinicians. Much of that reasoning never appears in any structured field of the electronic health record. It lives in the free-text clinical notes: the neurologist&#8217;s concern, the intensivist&#8217;s uncertainty, the nurse&#8217;s overnight observations, the record of a difficult family meeting. A new study published in Neurocritical Care asks whether machine learning can read between those lines, and whether the words clinicians write down might reveal not only a patient&#8217;s likely outcome but also the hidden logic, and possible biases, of prognostic decision-making itself.</p>
<p>The research, led by Furlow and colleagues, tackles one of the most consequential prediction problems in critical care. Up to 80 percent of patients who regain spontaneous circulation after cardiac arrest are comatose on admission, and their recovery is frequently uncertain. Because brain injury is the most common cause of death and disability in this population, accurate neuroprognostication is essential, yet reliable predictors remain limited and many apply only to narrow subsets of patients. Current guidelines emphasize a multimodal approach, combining examination findings, EEG patterns, imaging, and biomarkers rather than relying on any single signal. Even so, the process remains as much art as science, and the reasoning behind a withdrawal-of-care decision is often scattered across hundreds of narrative notes.</p>
<p>To build their model, the researchers extracted unstructured text from notes spanning the entire hospitalization of comatose cardiac arrest survivors treated at three hospitals. Natural language processing algorithms processed this text and identified potentially informative keywords, which were then used to train a logistic regression model to predict outcomes on a stratified Cerebral Performance Category scale, the standard framework for grading neurologic recovery after cardiac arrest. The baseline model performed well, achieving an area under the receiver operating characteristic curve of 0.9 and an area under the precision-recall curve of 0.89. Those figures place the note-based model in the range of respectable clinical prediction tools, though the authors caution that performance fell when the model was restricted to notes from only the first 24 hours of admission, with the decline especially pronounced for predicting good outcomes.</p>
<p>That last detail matters enormously, and it is where the accompanying commentary by Sarah Nutman and Christopher Horvat of the University of Pittsburgh sharpens the discussion. Post-cardiac arrest guidelines from the American Heart Association, the European Resuscitation Council and European Society of Intensive Care Medicine, and the Neurocritical Care Society recommend waiting at least 72 hours before attempting to prognosticate, because sedation, targeted temperature management, and the natural evolution of injury can cloud early assessment. A model trained on notes from the first day of admission is therefore being asked to make a judgment that guidelines explicitly say should not yet be made. Nutman and Horvat suggest that the drop in performance on early notes is unsurprising, and that a 72-hour window might have been more instructive for evaluating the model&#8217;s true clinical utility.</p>
<p>The commentary also raises a subtler statistical concern: temporal leakage. When a model is trained on notes drawn from the entire hospitalization, the patient&#8217;s eventual outcome can become progressively clearer as time passes, so the text may simply echo a conclusion that clinicians have already reached rather than genuinely predicting it. The authors of the original study worried that their models&#8217; weaker performance on good outcomes creates a risk of self-fulfilling prophecies, in which a pessimistic algorithmic prediction influences the decision to withdraw care, thereby making the prediction come true. Nutman and Horvat propose an alternative interpretation: the asymmetry may reflect clinicians incorporating information that never makes it into the written record, leaving the model blind to the very signals that would allow it to recognize patients on a path to recovery.</p>
<p>Perhaps the most provocative contribution of the work is not the AUROC at all, but the window it opens into the hidden architecture of prognostic decision-making. By identifying language patterns associated with withdrawal of life-sustaining treatment, the researchers demonstrated that natural language processing can do more than forecast outcomes; it can interrogate the clinical reasoning, and the possible biases, embedded in routine documentation. Two findings stand out. Keywords related to intimate partner violence were significantly more common in the notes of patients on whom life-sustaining treatment was withdrawn. And although several guidelines recommend against using myoclonus as a prognostic marker, myoclonus appeared significantly more often in the records of patients whose care was withdrawn, raising the possibility that the finding is still quietly influencing bedside decisions despite the guideline warnings.</p>
<p>These observations point toward a use case that extends well beyond prediction. Even without a predictive model, the approach suggests that natural language processing could be deployed to systematically audit differences in clinical outcomes using variables that structured data never captures. In an era when health systems are increasingly attentive to equity, the ability to mine the narrative record for evidence of inconsistent decision-making represents a genuinely novel form of quality control. The same technique that flags a guideline-discordant reliance on myoclonus could, in principle, reveal whether prognostic pessimism is distributed unevenly across patient populations, institutions, or clinical teams.</p>
<p>The study also highlights an underappreciated dimension of prognostication: whose observations count. Neuroprognostication has historically been driven by physician perspectives, yet nurses, therapists, and other health professionals spend far more continuous time at the bedside and often notice changes, in spontaneous movements, responsiveness to family, or sleep-wake cycling, that physicians document less frequently. Incorporating these complementary observations through interdisciplinary evaluation may improve predictions, and the note-mining approach provides a structured mechanism to capture and integrate perspectives from across the clinical team. In that sense, the model is not merely reading the chart; it is aggregating the collective perception of everyone who wrote in it.</p>
<p>Still, significant questions remain before such an algorithm could approach the bedside. Post-arrest guidelines insist on a multimodal approach rather than reliance on any single predictor or algorithm, so the critical unknown is the additive benefit of this tool when used alongside established predictors, including the clinical examination, EEG and neuroimaging findings, and serum biomarkers. The original study compared its model&#8217;s performance to other natural language processing models, but not to known neuroprognostic predictors, which limits how much can be said about its incremental value. It also remains unknown whether the algorithm captures anything genuinely new, or whether it simply rediscovers, in narrative form, the same information clinicians already extract from structured testing. A model trained on clinical notes may be learning patient biology, clinician interpretation, family preferences, institutional practice patterns, or an inseparable mixture of all four, and in post-arrest care that distinction is not academic.</p>
<p>The stakes of getting this right could hardly be higher. A falsely pessimistic prediction does not merely misclassify a patient; it can trigger decisions that make the prediction come true, converting an algorithmic error into an irreversible outcome. As Nutman and Horvat conclude, the next step is not to build better note-based models but to determine what these models are actually learning, when their predictions become reliable, and how they should be used alongside established multimodal prognostic tools. Used carefully, natural language processing may help clinicians extract real signals from the clinical record and strengthen one of the most difficult judgments in medicine. Used carelessly, it risks transforming documentation artifacts into clinical authority. The enduring promise of this work lies in making the invisible parts of prognostication visible, and then deciding, with appropriate humility, which of those signals deserve to guide care.</p>
<p><strong>Subject of Research:</strong> Natural language processing of clinical notes for neurologic outcome prediction after cardiac arrest</p>
<p><strong>Article Title:</strong> Reading Between the Lines After Cardiac Arrest</p>
<p><strong>Article References:</strong> Nutman, S. K., &amp; Horvat, C. M. (2026). Reading Between the Lines After Cardiac Arrest. <em>Neurocritical Care</em>. <a href="https://doi.org/10.1007/s12028-026-02649-2" rel="noopener noreferrer">https://doi.org/10.1007/s12028-026-02649-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12028-026-02649-2" rel="noopener noreferrer">10.1007/s12028-026-02649-2</a></p>
<p><strong>Keywords:</strong> cardiac arrest, neuroprognostication, natural language processing, machine learning, electronic health records, clinical notes, coma, withdrawal of life-sustaining treatment, EEG, critical care, predictive models, healthcare bias</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224978</post-id>	</item>
		<item>
		<title>AI Reads Clinical Notes to Predict Recovery After Cardiac Arrest</title>
		<link>https://scienmag.com/ai-reads-clinical-notes-to-predict-recovery-after-cardiac-arrest/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:21:54 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI-based clinical note analysis]]></category>
		<category><![CDATA[AI-driven assessment of neurological recovery]]></category>
		<category><![CDATA[automated prediction of cerebral performance]]></category>
		<category><![CDATA[cardiac arrest]]></category>
		<category><![CDATA[cerebral performance category]]></category>
		<category><![CDATA[clinical documentation analysis for patient outcomes]]></category>
		<category><![CDATA[clinical notes]]></category>
		<category><![CDATA[coma]]></category>
		<category><![CDATA[deep learning in neurocritical care]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[logistic regression]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for neurological prognosis]]></category>
		<category><![CDATA[medical text mining for brain injury]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing in medicine]]></category>
		<category><![CDATA[neural network applications in healthcare]]></category>
		<category><![CDATA[neurocritical care]]></category>
		<category><![CDATA[neurological outcome]]></category>
		<category><![CDATA[predicting recovery after cardiac arrest]]></category>
		<category><![CDATA[predictive medicine]]></category>
		<category><![CDATA[prognostic modeling for comatose patients]]></category>
		<category><![CDATA[prognostication]]></category>
		<category><![CDATA[unstructured medical record data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224286</guid>

					<description><![CDATA[Researchers trained a natural language processing model to automatically derive neurological outcome scores for comatose cardiac arrest patients from clinical notes, achieving 81 percent accuracy using the full hospital record.]]></description>
										<content:encoded><![CDATA[<p>When a patient survives a cardiac arrest but remains comatose, one of the hardest questions in medicine follows within days: will the brain recover? Clinicians currently answer it with a painstaking blend of neurological exams, EEG recordings, brain imaging, and biomarkers, and even then the answer is often uncertain. A new study published in Neurocritical Care offers a strikingly different approach. Instead of relying on structured test results alone, a team of researchers from the University of California, San Francisco, UC Berkeley, Massachusetts General Hospital, and Beth Israel Deaconess Medical Center trained a machine learning model to read the clinical notes that doctors, nurses, and consultants write throughout a hospitalization, and to derive from that unstructured text the patient&#8217;s neurological outcome at discharge. The results suggest that the everyday prose of the medical record carries a powerful prognostic signal, one that algorithms can extract with an accuracy approaching that of expert human review.</p>
<p>The outcome measure at the heart of the study is the Cerebral Performance Category, or CPC, a five-level scale that neurointensivists use to summarize how well a patient&#8217;s brain is functioning. Categories 1 through 3 denote good outcomes, ranging from a return to normal life or moderate disability, while categories 4 and 5 denote poor outcomes, spanning severe disability, a vegetative state, or death. Assigning a CPC score requires a trained human to synthesize days or weeks of clinical information, which makes it labor-intensive and difficult to standardize across hospitals. That bottleneck matters because large-scale research on cardiac arrest prognostication, including multi-center trials and quality-improvement programs, depends on consistently labeled outcomes for thousands of patients. If an algorithm could assign these categories automatically, researchers could analyze far larger cohorts and clinicians could track outcomes in near real time.</p>
<p>To test that idea, the team conducted a retrospective cohort study of adult patients who were comatose after either in-hospital or out-of-hospital cardiac arrest at three academic hospitals in the United States. The dataset comprised 357 patients in total, split into a training set of 249 patients on which the model learned, and a holdout set of 108 patients on which its performance was evaluated. The raw material was the full corpus of clinical notes in each patient&#8217;s electronic health record, from admission through discharge. Before any modeling, the text was de-identified to strip out protected health information, a critical step given the privacy stakes of mining free-text records. The pipeline then transformed the notes into numerical features and fed them into a logistic regression classifier, a relatively simple and transparent model chosen deliberately over more opaque deep learning architectures so that the researchers could inspect what the model was actually learning.</p>
<p>The performance figures are the study&#8217;s headline. Using all clinical notes across the entire hospitalization, the model achieved an overall accuracy of 81 percent in distinguishing good from poor neurological outcomes, with an area under the receiver operating characteristic curve of 0.90 and an area under the precision-recall curve of 0.89. In practical terms, an AUROC of 0.90 means that if you picked one patient with a good outcome and one with a poor outcome at random, the model would assign a higher risk score to the poorer-outcome patient nine times out of ten. The AUPRC figure is particularly meaningful in this setting because good and poor outcomes are not evenly balanced in the cohort; precision-recall performance is less inflated by class imbalance and therefore a stricter test of real-world utility. For a model built from nothing but free text, these numbers place it in the same performance territory as many prognostic tools built on carefully curated physiological data.</p>
<p>Perhaps the most provocative finding emerged when the researchers restricted the model&#8217;s input to notes documented only within the first 24 hours after hospitalization. Accuracy dipped to 74 percent, with the AUROC and AUPRC both falling to 0.83, but the model still performed well above chance. That means substantial prognostic information about a patient&#8217;s eventual neurological outcome is embedded in the very first day of clinical documentation, long before the outcome is known. The authors interpret this signal as a mixture of three ingredients: the patient&#8217;s underlying biology, which manifests in early exam findings and physiological derangements; the clinician&#8217;s structured assessment of those findings; and potentially a third, more troubling component, early prognostic framing, in which clinicians&#8217; initial expectations about recovery color the language they use and the care they document.</p>
<p>That third component connects to one of the most debated issues in neurocritical care: the role of clinician bias and self-fulfilling prophecy in cardiac arrest prognostication. Prior research has shown that early withdrawal of life-sustaining therapy is common after cardiac arrest and may result in deaths that would not otherwise occur, and that providers can be overconfident in their early outcome predictions. The new study found that the model demonstrated higher precision in predicting poor neurological outcomes particularly among patients who underwent withdrawal of life-sustaining therapy. In other words, the model was especially good at detecting poor outcomes in precisely the group where the outcome may have been shaped, at least in part, by the decision to withdraw care. This raises a subtle question about what the model is truly measuring: the patient&#8217;s intrinsic recovery potential, or the trajectory that clinical decision-making set in motion. The authors are careful on this point, and their framing of early prognostic signals as a blend of biology and clinician assessment acknowledges the entanglement rather than claiming the model has solved it.</p>
<p>Technically, the choice of logistic regression over a large language model is worth unpacking. Modern clinical NLP increasingly relies on transformer-based models that can capture context and nuance in text, but they are harder to interpret and validate in high-stakes medical settings. By using a simpler model on engineered text features, the team could examine which words and phrases drove predictions, an essential property when the goal is to understand what clinical documentation reveals about prognosis. The researchers have also made the entire NLP pipeline, including preprocessing, modeling, and evaluation code, publicly available on GitHub, which lowers the barrier for other groups to replicate the approach on their own patient populations and to scrutinize the method for hidden artifacts. Reproducibility of this kind is rare and valuable in clinical machine learning, where models often fail to generalize when moved between hospitals with different documentation practices.</p>
<p>The study is candid about its limitations and about the work that remains. The cohort of 357 patients from three academic centers is modest by machine learning standards, and the model&#8217;s performance will need validation in external, more diverse populations before any clinical deployment. The authors note that future research should incorporate multimodal data, combining text with EEG, imaging, and laboratory values, and should employ interpretable advanced language models to push accuracy higher toward fully automatable CPC derivation. There is also the question of what such a tool should be used for. The authors position it primarily as a research instrument, a way to generate consistently labeled outcome data at scale for prognostication studies and quality improvement, rather than as a bedside oracle that would tell families whether a loved one will wake up. Given the documented dangers of premature prognostication and early withdrawal of support, that restraint seems well judged.</p>
<p>Still, the broader implications are considerable. Cardiac arrest affects hundreds of thousands of people each year in the United States alone, and neurological outcome remains the dominant determinant of long-term quality of life among survivors. The American Heart Association has issued formal standards for how prognostication studies should be conducted, precisely because the field is littered with overconfident early predictions. An automated system that can derive standardized outcome categories from routine documentation could accelerate the search for better prognostic models, enable continuous auditing of how outcomes vary across hospitals and patient groups, and eventually support clinicians with a second, data-driven opinion grounded in the full record rather than a single exam. The finding that the first 24 hours of notes already carry most of the signal is both an opportunity, for earlier and better-informed decision-making, and a warning, that the language clinicians write on day one may be quietly shaping the outcomes they later record. Either way, the study makes a compelling case that the medical record&#8217;s unstructured text is not just documentation of care, but a rich, largely untapped data source for understanding and predicting recovery of the injured brain.</p>
<p><strong>Subject of Research:</strong> Automated prediction of neurological outcomes after cardiac arrest using natural language processing of electronic health record notes</p>
<p><strong>Article Title:</strong> Automated Derivation of Cerebral Performance Category at Hospital Discharge After Cardiac Arrest Using Natural Language Processing and Machine Learning</p>
<p><strong>Article References:</strong> Automated Derivation of Cerebral Performance Category at Hospital Discharge After Cardiac Arrest Using Natural Language Processing and Machine Learning. (n.d.). <a href="https://doi.org/10.1007/s12028-026-02650-9" rel="noopener noreferrer">https://doi.org/10.1007/s12028-026-02650-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12028-026-02650-9" rel="noopener noreferrer">10.1007/s12028-026-02650-9</a></p>
<p><strong>Keywords:</strong> cardiac arrest, natural language processing, machine learning, cerebral performance category, electronic health records, neurocritical care, prognostication, clinical notes, logistic regression, coma, neurological outcome, predictive medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224286</post-id>	</item>
		<item>
		<title>AI Pipeline Turns Messy Clinical Notes Into Research-Ready Data With Near-Perfect Accuracy</title>
		<link>https://scienmag.com/ai-pipeline-turns-messy-clinical-notes-into-research-ready-data-with-near-perfect-accuracy/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:12:50 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI-powered clinical note extraction]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[automated chart review]]></category>
		<category><![CDATA[CLASS pipeline]]></category>
		<category><![CDATA[clinical documentation analysis]]></category>
		<category><![CDATA[clinical notes]]></category>
		<category><![CDATA[Clinical Research]]></category>
		<category><![CDATA[data extraction]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[esophageal airway treatment surgery]]></category>
		<category><![CDATA[healthcare data accuracy]]></category>
		<category><![CDATA[hospital informatics solutions]]></category>
		<category><![CDATA[Johns Hopkins]]></category>
		<category><![CDATA[Johns Hopkins AI healthcare project]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models for medical data]]></category>
		<category><![CDATA[medical informatics]]></category>
		<category><![CDATA[medical record data structuring]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[natural language processing in healthcare]]></category>
		<category><![CDATA[pediatric surgery]]></category>
		<category><![CDATA[privacy-preserving medical AI]]></category>
		<category><![CDATA[research-ready electronic health records]]></category>
		<category><![CDATA[secure medical data processing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201580</guid>

					<description><![CDATA[Researchers at Johns Hopkins All Children's Hospital developed CLASS, a large language model pipeline that extracted pediatric surgical procedure data from unstructured clinical notes with near-perfect concordance to expert review.]]></description>
										<content:encoded><![CDATA[<p>Electronic medical records are often described as gold mines of clinical information, but most of that treasure is buried. While laboratory values, vital signs, and medication orders arrive neatly coded, the richest details of a patient&#8217;s story—the surgical nuances, the procedural variations, the clinical judgment calls—live inside free-text notes written by busy clinicians. For researchers and quality-improvement teams, extracting that information has long meant hours of painstaking manual chart review, a process that is slow, expensive, and nearly impossible to scale. Now, a team at Johns Hopkins All Children&#8217;s Hospital has shown that a carefully engineered large language model pipeline can do much of that work automatically, with accuracy so high that it approaches the ceiling of human agreement.</p>
<p>The system, called CLASS—short for Clinical LLM Abstraction &amp; Structuring System—is described in a new feasibility report published in the Journal of Medical Systems. Led by anesthesiologist and informatics researcher Frederick H. Kuo, the team built CLASS as a modular, Python-based pipeline that runs entirely within a secure institutional computing environment, keeping protected health information behind the hospital&#8217;s own walls rather than sending it to external services. That design choice reflects a growing consensus in medical informatics: the power of large language models can be harnessed for clinical data without compromising patient privacy, provided the infrastructure is configured correctly.</p>
<p>Technically, CLASS rests on three pillars. The first is a set of concept lists curated by subject matter experts—in this case, surgeons and informatics specialists who defined exactly which procedures and clinical details the system should look for. The second is a task-specific prompt suite, a collection of carefully worded instructions that steers the large language model toward consistent, clinically grounded interpretations of each note. The third is a schema-constrained output format, which forces the model to return its findings in a structured, predictable structure rather than free-flowing prose. The results are then exported to an interactive dashboard built for expert review, allowing clinicians to verify outputs, spot errors, and analyze the extracted data at scale.</p>
<p>One of the most innovative features of CLASS is its handling of the unknown. Rather than simply classifying notes against a fixed list of predefined concepts, the pipeline actively flags potential variants or entirely novel concepts that do not fit the existing schema, surfacing them for expert consideration. In specialized fields where terminology evolves quickly and procedures are often described in non-standard ways, this ability to propose expansions to the concept vocabulary could fundamentally change how clinical registries and research databases are built and maintained.</p>
<p>To test the system, the researchers applied CLASS to a retrospective corpus of pediatric esophageal airway treatment surgery, or EATS, operative notes at their single center. EATS is a demanding test case: these operative notes describe complex, highly individualized procedures in children with airway and esophageal abnormalities, and much of the procedural detail is not captured in standard billing codes. The team compared CLASS outputs against adjudication by an experienced surgeon on the twenty longest notes in the corpus, yielding 3,960 individual note-procedure pairs for evaluation.</p>
<p>The results were striking. Across those thousands of judgments, the observed concordance between the automated pipeline and the surgeon&#8217;s adjudication reached an F1 score of 0.9967—a near-perfect measure of precision and recall combined. In practical terms, the model almost never missed a procedure the surgeon identified, and almost never invented one that was not there. For a task as subtle as parsing operative prose about pediatric airway surgery, that level of agreement suggests that large language models, when properly constrained and prompted, can match expert-level abstraction performance in at least some specialized clinical domains.</p>
<p>The novelty-detection results were more nuanced, and arguably more interesting. CLASS proposed 28 candidate procedure variants or additions to the curated concept list, and the clinical team judged 18 of them—64.3 percent—to be genuinely useful. That is a meaningful yield: nearly two out of every three suggestions from the machine were worth a clinician&#8217;s time. At the same time, the surgeon identified 12 additional procedures that CLASS failed to surface, a reminder that the system works best as a collaborator rather than a replacement. The human expert still caught things the machine missed, and the machine still surfaced things the human might not have thought to codify.</p>
<p>The authors are careful to frame the study appropriately. This was an exploratory implementation at a single center, focused on a single surgical service, with evaluation limited to a small set of long notes. Generalizability to other institutions, other note types, and other clinical tasks remains unproven, and the team emphasizes that broader validation is needed before such pipelines could be trusted for high-stakes applications. The full production code and clinical data cannot be released publicly because they involve protected health information and institution-specific infrastructure, but the researchers have shared technical implementation details, template code, pseudocode, and de-identified prompt examples in the online supplementary materials, giving other informatics teams a practical roadmap for building similar systems.</p>
<p>Even with those caveats, the implications are considerable. Clinical research has long been throttled by the bottleneck of manual abstraction: cohort studies that could enroll thousands of patients are often limited to hundreds simply because there are only so many hours in a research coordinator&#8217;s day. Quality-improvement programs face the same constraint, unable to measure surgical outcomes comprehensively when the relevant data must be pulled by hand from narrative notes. If pipelines like CLASS can reliably convert unstructured text into analyzable data within standard institutional infrastructure, the effective sample sizes of clinical research could expand dramatically, and hospitals could monitor the quality of specialized care in near real time.</p>
<p>The study also highlights a shift in how medical informatics teams may work in the coming years. Instead of writing brittle rule-based extraction algorithms or training bespoke machine-learning models on small labeled datasets, teams can now curate expert concept lists, design prompts, and review machine-generated suggestions—a workflow in which clinicians define what matters and the model handles the linguistic heavy lifting. The CLASS experience suggests this human-machine partnership can work: the model performs the extraction with near-perfect fidelity, proposes useful vocabulary expansions, and leaves final judgment to the experts who bear clinical responsibility. As large language models continue to demonstrate their grasp of medical language, studies like this one offer a concrete, privacy-conscious template for turning the narrative richness of the medical record into structured knowledge—without a single chart being pulled by hand.</p>
<p><strong>Subject of Research:</strong> A large language model pipeline for extracting structured data from unstructured clinical notes</p>
<p><strong>Article Title:</strong> Exploratory Implementation and Feasibility Report of CLASS (Clinical LLM Abstraction &amp; Structuring System), A Large Language Model Pipeline for Extracting Unstructured Data From Clinical Notes</p>
<p><strong>Article References:</strong> Exploratory Implementation and Feasibility Report of CLASS (Clinical LLM Abstraction &amp; Structuring System), A Large Language Model Pipeline for Extracting Unstructured Data From Clinical Notes. (n.d.). <a href="https://doi.org/10.1007/s10916-026-02462-6" rel="noopener noreferrer">https://doi.org/10.1007/s10916-026-02462-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10916-026-02462-6" rel="noopener noreferrer">10.1007/s10916-026-02462-6</a></p>
<p><strong>Keywords:</strong> large language models, clinical notes, natural language processing, electronic health records, medical informatics, data extraction, pediatric surgery, artificial intelligence, CLASS pipeline, clinical research, esophageal airway treatment surgery, Johns Hopkins</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201580</post-id>	</item>
	</channel>
</rss>
