<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>automation of root cause analysis in medicine &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/automation-of-root-cause-analysis-in-medicine/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 08:02:06 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>automation of root cause analysis in medicine &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Joins the Safety Team: Language Models Tackle Root Cause Analysis in Radiation Oncology</title>
		<link>https://scienmag.com/ai-joins-the-safety-team-language-models-tackle-root-cause-analysis-in-radiation-oncology/</link>
		
		<dc:creator><![CDATA[Skylar Underwood]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 08:02:06 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AAPM guidelines]]></category>
		<category><![CDATA[AI in radiation oncology incident analysis]]></category>
		<category><![CDATA[AI-assisted medical error investigation]]></category>
		<category><![CDATA[AI-powered causal reasoning in healthcare]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[artificial intelligence in radiation therapy]]></category>
		<category><![CDATA[automation of root cause analysis in medicine]]></category>
		<category><![CDATA[digital health innovations in radiation oncology]]></category>
		<category><![CDATA[enhancing radiation therapy safety with AI]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[incident learning]]></category>
		<category><![CDATA[incident report analysis using AI]]></category>
		<category><![CDATA[language models for healthcare safety]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[machine learning for incident investigation]]></category>
		<category><![CDATA[medical physics]]></category>
		<category><![CDATA[patient safety]]></category>
		<category><![CDATA[PLOS Digital Health]]></category>
		<category><![CDATA[quality improvement]]></category>
		<category><![CDATA[radiation oncology]]></category>
		<category><![CDATA[radiation oncology safety systems]]></category>
		<category><![CDATA[RO-ILS]]></category>
		<category><![CDATA[root cause analysis]]></category>
		<category><![CDATA[root cause analysis in medical radiation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=252669</guid>

					<description><![CDATA[A new proof-of-concept study shows that large language models can perform credible root cause analyses of radiation oncology incident reports, though hallucination rates mean human experts must remain firmly in the loop.]]></description>
										<content:encoded><![CDATA[<p>Radiation oncology is one of the most technologically demanding disciplines in modern medicine. Linear accelerators, treatment planning systems, imaging pipelines and immobilization devices must all work in flawless concert to deliver precisely sculpted doses of ionizing radiation to tumors while sparing healthy tissue. When something goes wrong in this chain, the consequences can be serious, which is why the field has built elaborate systems for reporting and analyzing incidents. Yet the analysis itself, known as root cause analysis or RCA, remains a labor-intensive, expert-driven process that often stalls under the weight of its own paperwork. A new proof-of-concept study published in PLOS Digital Health suggests that large language models, the same class of artificial intelligence systems that power modern chatbots, may be able to shoulder part of that burden, performing structured causal reasoning on real incident reports with a degree of accuracy that surprised even the medical physicists who evaluated them.</p>
<p>The research team, led by Yuntao Wang and colleagues including Mariluz De Ornelas, Matthew T. Studenski, Elizabeth Bossart, Siamak P. Nejad-Davarani and Yunze Yang, drew its test material from the Radiation Oncology Incident Learning System, or RO-ILS, a national reporting database that collects narrative accounts of errors and near-misses from clinics across the country. Rather than feeding the models raw, unstructured text and hoping for the best, the investigators designed a carefully standardized prompt built on the root cause analysis guidelines of the American Association of Physicists in Medicine. Four state-of-the-art models were put through their paces: Gemini 2.5 Pro, GPT-4o, o3 and Grok 3. Each received the Background and Incident Overview sections from nineteen publicly available RO-ILS cases and was instructed to produce three deliverables for every case: the root causes of the incident, the lessons that should be learned from it, and a set of suggested corrective actions.</p>
<p>What makes this study methodologically interesting is the sheer ambition of its evaluation framework. Assessing the quality of an AI-generated causal analysis is not like checking arithmetic; there is no single correct answer key. The team therefore layered three distinct tiers of assessment on top of one another. At the objective level, they used semantic similarity metrics, computing cosine similarity between model outputs and reference analyses with a Sentence Transformer model, a technique that measures how closely two pieces of text mean the same thing even when they use different words. At a semi-subjective level, they calculated precision, recall and F1-scores, along with an expert-adjudicated positive predictive value, a hallucination rate, and performance criteria covering relevance, comprehensiveness, quality of justification and quality of the proposed solutions. Finally, five board-certified medical physicists provided subjective ratings of reasoning quality and overall performance, bringing genuine clinical judgment to bear on every output.</p>
<p>The headline finding is that the models performed satisfactorily across the board. All four demonstrated comparable baseline capabilities in extracting causal factors from incident narratives, and the analyses they produced were judged relevant and accurate, aligned with what the expert reviewers expected from a competent human analyst. Gemini 2.5 Pro emerged with the highest overall performance score among the four systems, though the differences between models were not uniform across every metric. Statistical testing revealed significant differences among the models in expert-adjudicated positive predictive value, in hallucination rate, and in the subjective ratings assigned by the physicist reviewers, with p-values below 0.05. In plain terms, the models were not interchangeable: some were noticeably more trustworthy than others when their outputs were scrutinized by people who do this work for a living.</p>
<p>That last point matters enormously, because the study did not shy away from the technology&#8217;s most notorious weakness. Every model exhibited some degree of hallucination, meaning it fabricated or distorted information not supported by the source material, and the rates ranged from a comparatively modest 11 percent to a troubling 61 percent. For a field where patient safety hangs in the balance, a hallucination rate anywhere near the upper end of that spectrum would be disqualifying for unsupervised use. The authors are careful to frame their results as evidence for assistive rather than autonomous deployment. The vision they describe is not a machine that replaces the safety committee but a machine that drafts the first pass of an analysis, surfaces plausible causal threads, and proposes candidate corrective actions that human experts can then verify, refine and own.</p>
<p>The implications for clinical practice could be substantial. Root cause analysis in radiation oncology typically requires convening multidisciplinary teams of physicists, dosimetrists, therapists, physicians and administrators to reconstruct what happened and why. These sessions are time-consuming, and the narrative reports that feed them are often long, ambiguous and written under stressful circumstances. A language model that can rapidly generate a structured preliminary analysis, organized according to established AAPM guidelines, could shorten the path from incident report to corrective action. It could also help smaller clinics that lack the staffing to conduct exhaustive analyses of every near-miss, potentially raising the baseline of safety surveillance across the entire field. In the aggregate, tools like this could strengthen incident learning systems by making the analytical step faster and more consistent.</p>
<p>The study also offers a template for how medical AI evaluations should be conducted more broadly. Rather than relying on a single metric or a single reviewer, the researchers triangulated across objective similarity measures, quantitative classification metrics and human expert judgment, and they applied statistical significance testing to distinguish genuine performance differences from noise. This layered approach acknowledges an uncomfortable truth about language models: text that looks plausible is not necessarily text that is correct, and similarity to a reference answer does not capture whether a proposed corrective action would actually work in a clinic. By combining expert-adjudicated precision with explicit hallucination measurement, the evaluation gets closer to the question that really matters for patient safety: can this system&#8217;s output be trusted, and under what supervision?</p>
<p>Caution remains warranted. Nineteen publicly available cases are a small sample, and publicly reported incidents may differ in complexity and completeness from the internal reports a clinic generates behind its own walls. The models were tested on narrative sections only, without access to the full investigative context that a real RCA team would possess. Hallucination rates, even at the low end, mean that every machine-generated statement about an incident would need verification before it informed any corrective decision. There are also governance questions the study does not resolve: how patient confidentiality would be protected when incident narratives are processed by commercial models, who bears responsibility for an AI-suggested action that proves inadequate, and how such tools would be validated and regulated before entering routine safety workflows.</p>
<p>Still, the direction of travel is clear and, for a field built on the principle of learning from error, quietly exciting. Radiation oncology was among the first medical specialties to confront the reality that complex technology fails in complex ways, and its incident learning infrastructure is among the most mature in healthcare. Injecting language models into that infrastructure, as assistive analysts that draft, summarize and propose while humans verify and decide, could compress the feedback loop between error and improvement from weeks to hours. The PLOS Digital Health study is a proof of concept, not a prescription, but it demonstrates that the reasoning capabilities of current frontier models are already close enough to expert expectations to be worth taking seriously. As hallucination rates fall and evaluation standards mature, the safety committee&#8217;s newest member may well be an algorithm, one that never gets tired of reading incident reports and never forgets a lesson once it has been written down.</p>
<p><strong>Subject of Research:</strong> Large language model-based root cause analysis of radiation oncology patient safety incidents</p>
<p><strong>Article Title:</strong> Augmenting patient safety surveillance in radiation oncology with large language model-based root cause analysis: A proof-of-concept study</p>
<p><strong>Article References:</strong> Wang, Y., De Ornelas, M., Studenski, M. T., Bossart, E., Nejad-Davarani, S. P., &amp; Yang, Y. (2026). Augmenting patient safety surveillance in radiation oncology with large language model-based root cause analysis: A proof-of-concept study. <em>PLOS Digital Health, 5</em>(9), e0001740. <a href="https://doi.org/10.1371/journal.pdig.0001740" rel="noopener noreferrer">https://doi.org/10.1371/journal.pdig.0001740</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pdig.0001740" rel="noopener noreferrer">10.1371/journal.pdig.0001740</a></p>
<p><strong>Keywords:</strong> large language models, radiation oncology, root cause analysis, patient safety, RO-ILS, incident learning, medical physics, hallucination, AAPM guidelines, artificial intelligence, PLOS Digital Health, quality improvement</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">252669</post-id>	</item>
	</channel>
</rss>
