<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>uncertainty-aware AI models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/uncertainty-aware-ai-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 19:04:21 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>uncertainty-aware AI models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Ultrasound Model Learns When Not to Decide, Deferring Hard Lymph Node Cases</title>
		<link>https://scienmag.com/ai-ultrasound-model-learns-when-not-to-decide-deferring-hard-lymph-node-cases/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 19:04:21 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AI decision abstention]]></category>
		<category><![CDATA[AI in clinical decision-making]]></category>
		<category><![CDATA[AI triage in radiology]]></category>
		<category><![CDATA[AI ultrasound]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[cervical lymph node assessment]]></category>
		<category><![CDATA[cervical lymphadenopathy]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[diagnostic AI]]></category>
		<category><![CDATA[external validation]]></category>
		<category><![CDATA[human-AI collaboration in radiology]]></category>
		<category><![CDATA[lymph node benign versus malignant classification]]></category>
		<category><![CDATA[lymphadenopathy diagnosis]]></category>
		<category><![CDATA[machine learning in head and neck imaging]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[medical imaging AI]]></category>
		<category><![CDATA[multimodal imaging]]></category>
		<category><![CDATA[reader study]]></category>
		<category><![CDATA[selective prediction]]></category>
		<category><![CDATA[triage]]></category>
		<category><![CDATA[ultrasound]]></category>
		<category><![CDATA[ultrasound-based cancer detection]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<category><![CDATA[uncertainty-aware AI models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197676</guid>

					<description><![CDATA[A locked uncertainty-aware AI triage model safely automated about 37 percent of cervical lymph node ultrasound cases while deferring the rest to senior reviewers in multicenter validation.]]></description>
										<content:encoded><![CDATA[<p>Every day, radiologists around the world face the same deceptively simple question: is this enlarged neck lymph node benign or malignant? The stakes could hardly be higher. Cervical lymphadenopathy is one of the most common reasons for head and neck ultrasound, and getting the call wrong in either direction carries real consequences — a missed metastasis can delay cancer treatment by weeks, while an unnecessary biopsy or surgery imposes cost, anxiety and physical risk on a patient who never needed them. Now, a multicenter research team from Fujian, China, has taken a step toward a new kind of artificial intelligence triage tool, one that is designed not only to answer that question but also to know when it should decline to answer at all. The work, published as an open-access article in BMC Medical Imaging, introduces an uncertainty-aware selective routing framework for ultrasound-based assessment of cervical lymph nodes, and its central finding is as candid as it is technically notable: the system safely automated only a minority of cases, and it deliberately deferred the majority to experienced human reviewers.</p>
<p>The study stands out for a methodological choice that most clinical AI deployments still lack. Conventional binary classifiers — the workhorses of medical machine learning — produce a single probability for every case, and a fixed threshold converts that probability into a decision of malignant or benign. Such systems never signal doubt. They output a number even when the input image is ambiguous, the acquisition quality is poor, or the case falls far outside anything they were trained on. The research team, led by first authors Hang Ling, Cailing Lin and Jing Ning, with corresponding author Ziwei Zhang, built their framework around a different premise: that a diagnostic algorithm should be able to abstain. Their selective triage rule combines three ingredients — a fusion prediction model, a composite uncertainty score, and a pre-locked routing policy that sends confident cases down automatic pathways and routes uncertain ones to senior human review.</p>
<p>The technical architecture is worth unpacking. The fusion model integrates three distinct information streams: structured clinical data, features extracted from multimodal ultrasound imaging, and variables mined from structured ultrasound reports. Multimodal ultrasound is an important component here, since modern neck ultrasound is not a single image but a constellation of modalities — grayscale morphological features such as echogenicity and border characteristics, Doppler vascular patterns, and in some settings contrast-enhanced ultrasound dynamics. By fusing these with clinical context and report-derived information, the model aims to approximate the holistic judgment a seasoned sonographer applies, rather than relying on pixel patterns alone.</p>
<p>Just as important is what the researchers did with uncertainty. Rather than trusting the model&#8217;s raw confidence, they constructed a composite uncertainty measure from two complementary signals: predictive entropy, normalized against its distribution in the development cohort, and probability dispersion estimated through bootstrap resampling of the model&#8217;s fits. Entropy captures how peaked or flat the model&#8217;s output distribution is on a given case, while bootstrap dispersion captures how sensitive the prediction is to perturbations in the training data — a proxy for how far the case sits from the model&#8217;s comfort zone. Combining the two, the team fixed an uncertainty threshold of U = 0.700. Critically, every free parameter was locked during development: the binary classification threshold at p = 0.537, the selective-routing probability boundaries at p = 0.320 and p = 0.660, and the uncertainty cutoff. Once locked, the entire pipeline was applied to the internal-validation and two external-validation cohorts without any retuning whatsoever — a design that guards against the subtle overfitting that plagues many retrospective AI studies.</p>
<p>The study population comprised 518 patients, with one index lymph node analyzed per person: 206 in the development cohort, 88 in internal validation, and 112 in each of two independent external cohorts. Discrimination remained remarkably stable across sites, a finding that in itself deserves attention because performance degradation at external sites is the most common failure mode of published clinical prediction models. The area under the receiver operating characteristic curve was 0.839 in internal validation, 0.850 in external validation cohort 1, and 0.842 in external validation cohort 2. At the locked binary threshold, sensitivity and specificity were 0.595 and 0.902 in internal validation, 0.657 and 0.911 in the first external cohort, and 0.817 and 0.808 in the second. Those numbers describe a competent but not extraordinary classifier — which is precisely the point, because the selective framework was engineered to compensate for the model&#8217;s fallibility rather than to hide it.</p>
<p>The heart of the paper lies in its selective-triage results. In the pooled external-validation population of 224 patients, 38 were routed to a lower-risk automatic pathway, 44 to a higher-risk automatic pathway, and 142 — more than sixty percent — were deferred for senior review. Automatic coverage was therefore 0.366, while selective accuracy among the automated cases reached 0.927, with a 95 percent confidence interval of 0.849 to 0.966. The accepted errors within the automated subset were small in absolute terms: two false negatives and four false positives. Among patients funneled into the lower-risk automatic pathway, the negative predictive value was 0.947, meaning the residual probability of malignancy in that group was 5.3 percent — a figure the authors report transparently with a wide confidence interval spanning roughly 1.5 to 17.3 percent. The higher-risk pathway achieved a positive predictive value of 0.909. A risk–coverage analysis, summarized by a partial area under the risk–coverage curve of 0.037 in the pooled external data, quantified how selective accuracy behaved as coverage was expanded or contracted.</p>
<p>The authors are unusually explicit about the limits of these numbers. Only about 37 percent of external cases were eligible for automatic routing, and the accepted false-negative and false-positive counts, while modest, are not zero. Two of the 38 patients automatically assigned to the lower-risk pathway turned out to have malignant nodes. In a separate per-case analysis, the composite uncertainty score was only moderately effective at flagging binary-model errors, achieving an area under the curve of 0.581 for error detection and an area under the precision-recall curve of 0.247 — numbers that indicate real room for improvement in the uncertainty estimation itself. The team concludes that the framework should be interpreted as a retrospective, research-stage selective-routing demonstration rather than an established clinical safety or workflow tool, a disclaimer that rare in a field where press releases routinely outrun the evidence.</p>
<p>Where the system showed its most immediately practical benefit was in the eight-reader study. Using crossed reader-by-case bootstrap resampling — the gold-standard multi-reader multi-case methodology for interpreting diagnostic accuracy studies — the researchers tested whether access to the model&#8217;s output changed reader performance. Junior readers improved their accuracy by 0.055, a statistically significant gain with a confidence interval of 0.022 to 0.089 and a p-value of 0.002. Middle-level and senior readers showed no statistically significant change, a pattern consistent with the intuition that the tool functions as a form of expert guidance for less experienced practitioners while adding little for those who already possess the pattern-recognition skills it encodes. If the framework ultimately translates to the clinic, its clearest value proposition may be compressing the training gap between junior and senior diagnosticians, rather than replacing expert judgment outright.</p>
<p>The broader significance of the study lies in its modeling of what responsible clinical AI could look like. Rather than chasing headline accuracy figures, the researchers built their entire evaluation around the questions that actually matter for deployment: when should an algorithm be allowed to act autonomously, how much residual risk is acceptable within the automated zone, and how much workload is genuinely deferred to humans. The reported numbers answer those questions soberly. Roughly a third of cases could be automated with a combined accuracy above 92 percent, but the framework&#8217;s own honesty mechanisms pushed nearly two-thirds of patients toward senior review, and even the automated pathway retained a nonzero malignancy risk. The team also adhered to modern reporting and risk-of-bias standards, referencing the TRIPOD+AI and PROBAST+AI frameworks, and the study received ethics approval from Fujian Provincial Hospital with center-specific authorization for the external cohorts. Funded by the Natural Science Foundation of Fujian Province, the work offers a template — conservative, externally validated, and uncertainty-aware — for a generation of diagnostic AI systems whose most important capability may be knowing what they do not know.</p>
<p><strong>Subject of Research:</strong> An uncertainty-aware selective artificial intelligence triage model for distinguishing benign from malignant cervical lymphadenopathy on multimodal ultrasound</p>
<p><strong>Article Title:</strong> An uncertainty-aware ultrasound triage model for cervical lymphadenopathy: a retrospective multicenter development and external validation study</p>
<p><strong>Article References:</strong> An uncertainty-aware ultrasound triage model for cervical lymphadenopathy: a retrospective multicenter development and external validation study. (n.d.). <a href="https://doi.org/10.1186/s12880-026-02780-8" rel="noopener noreferrer">https://doi.org/10.1186/s12880-026-02780-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12880-026-02780-8" rel="noopener noreferrer">10.1186/s12880-026-02780-8</a></p>
<p><strong>Keywords:</strong> cervical lymphadenopathy, ultrasound, artificial intelligence, uncertainty quantification, selective prediction, multimodal imaging, external validation, triage, diagnostic AI, reader study, medical imaging, clinical decision support</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197676</post-id>	</item>
		<item>
		<title>New Framework Enhances AI Trustworthiness in Cancer Subtyping</title>
		<link>https://scienmag.com/new-framework-enhances-ai-trustworthiness-in-cancer-subtyping/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Tue, 23 Jun 2026 09:46:23 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[adaptive AI frameworks for healthcare]]></category>
		<category><![CDATA[AI error reduction in pathology]]></category>
		<category><![CDATA[AI generalizability in medical imaging]]></category>
		<category><![CDATA[AI integration in clinical diagnostics]]></category>
		<category><![CDATA[AI trustworthiness in cancer diagnostics]]></category>
		<category><![CDATA[digital pathology AI systems]]></category>
		<category><![CDATA[enhancing AI confidence calibration]]></category>
		<category><![CDATA[improving reliability of AI cancer subtyping]]></category>
		<category><![CDATA[neural network uncertainty detection]]></category>
		<category><![CDATA[TRUECAM AI framework]]></category>
		<category><![CDATA[uncertainty quantification in medical AI]]></category>
		<category><![CDATA[uncertainty-aware AI models]]></category>
		<guid isPermaLink="false">https://scienmag.com/new-framework-enhances-ai-trustworthiness-in-cancer-subtyping/</guid>

					<description><![CDATA[In the rapidly evolving landscape of medical artificial intelligence (AI), one of the most critical hurdles remains uncertainty quantification—a challenge that, if unresolved, can undermine the reliability of AI-driven diagnostics. Traditional artificial neural networks often lack the capacity to adequately recognize when they encounter unfamiliar input data beyond the scope of their training. This limitation [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the rapidly evolving landscape of medical artificial intelligence (AI), one of the most critical hurdles remains uncertainty quantification—a challenge that, if unresolved, can undermine the reliability of AI-driven diagnostics. Traditional artificial neural networks often lack the capacity to adequately recognize when they encounter unfamiliar input data beyond the scope of their training. This limitation can lead to dangerous overconfidence in erroneous outputs. For instance, a model trained exclusively to classify African mammals might mistakenly identify a South American jaguar as a leopard, confident in an answer that falls disastrously short of reality. Addressing this issue is paramount for medical AI, where errors have direct consequences on patient care.</p>
<p>Recently, a breakthrough was reported on June 23 in the prestigious journal Nature Biomedical Engineering. A collaborative team of researchers from Vanderbilt Health and institutions in Hong Kong unveiled a novel AI framework named TRUECAM, an uncertainty-aware wrapper designed to enhance the trustworthiness and generalizability of digital pathology AI systems. Unlike conventional AI models, this wrapper functions as an adaptive interface layer, seamlessly integrating with existing neural networks to better signal when input data is outside the model’s domain or when image quality is insufficient for reliable classification.</p>
<p>TRUECAM’s innovation lies in its dual capability: it not only identifies “out-of-scope” inputs—cases where the AI should prudently refrain from making definitive diagnoses—but also actively filters out noninformative or misleading data segments, such as normal tissue or poorly stained areas that can adversely affect whole-slide image analysis. This filtering is critical because pathology slides often contain heterogeneous regions, and focusing on diagnostically relevant tissue is essential for accurate subtype classification, especially in complex diseases like non-small cell lung cancer (NSCLC).</p>
<p>The research team demonstrated TRUECAM’s effectiveness primarily through NSCLC subtyping, leveraging whole-slide image data sourced from two geographically diverse cancer research consortia. This rigorous testing environment also included a thoughtfully constructed dataset of clinically important “out-of-scope” images, mimicking scenarios that typically challenge AI reliability. Further validation involved real-world images obtained from Queen Mary Hospital in Hong Kong, extending the framework&#8217;s robustness across various clinical settings. Intriguingly, the researchers also tested TRUECAM on cancer tissue from additional organs, including breast, brain, and kidney, underscoring the model’s versatility.</p>
<p>When benchmarked against prevailing digital pathology uncertainty quantification methods, TRUECAM outperformed in several dimensions: accuracy, processing speed, efficiency, and cost-effectiveness. Its streamlined architecture ensures that the enhancement of diagnostic certainty does not come at the expense of increased computational load or resource demands. This balance positions TRUECAM uniquely for clinical adoption, where timely and reliable results are not negotiable.</p>
<p>Professor Bradley Malin, PhD, a noted authority in biomedical informatics and one of the study’s corresponding authors, emphasized the imperative for trustworthy AI within the medical domain. He highlighted the multifaceted sources of variation that impede AI performance—not only the inherent diversity in patient profiles but also institutional variability in specimen collection, staining techniques, and unavoidable tissue preparation artifacts. TRUECAM addresses these variables comprehensively, providing a safeguard against confidently wrong AI conclusions that could otherwise jeopardize patient outcomes.</p>
<p>More than a mere diagnostic enhancer, TRUECAM embodies a paradigm shift by imparting customizable accuracy guarantees for cancer subtype classifications. This opens new pathways where clinicians can specify confidence thresholds tailored to clinical contexts, with AI systems transparently communicating their level of certainty and deferring ambiguous cases to human experts. Such abstention mechanisms are pivotal in integrating AI harmoniously into clinical workflows while maintaining patient safety.</p>
<p>TRUECAM&#8217;s approach to filtering irrelevant tissue regions yields a practical and scientific advantage. Chao Yan, PhD, MS, a lead author of the study, explained how this targeted elimination of “noise” allows the AI to concentrate its analytic power on relatively small, yet diagnostically crucial, patches of pathology slides. This refined focus aligns closely with the attention of human pathologists, who typically assess specific tissue regions for subtype determination, thereby enhancing both the accuracy and fairness of AI-driven pathology assessments.</p>
<p>Furthermore, the study addresses issues of equity in AI performance. TRUECAM demonstrated improved fairness metrics across sex and racial groups—a critical consideration as AI systems are increasingly scrutinized for potential biases that perpetuate healthcare disparities. The framework’s ability to generalize beyond lung cancer to other tissue types signals its broad applicability and potential to standardize trustworthy AI interpretations across oncology.</p>
<p>The authors hail from a diverse international coalition. Hong Kong Polytechnic University contributed lead author and corresponding author Xiaoge Zhang, PhD, an alumnus of Vanderbilt University, and PhD student Tao Wang. The University of Hong Kong provided corresponding author Maximus C.F. Yeung, MBBS, MSc. Vanderbilt Health’s Fedaa Najdawi, MBBS, Assistant Professor of Pathology, Microbiology, and Immunology, also participated actively. The study received partial funding from the National Institutes of Health (NIH) through award K99LM014428, underscoring government support for advancing safe AI integration in medicine.</p>
<p>TRUECAM’s unveiling marks a significant advancement towards the responsible implementation of AI in digital pathology and beyond. Its design reflects a deeper understanding of the complexities underlying medical image analysis and the necessity of embedding uncertainty awareness into AI workflows. As healthcare systems continue to adopt AI tools, frameworks like TRUECAM may become the standard bearers, ensuring that automated decisions are transparent, reliable, and aligned with clinical realities.</p>
<p>This development could revolutionize how pathologists and oncology teams employ AI in diagnostics, making it an indispensable ally rather than a source of unmitigated risk. By effectively flagging uncertainty and filtering out noise, TRUECAM not only enhances diagnostic precision but also fosters confidence among clinicians and patients alike—a crucial step towards widespread trust and adoption.</p>
<p>In sum, the work done by researchers at Vanderbilt Health and Hong Kong epitomizes the next frontier of medical AI: systems that are powerful yet prudent, capable of delivering accurate classifications while explicitly acknowledging their limits. Such innovation heralds a future where AI augments human expertise with clarity and caution, ultimately improving patient outcomes across diverse and challenging clinical environments.</p>
<p>Subject of Research:<br />
Not applicable</p>
<p>Article Title:<br />
Implementing trust in non-small cell lung cancer diagnosis with a conformalized uncertainty-aware AI framework</p>
<p>News Publication Date:<br />
23-Jun-2026</p>
<p>Web References:<br />
https://www.nature.com/articles/s41551-026-01694-8<br />
http://dx.doi.org/10.1038/s41551-026-01694-8</p>
<p>References:<br />
None provided</p>
<p>Image Credits:<br />
None provided</p>
<p>Keywords:<br />
Artificial intelligence, Pathology, Histology, Histological analysis, Cancer</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">167806</post-id>	</item>
	</channel>
</rss>
