<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>evaluation of AI confidence levels in medicine &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/evaluation-of-ai-confidence-levels-in-medicine/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 00:19:08 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>evaluation of AI confidence levels in medicine &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors</title>
		<link>https://scienmag.com/confident-but-wrong-why-accurate-medical-ai-can-still-mislead-doctors/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 00:19:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[challenges in deploying AI for clinical decision-making]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[clinical trust in AI diagnostics]]></category>
		<category><![CDATA[ConvNeXt]]></category>
		<category><![CDATA[deep learning architectures for medical image classification]]></category>
		<category><![CDATA[ensemble deep learning models for healthcare]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[ensemble strategies in medical AI]]></category>
		<category><![CDATA[evaluation of AI confidence levels in medicine]]></category>
		<category><![CDATA[Expected Calibration Error]]></category>
		<category><![CDATA[importance of model probability calibration]]></category>
		<category><![CDATA[medical AI calibration]]></category>
		<category><![CDATA[medical image classification]]></category>
		<category><![CDATA[MedMnist]]></category>
		<category><![CDATA[MedMNIST benchmark datasets for medical AI]]></category>
		<category><![CDATA[model calibration]]></category>
		<category><![CDATA[neural network calibration methods]]></category>
		<category><![CDATA[predictive entropy]]></category>
		<category><![CDATA[predictive uncertainty in medical imaging]]></category>
		<category><![CDATA[risks of overconfidence in AI diagnostics]]></category>
		<category><![CDATA[soft voting]]></category>
		<category><![CDATA[stacked generalization]]></category>
		<category><![CDATA[uncertainty quantification]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=232714</guid>

					<description><![CDATA[A new study of ensemble deep learning on medical image benchmarks shows that the most accurate models are rarely the best calibrated, and argues that uncertainty and calibration must guide clinical deployment.]]></description>
										<content:encoded><![CDATA[<p>In medical imaging, a prediction that is right is not automatically a prediction that can be trusted. A deep learning model that declares a skin lesion benign with 99 percent confidence is clinically very different from one that flags the same lesion as uncertain, even if both end up on the correct side of the diagnostic line. Yet most benchmarks that rank artificial intelligence systems for medical image classification score them on discrimination alone, ignoring whether their stated probabilities actually mean what they claim. A new study published in Neural Computing and Applications confronts that blind spot directly, arguing that calibration and the structure of predictive uncertainty deserve equal billing with accuracy when ensemble methods are judged for clinical use.</p>
<p>The research, conducted by Hrushikesh Sanap, an independent researcher based in Chhatrapati Sambhajinagar, India, evaluates four architecturally diverse deep networks alongside two ensemble strategies across four datasets drawn from the MedMNIST v2 benchmark collection. The individual models span the modern computer vision landscape: ConvNeXt-Base, a convolutional network redesigned with contemporary training recipes; Vision Transformer Base, an attention-based architecture that treats image patches as sequences; EfficientNetV2-M, a compact and efficient convolutional family; and InceptionResNetV2, an older but still competitive hybrid of inception modules and residual connections. All four were fine-tuned from ImageNet pre-trained weights at a resolution of 224 by 224 pixels, a deliberately realistic setting for clinical pipelines where computational budgets matter.</p>
<p>The four tasks were chosen to cover a spread of modalities, class counts and imbalance regimes. BloodMNIST requires eight-class classification of blood cell microscopy images and is close to saturation in the literature. BreastMNIST is a binary malignancy detection problem on breast ultrasound, and notably the smallest dataset in the study. DermaMNIST involves seven-class dermatoscopic lesion classification with pronounced class imbalance, and OrganAMNIST asks models to identify one of eleven abdominal organs on computed tomography slices. Every trained configuration exceeded the previously reported benchmark accuracy on all four datasets, with margins ranging from a fraction of a percentage point on the near-saturated blood cell task to roughly fifteen percentage points on the dermatoscopy data.</p>
<p>Two ensemble strategies were then layered on top. Soft Voting simply averages the predicted probabilities of the member networks, a technique with roots stretching back to foundational work on neural network ensembles in 1990. Rigorous Stacking, by contrast, trains a meta-learner on the outputs of the base models, following the stacked generalization framework introduced by David Wolpert in 1992. The ensembles matched or exceeded the strongest individual model in almost every case, with a single exception: Rigorous Stacking on the small BreastMNIST set, where the meta-learner presumably lacked enough data to learn reliable combination weights.</p>
<p>The central and most striking finding is that accuracy and calibration are dissociated. Calibration, measured here through reliability diagrams and the expected calibration error, asks whether a model&#8217;s stated confidence matches its empirical hit rate: of all the cases a model labels with 80 percent confidence, roughly 80 percent should actually be correct. Across the experiments, the most accurate configuration was rarely the best calibrated. Soft Voting attained the lower expected calibration error on all four datasets and the higher accuracy on three of them, while the accuracy advantage of Rigorous Stacking was confined to the most imbalanced dataset, DermaMNIST, and even there it came at the cost of weaker, less usable probabilities. Overconfidence, a well-documented pathology of modern neural networks, proved pervasive among the individual models and was reduced, but not eliminated, by ensembling.</p>
<p>Predictive uncertainty was quantified through the Shannon entropy of each prediction&#8217;s probability distribution, comparing the entropy distributions of correct against incorrect answers. Throughout the experiments, incorrect predictions sat at higher entropy than correct ones, which is the essential precondition for confidence-based triage, the practice of routing uncertain cases to human experts while letting confident, correct predictions pass through automatically. But the study adds a crucial caveat: a strong aggregate calibration score can still mask a collapsed, non-informative per-prediction uncertainty distribution. In other words, a model can look well calibrated on average while being nearly useless at telling clinicians which individual cases deserve scrutiny. Soft Voting held the most consistent separation between the entropy distributions of correct and incorrect predictions, reinforcing its case as the dependable default.</p>
<p>Interpretability analyses using gradient-weighted class activation mapping for the convolutional models and attention-based methods for the transformer complemented the quantitative results, and the residual errors clustered at clinically recognized decision boundaries. In dermatoscopy, the melanoma-to-nevus distinction remained the dominant failure mode, and even the largest accuracy gain left roughly a quarter of melanomas assigned to the benign nevus category. In breast ultrasound, the dangerous errors were malignant false negatives, the cases where a cancer is waved through as benign. In abdominal CT, the models struggled with kidney laterality, a reminder that spatial reasoning failures can persist even when overall organ identification accuracy looks impressive.</p>
<p>The practical implications are pointed. The study argues that post-hoc calibration should be treated as a prerequisite for safe deployment rather than an optional refinement, and that ensemble methods for medical imaging should be evaluated on calibration and uncertainty structure as much as on discrimination accuracy. For hospitals and regulators weighing which systems to trust, this reframes the leaderboard: a marginally less accurate but better calibrated ensemble that knows when it does not know may save more lives than a sharper classifier that is confidently wrong at the worst moments. The finding that simple probability averaging outperformed a learned stacking layer on calibration, while costing almost nothing in accuracy, is a rare case of the cheaper and simpler option also being the safer one.</p>
<p>The author is careful to delineate the limits of the evidence. All results derive from single-source benchmark data at 224 by 224 resolution, from a single training run under one fixed random seed, so run-to-run variance and dataset shift remain untested. Higher-resolution inputs and prospective external validation on real clinical data are flagged as the necessary next steps toward deployment. The datasets themselves are publicly available through the MedMNIST v2 collection, and the model prediction outputs, ensemble implementation and evaluation code have been released in an open GitHub repository, allowing independent verification of the per-class, entropy and reliability values underlying the analysis. For a field racing to put diagnostic AI in front of patients, the message is uncomfortable but timely: before asking how often a model is right, ask whether it knows the difference between knowing and guessing.</p>
<p><strong>Subject of Research:</strong> Calibration and uncertainty quantification in multi-architecture deep learning ensembles for medical image classification on MedMNIST benchmarks</p>
<p><strong>Article Title:</strong> Beyond accuracy: calibration and uncertainty in multi-architecture ensembles for MedMNIST classification</p>
<p><strong>Article References:</strong> Sanap, H. (2026). Beyond accuracy: calibration and uncertainty in multi-architecture ensembles for MedMNIST classification. <em>Neural Computing and Applications, 38</em>(19), Article 776. <a href="https://doi.org/10.1007/s00521-026-12521-1" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12521-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12521-1" rel="noopener noreferrer">10.1007/s00521-026-12521-1</a></p>
<p><strong>Keywords:</strong> medical image classification, ensemble learning, model calibration, uncertainty quantification, MedMNIST, expected calibration error, vision transformers, ConvNeXt, soft voting, stacked generalization, predictive entropy, clinical decision support</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">232714</post-id>	</item>
	</channel>
</rss>
