<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>self-explaining AI models in healthcare &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/self-explaining-ai-models-in-healthcare/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 06 Oct 2026 12:04:20 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>self-explaining AI models in healthcare &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI That Explains Itself: New Benchmark Reveals Which Neural Networks Truly See Pneumonia</title>
		<link>https://scienmag.com/ai-that-explains-itself-new-benchmark-reveals-which-neural-networks-truly-see-pneumonia/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 12:04:20 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[benchmarking convolutional neural networks for chest X-ray analysis]]></category>
		<category><![CDATA[chest X-ray]]></category>
		<category><![CDATA[chest X-ray image classification for pneumonia]]></category>
		<category><![CDATA[clinical AI]]></category>
		<category><![CDATA[comparative study of CNN architectures for pneumonia diagnosis]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[DenseNet121]]></category>
		<category><![CDATA[explainability and accuracy trade-offs in AI for radiology]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[generalization]]></category>
		<category><![CDATA[impact of neural network architecture on medical diagnosis]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[MobileNetV2]]></category>
		<category><![CDATA[model interpretability]]></category>
		<category><![CDATA[neural network explainability in medical imaging]]></category>
		<category><![CDATA[performance of MobileNetV2 in medical imaging]]></category>
		<category><![CDATA[pneumonia detection]]></category>
		<category><![CDATA[pneumonia detection using deep learning]]></category>
		<category><![CDATA[robustness of neural networks on unseen patient data]]></category>
		<category><![CDATA[self-explaining AI models in healthcare]]></category>
		<category><![CDATA[standardized training protocols for medical AI models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=241270</guid>

					<description><![CDATA[A standardized benchmark of six convolutional neural networks shows that the most accurate pneumonia-detecting models also produce the most faithful and stable explanations, a coupling that holds even when tested on an adult population the networks never saw during training.]]></description>
										<content:encoded><![CDATA[<p>Pneumonia remains one of the world&#8217;s deadliest infections, and chest X-rays are usually the first line of defense. But as radiology departments drown in images, hospitals are increasingly turning to deep learning to help flag suspicious scans. A new benchmarking study published in Applied Intelligence by researchers at Universidad Pablo de Olavide in Seville has now delivered a finding that could reshape how clinicians choose their AI: the best-performing neural networks are also the ones that explain themselves most faithfully, and that coupling survives even when the models are confronted with patients they were never trained on.</p>
<p>The research team, led by Francisco Gómez-Vela together with Aurelio López-Fernandez, Federico Divina and Miguel García-Torres, put six convolutional neural network architectures through an identical gauntlet: VGG16, ResNet50, DenseNet121, MobileNetV2, EfficientNetB0 and ConvNeXt-Tiny. Each was trained on the same pediatric chest X-ray dataset of 5,863 radiographs from Guangzhou Women and Children&#8217;s Medical Center, using the same optimizer, learning rate, batch size and early-stopping rules. That standardization matters, because most previous studies compared models trained under different conditions, making it impossible to tell whether performance differences came from the architecture or the training recipe.</p>
<p>The results upended some expectations. MobileNetV2, a lightweight network designed for smartphones rather than supercomputers, achieved the highest area under the ROC curve at 0.99, while DenseNet121, famous for its densely connected layers that recycle features across the network, matched it statistically across every metric. More surprising still, the venerable VGG16, a 2014 design with roughly 138 million parameters and no residual connections at all, took third place, beating both ResNet50 and the transformer-inspired ConvNeXt-Tiny. The researchers attribute this to the frozen-backbone transfer learning protocol: when the task does not demand deep domain-specific adaptation, architectural sophistication does not necessarily buy better predictions.</p>
<p>But raw accuracy was only half the story. The team went beyond the usual practice of eyeballing a few heatmaps and instead quantified explanation quality using five metrics: Deletion and Insertion AUC, which test whether the regions a model highlights actually drive its predictions; Sparsity and Entropy, which measure how focused those highlights are; and Stability SSIM, which checks whether explanations stay consistent when the input is perturbed. DenseNet121 came out on top across the board, producing attribution maps that were causally meaningful, compactly localized on lung tissue, and stable under noise.</p>
<p>The contrast cases were instructive. ResNet50 produced attribution maps that spread relevance almost uniformly across the image, so its explanations were numerically stable but essentially uninformative. EfficientNetB0 fared even worse: it collapsed deterministically to predicting a single class across all cross-validation folds, yielding a Matthews correlation coefficient of zero, and its gradient-based saliency maps came out entirely black. Perfect stability, the authors caution, can signal degenerate behavior rather than genuine robustness, a warning for anyone who equates consistent explanations with trustworthy ones.</p>
<p>To formalize the trade-off, the researchers introduced a Performance-Interpretability Index that multiplies a model&#8217;s AUC by the average of its Deletion and Insertion AUC scores. The multiplicative form deliberately penalizes models that classify well but explain poorly. Under this composite criterion, DenseNet121 ranked first, MobileNetV2 second, and VGG16 third, while ConvNeXt-Tiny and EfficientNetB0 sank to the bottom despite their modern pedigrees. The index offers hospitals a practical shortcut: instead of weighing accuracy against explainability as competing goals, they can select architectures that maximize both simultaneously.</p>
<p>The study&#8217;s most stringent test came from an unusual experimental design. The models were trained exclusively on pediatric X-rays, in which pneumonia typically appears as lobar consolidations or perihilar infiltrates, and then validated without any retraining on an independent adult dataset of over 15,000 images, where viral pneumonia manifests as bilateral ground-glass opacities in the lung periphery. This deliberate distributional shift simulates the demographic mismatch AI systems face in real deployment. DenseNet121 led the transfer with an external AUC of 0.83 and a recall of 0.98, while MobileNetV2 posted the best accuracy and F1-score with an AUC of 0.81. Crucially, the performance-interpretability ranking was preserved across populations.</p>
<p>The external validation also exposed a false friend. ResNet50, which had scored a respectable 0.96 internally, fell to an AUC of 0.43 on adult data, below random chance. The authors argue this is exactly the kind of hidden failure that explainability metrics can predict: ResNet50&#8217;s diffuse, unfocused attribution maps had already revealed that it was leaning on dataset-specific cues rather than transferable pathological features. High test performance without stable, localized explanations, they conclude, is a false positive of reliability.</p>
<p>The clinical implications are concrete. In triage settings where sensitivity is paramount, a model like DenseNet121, which rarely misses a true case, is the natural choice. In resource-constrained or point-of-care environments, MobileNetV2 offers nearly the same diagnostic power at a fraction of the computational cost, making it suitable for portable and real-time systems. VGG16&#8217;s strong internal showing but weaker cross-population generalization suggests its sheer parameter count encourages overfitting to the training distribution rather than learning transferable representations, a caution against equating model size with robustness.</p>
<p>The authors are careful to note the limits of their framework. All the explainability metrics are model-centric proxies that measure faithfulness to the network&#8217;s own reasoning, not to radiologist-annotated ground truth, and the study is a methodological benchmark rather than a clinical validation. Future work will extend the comparison to Vision Transformers, pursue prospective validation with radiologist annotations, and test how the lightweight architectures fare on edge devices. All code and data are publicly available, making this one of the first end-to-end reproducible pipelines that treats interpretability not as an afterthought but as a core criterion for deciding which AI deserves a place in the clinic.</p>
<p><strong>Subject of Research:</strong> Benchmarking deep learning architectures and explainable AI for pneumonia detection in chest X-rays</p>
<p><strong>Article Title:</strong> Performance-interpretability trade-offs and generalization in deep learning for pneumonia detection: A benchmarking study</p>
<p><strong>Article References:</strong> Gómez-Vela, F., López-Fernandez, A., Divina, F., &amp; García-Torres, M. (2026). Performance-interpretability trade-offs and generalization in deep learning for pneumonia detection: A benchmarking study. <em>Applied Intelligence, 56</em>(14), Article 414. <a href="https://doi.org/10.1007/s10489-026-07398-5" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07398-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07398-5" rel="noopener noreferrer">10.1007/s10489-026-07398-5</a></p>
<p><strong>Keywords:</strong> deep learning, pneumonia detection, chest X-ray, explainable AI, convolutional neural networks, model interpretability, benchmarking, generalization, medical imaging, DenseNet121, MobileNetV2, clinical AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">241270</post-id>	</item>
	</channel>
</rss>
