Sunday, October 4, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors

October 4, 2026
in Technology and Engineering
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 4 mins read
0
Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors

Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors

Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

In medical imaging, a prediction that is right is not automatically a prediction that can be trusted. A deep learning model that declares a skin lesion benign with 99 percent confidence is clinically very different from one that flags the same lesion as uncertain, even if both end up on the correct side of the diagnostic line. Yet most benchmarks that rank artificial intelligence systems for medical image classification score them on discrimination alone, ignoring whether their stated probabilities actually mean what they claim. A new study published in Neural Computing and Applications confronts that blind spot directly, arguing that calibration and the structure of predictive uncertainty deserve equal billing with accuracy when ensemble methods are judged for clinical use.

The research, conducted by Hrushikesh Sanap, an independent researcher based in Chhatrapati Sambhajinagar, India, evaluates four architecturally diverse deep networks alongside two ensemble strategies across four datasets drawn from the MedMNIST v2 benchmark collection. The individual models span the modern computer vision landscape: ConvNeXt-Base, a convolutional network redesigned with contemporary training recipes; Vision Transformer Base, an attention-based architecture that treats image patches as sequences; EfficientNetV2-M, a compact and efficient convolutional family; and InceptionResNetV2, an older but still competitive hybrid of inception modules and residual connections. All four were fine-tuned from ImageNet pre-trained weights at a resolution of 224 by 224 pixels, a deliberately realistic setting for clinical pipelines where computational budgets matter.

The four tasks were chosen to cover a spread of modalities, class counts and imbalance regimes. BloodMNIST requires eight-class classification of blood cell microscopy images and is close to saturation in the literature. BreastMNIST is a binary malignancy detection problem on breast ultrasound, and notably the smallest dataset in the study. DermaMNIST involves seven-class dermatoscopic lesion classification with pronounced class imbalance, and OrganAMNIST asks models to identify one of eleven abdominal organs on computed tomography slices. Every trained configuration exceeded the previously reported benchmark accuracy on all four datasets, with margins ranging from a fraction of a percentage point on the near-saturated blood cell task to roughly fifteen percentage points on the dermatoscopy data.

Two ensemble strategies were then layered on top. Soft Voting simply averages the predicted probabilities of the member networks, a technique with roots stretching back to foundational work on neural network ensembles in 1990. Rigorous Stacking, by contrast, trains a meta-learner on the outputs of the base models, following the stacked generalization framework introduced by David Wolpert in 1992. The ensembles matched or exceeded the strongest individual model in almost every case, with a single exception: Rigorous Stacking on the small BreastMNIST set, where the meta-learner presumably lacked enough data to learn reliable combination weights.

The central and most striking finding is that accuracy and calibration are dissociated. Calibration, measured here through reliability diagrams and the expected calibration error, asks whether a model’s stated confidence matches its empirical hit rate: of all the cases a model labels with 80 percent confidence, roughly 80 percent should actually be correct. Across the experiments, the most accurate configuration was rarely the best calibrated. Soft Voting attained the lower expected calibration error on all four datasets and the higher accuracy on three of them, while the accuracy advantage of Rigorous Stacking was confined to the most imbalanced dataset, DermaMNIST, and even there it came at the cost of weaker, less usable probabilities. Overconfidence, a well-documented pathology of modern neural networks, proved pervasive among the individual models and was reduced, but not eliminated, by ensembling.

Predictive uncertainty was quantified through the Shannon entropy of each prediction’s probability distribution, comparing the entropy distributions of correct against incorrect answers. Throughout the experiments, incorrect predictions sat at higher entropy than correct ones, which is the essential precondition for confidence-based triage, the practice of routing uncertain cases to human experts while letting confident, correct predictions pass through automatically. But the study adds a crucial caveat: a strong aggregate calibration score can still mask a collapsed, non-informative per-prediction uncertainty distribution. In other words, a model can look well calibrated on average while being nearly useless at telling clinicians which individual cases deserve scrutiny. Soft Voting held the most consistent separation between the entropy distributions of correct and incorrect predictions, reinforcing its case as the dependable default.

Interpretability analyses using gradient-weighted class activation mapping for the convolutional models and attention-based methods for the transformer complemented the quantitative results, and the residual errors clustered at clinically recognized decision boundaries. In dermatoscopy, the melanoma-to-nevus distinction remained the dominant failure mode, and even the largest accuracy gain left roughly a quarter of melanomas assigned to the benign nevus category. In breast ultrasound, the dangerous errors were malignant false negatives, the cases where a cancer is waved through as benign. In abdominal CT, the models struggled with kidney laterality, a reminder that spatial reasoning failures can persist even when overall organ identification accuracy looks impressive.

The practical implications are pointed. The study argues that post-hoc calibration should be treated as a prerequisite for safe deployment rather than an optional refinement, and that ensemble methods for medical imaging should be evaluated on calibration and uncertainty structure as much as on discrimination accuracy. For hospitals and regulators weighing which systems to trust, this reframes the leaderboard: a marginally less accurate but better calibrated ensemble that knows when it does not know may save more lives than a sharper classifier that is confidently wrong at the worst moments. The finding that simple probability averaging outperformed a learned stacking layer on calibration, while costing almost nothing in accuracy, is a rare case of the cheaper and simpler option also being the safer one.

The author is careful to delineate the limits of the evidence. All results derive from single-source benchmark data at 224 by 224 resolution, from a single training run under one fixed random seed, so run-to-run variance and dataset shift remain untested. Higher-resolution inputs and prospective external validation on real clinical data are flagged as the necessary next steps toward deployment. The datasets themselves are publicly available through the MedMNIST v2 collection, and the model prediction outputs, ensemble implementation and evaluation code have been released in an open GitHub repository, allowing independent verification of the per-class, entropy and reliability values underlying the analysis. For a field racing to put diagnostic AI in front of patients, the message is uncomfortable but timely: before asking how often a model is right, ask whether it knows the difference between knowing and guessing.

Subject of Research: Calibration and uncertainty quantification in multi-architecture deep learning ensembles for medical image classification on MedMNIST benchmarks

Article Title: Beyond accuracy: calibration and uncertainty in multi-architecture ensembles for MedMNIST classification

Article References: Sanap, H. (2026). Beyond accuracy: calibration and uncertainty in multi-architecture ensembles for MedMNIST classification. Neural Computing and Applications, 38(19), Article 776. https://doi.org/10.1007/s00521-026-12521-1

Image Credits: AI Generated

DOI: 10.1007/s00521-026-12521-1

Keywords: medical image classification, ensemble learning, model calibration, uncertainty quantification, MedMNIST, expected calibration error, vision transformers, ConvNeXt, soft voting, stacked generalization, predictive entropy, clinical decision support

Cite Scienmag News

Blake Davidson. (October 4, 2026). Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors. Scienmag. https://scienmag.com/confident-but-wrong-why-accurate-medical-ai-can-still-mislead-doctors/

Blake Davidson. "Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors." Scienmag, 4 October 2026, https://scienmag.com/confident-but-wrong-why-accurate-medical-ai-can-still-mislead-doctors/. Accessed 4 October 2026.

Blake Davidson. "Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors." Scienmag. October 4, 2026. https://scienmag.com/confident-but-wrong-why-accurate-medical-ai-can-still-mislead-doctors/

Tags: challenges in deploying AI for clinical decision-makingclinical decision supportclinical trust in AI diagnosticsConvNeXtdeep learning architectures for medical image classificationensemble deep learning models for healthcareensemble learningensemble strategies in medical AIevaluation of AI confidence levels in medicineExpected Calibration Errorimportance of model probability calibrationmedical AI calibrationmedical image classificationMedMnistMedMNIST benchmark datasets for medical AImodel calibrationneural network calibration methodspredictive entropypredictive uncertainty in medical imagingrisks of overconfidence in AI diagnosticssoft votingstacked generalizationuncertainty quantificationVision Transformers
Share26Tweet16
Previous Post

Numbing the Uterus: Mepivacaine Instillation Eases Pain of IUD Placement in Trial

Next Post

When Machines Make the Forms: AI Loosens Art’s Oldest Coupling

Related Posts

When Machines Make the Forms: AI Loosens Art’s Oldest Coupling
Technology and Engineering

When Machines Make the Forms: AI Loosens Art’s Oldest Coupling

October 4, 2026
Whales and Swarms Help AI Read Cancer Slides With Record Accuracy
Technology and Engineering

Whales and Swarms Help AI Read Cancer Slides With Record Accuracy

October 4, 2026
Hospitals Can Now Train AI Together Without Sharing Patient Data
Technology and Engineering

Hospitals Can Now Train AI Together Without Sharing Patient Data

October 4, 2026
AI Reads the Market: GPT-Powered Risk Budgeting Beats Classic Portfolio Strategies
Technology and Engineering

AI Reads the Market: GPT-Powered Risk Budgeting Beats Classic Portfolio Strategies

October 3, 2026
Tiny Synthetic Probes Are Rewriting How Scientists See Inside Living Bodies
Technology and Engineering

Tiny Synthetic Probes Are Rewriting How Scientists See Inside Living Bodies

October 3, 2026
Electron Beam Geometry Rewrites the Surface Properties of Titanium Alloy
Technology and Engineering

Electron Beam Geometry Rewrites the Surface Properties of Titanium Alloy

October 3, 2026
Next Post
When Machines Make the Forms: AI Loosens Art’s Oldest Coupling

When Machines Make the Forms: AI Loosens Art's Oldest Coupling

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • How Much Tumor Is Enough to Leave Behind? Volumetric Study Redefines Craniopharyngioma Surgery Risk
  • When Machines Make the Forms: AI Loosens Art’s Oldest Coupling
  • Confident but Wrong: Why Accurate Medical AI Can Still Mislead Doctors
  • Numbing the Uterus: Mepivacaine Instillation Eases Pain of IUD Placement in Trial

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading