Saturday, September 12, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know

September 12, 2026
in Technology and Engineering
Ophelia Keating
By Ophelia Keating Scienmag Editorial Profile - Health Services Research
Reading Time: 5 mins read
0
New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know

New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know

New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Deep learning models can now spot malaria parasites in blood smears, grade diabetic retinopathy from retinal photographs, and detect the earliest structural signatures of Alzheimer’s disease on brain MRI scans, often matching the performance of experienced clinicians. Yet a persistent problem has kept many of these systems out of routine clinical use: they deliver confident-looking answers without any reliable way of communicating when those answers, and the explanations behind them, should not be trusted. A new open-access study published in Machine Learning with Applications by Akshat Dubey, Aleksandar Anžel, Bahar İlgen, and Georges Hattab tackles this trust gap head-on with a framework called UbiQVision, which converts the explanations produced by deep vision models into formal mathematical evidence that can be weighed, fused, and, crucially, flagged as unreliable.

The core insight behind UbiQVision is that explainable artificial intelligence, or XAI, and uncertainty quantification have usually been treated as separate problems, when in fact they are inseparable in high-stakes medicine. The dominant explanation technique for medical imaging is SHAP, short for SHapley Additive exPlanations, a game-theoretic method that assigns each pixel a contribution score indicating how much it pushed the model toward or away from a diagnosis. SHAP produces visually compelling heatmaps that clinicians find intuitive. But the method carries hidden assumptions. Standard SHAP formulations effectively treat features as independent, while pixels in medical images are strongly correlated. When the underlying data distribution is misspecified or estimated from small, biased samples, SHAP values can become unstable, producing misleading rankings of imaging biomarkers or spurious emphasis on artifacts. Clinicians, susceptible to automation bias, may over-trust visually appealing heatmaps that do not faithfully reflect the model’s true reasoning.

UbiQVision addresses this by unifying three mathematical disciplines into a single pipeline. First, the researchers constructed a heterogeneous ensemble of three distinct neural network architectures: a lightweight custom convolutional neural network, the widely used residual network ResNet-18, and a Vision Transformer pre-trained on ImageNet. Architectural diversity matters because it ensures the models’ errors are not perfectly correlated, a prerequisite for meaningful evidence fusion. Second, instead of averaging the ensemble’s predictions uniformly, the framework applies Bayesian meta-learning. Each model’s reliability is modeled as a random variable following a Dirichlet distribution, updated with validation performance scores such as F1 metrics. A temperature parameter controls how sharply the weighting favors the strongest model, and sampling from this posterior gives each model a probabilistic vote that rewards robust performers while preserving the influence of weaker models that may have learned strong local evidence.

The third and most novel component is the transformation of SHAP attributions into basic probability assignments within Dempster–Shafer evidence theory, a classical framework for reasoning under uncertainty. Using a hyperbolic tangent transformation scaled by a sensitivity parameter, the framework maps unbounded, real-valued SHAP scores into bounded evidential masses. Positive attributions become mass supporting the target diagnosis, negative attributions become mass supporting its negation, and any leftover mass is assigned to the universal set, representing total epistemic ignorance. Dempster’s rule of combination then fuses the weighted masses from all three models into pixel-level maps of belief, plausibility, and uncertainty. A conflict coefficient, computed during fusion, explicitly quantifies where the models disagree, rather than smoothing that disagreement away as conventional ensemble averaging does.

The resulting outputs map directly onto clinical concepts. The belief map marks regions where the ensemble has reached confirmed consensus, such as the dark, ring-like chromatin structures of a malaria parasite inside an infected red blood cell. The plausibility map captures the upper bound of what could be true, exposing internal conflict when, for example, the noisy ResNet model highlights random tissue as pathological while the other models disagree. The uncertainty map quantifies total ignorance: bright yellow regions signal that the model genuinely knows nothing, correctly covering empty slide background or out-of-distribution inputs, while dark purple regions indicate the model has sufficient evidence to decide. This explicit separation of confirmed disease, conflicting opinions, and insufficient data is precisely what standard softmax classifiers, which force every pixel into a category, cannot provide.

The team evaluated the framework across three publicly available medical imaging datasets spanning histology, neuroimaging, and ophthalmology. On the NIH malaria dataset of 27,558 balanced blood smear images, the Bayesian weighting identified the custom CNN as the primary expert with a posterior weight of roughly 0.37, and the fused belief maps performed what amounts to semantic segmentation of the parasite, filtering out the cell wall and cytoplasm as irrelevant background. Ten-fold stratified cross-validation showed highly consistent macro F1 scores: ResNet averaged 96.2 percent, with the custom CNN and Vision Transformer close behind at 95.7 percent. Local Lipschitz stability analysis confirmed that the SHAP attributions feeding the fusion were mathematically stable, with all three architectures scoring below 0.0012, indicating the maps reflect genuine features rather than unstable gradient noise.

The Alzheimer’s disease experiments revealed perhaps the most clinically resonant behavior. Using T1-weighted MRI scans graded across four dementia stages, the framework captured the non-linear progression of brain atrophy by modulating its evidential confidence with disease severity. In moderate dementia cases, positive attributions aligned precisely with enlarged ventricular boundaries, and the belief map showed dense, localized clusters of confirmed pathological evidence. For very mild dementia, where atrophy is subtle and easily confused with healthy aging, the uncertainty maps showed widespread high entropy, mirroring the genuine diagnostic difficulty that human radiologists face. Notably, the framework exposed a well-known weakness in the field: the very mild dementia class produced the highest mean fused uncertainty, correctly signaling that the ensemble was operating near the limits of its knowledge rather than masking that limitation behind a confident label.

On the diabetic retinopathy dataset from the EyePACS Kaggle competition, the framework faced its hardest test, a five-class ordinal grading problem with subtle transitions between severity levels. Here the custom CNN struggled, achieving a mean macro F1 of only 46.1 percent, while the Vision Transformer and ResNet reached 68.7 and 67.9 percent respectively. The framework adapted, and its uncertainty behavior tracked clinical reality: severe diabetic retinopathy, characterized by massive hemorrhages and extensive ischemia, elicited the lowest median uncertainty, while proliferative disease with its ambiguous, newly forming vascular anomalies produced the highest. Ablation studies across all three datasets confirmed that progressive Gaussian blur, which destroys anatomical structure, caused mean fused uncertainty to rise monotonically, demonstrating that the framework’s ignorance estimates genuinely track epistemic uncertainty arising from missing structural information.

Beyond the maps themselves, selective prediction risk-coverage analysis showed that UbiQVision provides superior uncertainty calibration compared with deep ensemble variance, Monte Carlo dropout, and integrated gradients baselines. On the malaria dataset, the framework maintained a residual error rate of zero up to roughly 35 percent coverage, while baseline methods exhibited dangerous overconfidence spikes at lower coverage levels. The framework is entirely post-hoc and model-agnostic at the ensemble level, requiring no modification to validated training pipelines, which distinguishes it from evidential deep learning approaches that demand specialized loss functions. The authors acknowledge real limitations: computational cost is substantial, with inference times of 0.55 to 1.03 seconds per image and peak memory demands of 7.5 to 7.7 gigabytes, and image resolution was constrained to 128 by 128 pixels for most models due to the memory requirements of pixel-wise SHAP computation. Shared blind spots among models trained on identical data could also undermine the uncertainty estimates under adversarial conditions.

Even so, the implications for safety-critical medical AI are considerable. By making the unknown unknowns visible, the framework allows clinical workflows to route high-confidence predictions for expedited validation while directing uncertain or contested cases to expert review, a distinction directly relevant to regulatory requirements under the EU AI Act, which mandates transparency, robustness, and explainability in high-risk medical systems. The researchers envision extending the evidential fusion to multi-modal and longitudinal settings, tracking belief and ignorance at the patient level over time, and using the uncertainty outputs to drive active learning. The code is publicly available on GitHub, and the framework’s deeper contribution may be conceptual: it reframes medical AI from a binary classifier that masquerades confidence as certainty into a risk assessment tool that communicates, pixel by pixel, exactly how much it knows, how much it doubts, and where it is simply guessing.

Subject of Research: Uncertainty-aware explainable AI framework for reliable deep learning ensembles in medical imaging

Article Title: UbiQVision: Spatial Dempster-Shafer fusion of XAI attributions for reliable deep vision ensembles

Article References: Dubey, A., Anžel, A., İlgen, B., & Hattab, G. (2026). UbiQVision: Spatial Dempster–Shafer fusion of XAI attributions for reliable deep vision ensembles. Machine Learning with Applications, 25, Article 101000. https://doi.org/10.1016/j.mlwa.2026.101000

Image Credits: AI Generated

DOI: 10.1016/j.mlwa.2026.101000

Keywords: explainable AI, uncertainty quantification, Dempster–Shafer theory, medical imaging, deep learning ensembles, SHAP, Bayesian meta-learning, malaria detection, Alzheimer's disease, diabetic retinopathy, Vision Transformers, clinical decision support

Cite Scienmag News

Ophelia Keating. (September 12, 2026). New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know. Scienmag. https://scienmag.com/new-ai-framework-weighs-evidence-to-reveal-when-medical-vision-models-truly-know/

Ophelia Keating. "New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know." Scienmag, 12 September 2026, https://scienmag.com/new-ai-framework-weighs-evidence-to-reveal-when-medical-vision-models-truly-know/. Accessed 12 September 2026.

Ophelia Keating. "New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know." Scienmag. September 12, 2026. https://scienmag.com/new-ai-framework-weighs-evidence-to-reveal-when-medical-vision-models-truly-know/

Tags: AI transparency in clinical applicationsAlzheimer's diseaseBayesian meta-learningclinical decision supportdeep learning ensemblesdeep learning models for disease detectionDempster–Shafer theorydiabetic retinopathyexplainable AIexplainable AI in healthcareformal evidence generation for AI model explanationshigh-stakes medical AI decision reliabilityimproving trust in AI-driven medical diagnosesintegrating explainability and uncertainty in medical diagnosismalaria detectionMedical Imagingmedical vision model trustworthinessreliable AI explanations in medicineSHAPSHAP explainability method for medical imagesUbiQVision framework for medical AIuncertainty quantificationuncertainty quantification in medical imagingVision Transformers
Share26Tweet16
Previous Post

Soil Secrets Along Syria’s Elevation Gradient Reveal Fertility Clues

Next Post

Neighborhood Parks Tied to More Children and Weekday Foot Traffic in Tokyo

Related Posts

Fractional Calculus Meets Interval Optimization in New Study of Constrained Systems
Technology and Engineering

Fractional Calculus Meets Interval Optimization in New Study of Constrained Systems

September 12, 2026
New AI Network Weighs Which Senses to Trust When Reading Human Emotion
Technology and Engineering

New AI Network Weighs Which Senses to Trust When Reading Human Emotion

September 12, 2026
Gold Nanospheres and Grated CdS Layers Push Polymer Solar Cell Efficiency Toward 44%
Technology and Engineering

Gold Nanospheres and Grated CdS Layers Push Polymer Solar Cell Efficiency Toward 44%

September 12, 2026
Sponge Cities Cut Flood-Linked Dysentery Risk in China, Study Finds
Technology and Engineering

Sponge Cities Cut Flood-Linked Dysentery Risk in China, Study Finds

September 12, 2026
Crystal Symmetry Revealed as Hidden Architect of Zirconia’s Electronic and Optical Behavior
Technology and Engineering

Crystal Symmetry Revealed as Hidden Architect of Zirconia’s Electronic and Optical Behavior

September 12, 2026
Wildfire Smoke Study Reveals Hidden Toxic Chemicals in Reno’s Air
Technology and Engineering

Wildfire Smoke Study Reveals Hidden Toxic Chemicals in Reno’s Air

September 12, 2026
Next Post
Neighborhood Parks Tied to More Children and Weekday Foot Traffic in Tokyo

Neighborhood Parks Tied to More Children and Weekday Foot Traffic in Tokyo

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Neighborhood Parks Tied to More Children and Weekday Foot Traffic in Tokyo
  • New AI Framework Weighs Evidence to Reveal When Medical Vision Models Truly Know
  • Soil Secrets Along Syria’s Elevation Gradient Reveal Fertility Clues
  • Fractional Calculus Meets Interval Optimization in New Study of Constrained Systems

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading