Diabetic foot ulcers are one of the most feared complications of diabetes, affecting between 15 and 25 percent of people with the disease during their lifetime and contributing to as many as 85 percent of diabetes-related lower-limb amputations worldwide. With the International Diabetes Federation projecting that the number of adults living with diabetes will climb from more than 463 million in 2019 to roughly 700 million by 2045, the pressure to find fast, scalable and reliable ways of detecting these wounds before they spiral into infection and amputation has never been greater. A new systematic review published in Cluster Computing takes an unusually hard-nosed look at whether artificial intelligence is actually ready to meet that challenge, and its conclusions are a sobering corrective to the field’s headline numbers.
Led by Hadil Aldhubiea of Qatar University, together with colleagues in Qatar and Bangladesh, the review followed the PRISMA 2020 reporting standard and combed through PubMed, IEEE Xplore and Scopus for studies published between January 2019 and December 2024. From 141 records screened at the title-and-abstract stage, the team ultimately included 48 distinct deep learning study configurations drawn from 46 publications, spanning the full spectrum of diabetic foot ulcer analysis: image classification, wound grading, lesion detection, pixel-level segmentation, and multimodal fusion of different data streams. The search covered everything from convolutional neural networks and EfficientNet backbones to transformer hybrids, Siamese networks, generative adversarial networks and recurrent architectures applied to longitudinal patient data.
The technical landscape the reviewers mapped is dominated by convolutional neural networks, which appeared in roughly 60 percent of the included models. These networks learn hierarchical visual features, from edges to textures to wound patterns, and are well suited to capturing the fine-grained cues that matter in ulcer imagery, such as irregular wound edges, granulation tissue and the redness of surrounding skin. Transfer learning from large natural-image datasets like ImageNet remains the workhorse strategy, because most diabetic foot ulcer datasets are simply too small to train deep networks from scratch. More recent designs push further: DFU-SIAM combines an EfficientNetV2 backbone for local feature extraction with a BEiT transformer for global context, paired with a Large Margin Cotangent Loss to separate clinically similar categories under class imbalance, while DFU VIRNet fuses visible and infrared streams to catch physiological signals that ordinary photographs miss.
That multimodal angle is one of the most promising threads in the literature. Thermal imaging can reveal temperature gradients across the plantar surface that reflect inflammation, ischemia and neuropathic risk, potentially flagging danger before any wound is visible to the naked eye. The review found that within individual studies, multimodal pipelines combining RGB with thermal data beat their own unimodal baselines by a median of about 2.6 percentage points in classification accuracy and 3.1 points in Dice score for segmentation. Those gains are real but modest, and the authors caution that they are correlational, confounded with dataset choice, acquisition hardware and validation protocol. More troubling, many studies omit the engineering details, geometric registration, radiometric calibration, synchronization and normalization, that determine whether a fusion pipeline can be reproduced at another site or on another camera.
To move beyond anecdote, the reviewers built a 13-criterion quality rubric and computed a Weighted Methodological Rigor Score for every study, giving the heaviest weight to external validation, the single strongest indicator that reported performance is not an artifact of one dataset. A sensitivity analysis across five alternative weighting schemes confirmed the rankings were robust, with Spearman correlations of at least 0.92. The results were unflattering: the mean weighted score was just 49.0 percent, ranging from 23.8 to 95.2 percent. Only 27.1 percent of studies reported any external validation, a mere 8.3 percent released code, and just 16.7 percent discussed mobile or web deployment constraints. Just over half used publicly available datasets, with the Kaggle DFU collection and the DFUC grand challenge series the most common.
The most striking finding is a simple cross-analysis with big implications. Among the 36 configurations reporting a headline classification accuracy, the 11 studies that had undergone external validation reported a mean accuracy of 91.0 percent, while the 25 studies evaluated only on internal splits reported 94.3 percent. That gap of roughly 3.3 percentage points is exactly what you would expect if internal-only evaluation inflates performance, and it aligns with a broader pattern: 31 of those 36 configurations cluster between 90 and 99 percent accuracy, a suspiciously tight and optimistic distribution consistent with positive-outcome publication bias. The authors are careful to frame this as descriptive rather than causal, but the message is hard to ignore, headline accuracies in this field likely overstate what models will do in the real world.
Underlying many of these problems is the data itself. The review documents how lesion-to-image area ratios vary by more than two orders of magnitude across datasets, meaning a model tested on tightly cropped 150-by-150-pixel patches where the ulcer fills the frame faces a fundamentally easier task than one analyzing full-foot photographs where the wound occupies less than 5 percent of the image. Class imbalance was explicitly reported in 37.5 percent of studies, regional bias in 45.8 percent, and limited demographic diversity in 16.7 percent. Patch-based pipelines introduce a further hazard: if patches from the same patient appear in both training and test sets, leakage silently inflates results. Only about 62.5 percent of studies reported complete demographic information, and no publicly available dataset currently distributes paired images and clinical metadata under a redistributable license, a gap the reviewers identify as a priority for patient-aware modeling.
Evaluation practice compounds the confusion. Accuracy alone is misleading when severe categories like ischemia and gangrene are rare, and half of the included studies framed the task as binary ulcer-versus-healthy classification, which the authors argue has limited standalone clinical value because treatment decisions hinge on infection, ischemia, depth and Wagner grade. Segmentation studies report Dice and intersection-over-union inconsistently, detection papers often omit the IoU thresholds underlying their mean average precision figures, and pixel accuracy can look impressive simply because ulcers occupy so little of the image. The review proposes a minimum reporting checklist, including confusion matrices, macro-averaged metrics, boundary-aware segmentation measures, explicit patient-level splits and confidence intervals, as a baseline for future work.
The deployment picture is equally underdeveloped. Real-world use, especially home monitoring by patients with reduced foot sensation and possible visual impairment, demands pose-tolerant detection, robustness to mixed indoor lighting, and in-app capture guidance. Thermal pipelines carry their own discipline: plantar temperature is sensitive to recent footwear, ambient conditions and activity, and established protocols require at least 15 minutes of foot acclimatization in a controlled room at roughly 22 to 24 degrees Celsius before imaging. Models trained on acclimatized data will drift out of distribution the moment those conditions are not met, which is why the reviewers recommend contralateral-foot reference checks and restraint about high-confidence predictions on out-of-protocol captures. Not a single one of the 48 studies provided a complete specification for integration with electronic health records, a translational gap the authors call substantial.
The review’s closing roadmap is clear-eyed rather than defeatist. It calls for leakage-resistant benchmarking with patient-level separation and cross-dataset testing, standardized acquisition and calibration guidance for thermal imaging, lightweight and energy-aware model design for edge devices, clinically interpretable explainability extending to modality-specific contributions, and open release of code, weights and evaluation pipelines. The authors’ central conclusion reframes the field’s problem: progress in diabetic foot ulcer AI is currently limited less by the availability of powerful architectures than by the absence of standardized, transparent and clinically grounded evaluation. Until external validation, honest reporting and reproducible engineering become routine, the impressive accuracy figures filling the literature should be read as upper bounds, not guarantees, of what these systems will deliver at the bedside.
Subject of Research: Deep learning methods for diabetic foot ulcer detection, segmentation and multimodal assessment
Article Title: Deep learning for diabetic foot ulcer: a systematic review of segmentation, classification and multimodal integration
Article References: Aldhubiea, H., Fetais, N., Newaz, M., Islam, M. S. B., Chowdhury, M. E. H., Suganthan, P. N., & Kiranyaz, S. (2026). Deep learning for diabetic foot ulcer: a systematic review of segmentation, classification and multimodal integration. Cluster Computing, 29(14), Article 826. https://doi.org/10.1007/s10586-026-06619-3
Image Credits: AI Generated
DOI: 10.1007/s10586-026-06619-3
Keywords: diabetic foot ulcer, deep learning, systematic review, convolutional neural networks, image segmentation, thermal imaging, multimodal fusion, external validation, reproducibility, medical imaging, clinical deployment, PRISMA
Cite Scienmag News
Blake Davidson. (October 6, 2026). AI Can Spot Diabetic Foot Ulcers, But Most Studies May Be Overstating How Well It Works. Scienmag. https://scienmag.com/ai-can-spot-diabetic-foot-ulcers-but-most-studies-may-be-overstating-how-well-it-works/
Blake Davidson. "AI Can Spot Diabetic Foot Ulcers, But Most Studies May Be Overstating How Well It Works." Scienmag, 6 October 2026, https://scienmag.com/ai-can-spot-diabetic-foot-ulcers-but-most-studies-may-be-overstating-how-well-it-works/. Accessed 6 October 2026.
Blake Davidson. "AI Can Spot Diabetic Foot Ulcers, But Most Studies May Be Overstating How Well It Works." Scienmag. October 6, 2026. https://scienmag.com/ai-can-spot-diabetic-foot-ulcers-but-most-studies-may-be-overstating-how-well-it-works/

