Skin cancer remains one of the most common and deadly malignancies worldwide, and the difference between early detection and late diagnosis can be measured in lives. Dermoscopy, the examination of skin lesions with a specialized magnifying instrument, has produced vast libraries of clinical images that artificial intelligence systems can learn from. Yet the very richness of these dermoscopic images has become a bottleneck for the convolutional neural networks designed to interpret them. High-dimensional image data carry duplicated and redundant features that push models toward overfitting, meaning they memorize quirks of the training set rather than learning the true visual signatures of malignancy. The result is diagnostic uncertainty exactly where clinicians need confidence most. A new study published in Cluster Computing by Amruta Thorat and Chaya R. Jadhav of Dr. D. Y. Patil Institute of Technology in Pune, India, tackles this problem not by redesigning network architectures but by rethinking the optimizer that trains them.
The researchers introduce an optimization-based embedded feature selection framework powered by a hybrid optimizer that combines ADAMW, a variant of the popular ADAM algorithm with decoupled weight decay, with stochastic gradient descent, the classic workhorse of deep learning. The central idea is that the choice of optimization algorithm does more than simply minimize a loss function; it shapes which features the network effectively relies on. By progressively stabilizing gradients and suppressing noisy activations during training, the hybrid approach acts as an implicit filter, weeding out redundant features while the network learns. This embedded form of feature selection is woven directly into the training loop rather than applied as a separate preprocessing step, which the authors argue makes the resulting models both leaner and more robust when confronted with the noisy, high-dimensional character of real dermoscopic data.
To understand why this matters, it helps to recall how the two parent optimizers behave. ADAM and its derivatives maintain adaptive per-parameter learning rates by tracking running estimates of the first and second moments of the gradients, which speeds convergence but can settle into sharp, poorly generalizing minima. Stochastic gradient descent, by contrast, takes noisier but more exploratory steps that often lead to flatter minima associated with better generalization, though it converges more slowly and is sensitive to learning-rate schedules. The hybrid scheme alternates or blends these update rules so that the network benefits from ADAMW’s fast, stable early progress and SGD’s generalization-friendly late-stage refinement. In the authors’ formulation, this progression also dampens noisy activations, effectively pruning the influence of spurious image features that would otherwise inflate the model’s confidence without improving its diagnostic judgment.
The evaluation was deliberately broad. The framework was tested on three widely used dermoscopic datasets: ISIC-2020, a large patient-centric collection of images and metadata assembled for melanoma identification; a Malignant-Benign dataset of skin lesion photographs; and PH2, a smaller but carefully annotated benchmark of dermoscopic images. Across these datasets the authors trained four distinct backbone architectures, EfficientNet-B0, ResNet-50, InceptionV3, and ConvNeXt-Tiny, spanning several generations of convolutional network design, from compact mobile-friendly models to modern transformer-influenced convnets. Every configuration was assessed with five-fold cross-validation, a protocol that partitions the data into five subsets and rotates them through training and testing to ensure that reported performance is not an artifact of a single lucky split. This combination of multiple datasets, multiple architectures, and repeated validation gives the comparison unusual breadth for an optimizer study.
The headline numbers are striking. The hybrid ADAMW plus SGD optimizer significantly outperformed ADAMW, ADAM, and SGD used alone, reaching 95.4 percent accuracy and an area under the receiver operating characteristic curve of 96.6 percent on ISIC-2020, and 97.4 percent accuracy with 98.3 percent AUC on the PH2 dataset. AUC is a particularly meaningful metric in medical diagnosis because it summarizes the trade-off between sensitivity and specificity across all decision thresholds; a value above 98 percent indicates that the model ranks a randomly chosen malignant lesion above a randomly chosen benign one almost every time. The improvements were not left to visual inspection of a few numbers. The authors confirmed statistical significance with paired t-tests and Wilcoxon signed-rank tests, both returning p-values below 0.05, meaning the gains over the standard optimizers are unlikely to be chance fluctuations in the validation folds.
Accuracy figures alone can be misleading if a model is right for the wrong reasons, so the study goes a step further into explainability. The researchers employed Grad-CAM++, a gradient-based visualization technique that produces heatmaps highlighting which regions of an image most influenced the network’s decision. When trained with the hybrid optimizer, the models showed enhanced localization of malignant regions in the lesions, suggesting that the implicit feature selection was steering attention toward clinically relevant structures rather than background artifacts such as rulers, ink markers, or hair. To quantify this, the team applied activation similarity metrics including the Structural Similarity Index, the Pearson Correlation Coefficient, and Mean Absolute Deviation, comparing the models’ internal activation patterns and saliency maps. These measurements validated the claim that the optimizer-driven feature selection genuinely reshapes what the network learns to look at, not merely how confidently it guesses.
The work sits within a decade-long arc that began when deep neural networks first matched dermatologist-level performance on skin cancer classification, a landmark demonstrated in a 2017 Nature study by Andre Esteva and colleagues. Since then, a crowded literature has explored transfer learning, attention mechanisms, ensemble models, and metaheuristic optimizers such as grey wolf optimization to squeeze more accuracy from dermoscopic images. Thorat and Jadhav had previously surveyed many of these models in a comprehensive analysis of skin cancer detection approaches. Their new contribution is distinctive in treating the optimizer itself as the lever for feature selection, rather than bolting on a separate selection stage or a costly evolutionary search. This framing connects to a broader current in machine learning research, where systematic reviews of optimizers in medical deep learning have highlighted how much diagnostic performance depends on seemingly mundane training choices.
The clinical implications are considerable. A classifier that overfits redundant features may perform brilliantly on curated benchmarks yet falter on images from a different hospital, camera, or patient population, which is precisely the distribution shift that real-world deployment entails. By suppressing noisy activations and favoring stable, discriminative features, the hybrid optimizer produced models the authors describe as more robust and more explainable, two properties that regulators and clinicians consistently demand from computer-aided diagnosis tools. Better localization of malignant regions, visualized through Grad-CAM++, also gives dermatologists a way to sanity-check the machine’s reasoning, comparing the highlighted areas against their own reading of the lesion. In screening scenarios where a model triages thousands of images and flags suspicious ones for expert review, both the measured AUC and the interpretability of its attention directly affect patient outcomes.
The authors are candid about the road ahead. Their future research agenda centers on multimodal fusion, combining dermoscopic images with clinical metadata such as patient age, lesion location, and history, an approach other studies have shown can substantially improve classification, and on lightweight deployment on edge devices, which would allow the trained models to run on inexpensive hardware in clinics that lack high-performance computing. The datasets used in the study, including ISIC-2020, the Kaggle-hosted Malignant versus Benign collection, and PH2, are publicly available benchmarks, so other groups can readily test whether the hybrid ADAMW plus SGD recipe transfers to their own pipelines. No datasets were generated or analyzed beyond these existing resources, and the authors declare no competing interests.
For a field that has often chased accuracy through ever-larger architectures and ever-bigger datasets, the message of this study is refreshingly economical: how you train can matter as much as what you train. A carefully engineered blend of two classic optimization strategies, applied consistently across four architectures and three datasets, delivered statistically significant gains and visibly better attention maps without adding parameters or annotation burden. If the reported robustness holds up in prospective clinical evaluations, optimizer-driven feature selection could become a standard ingredient in the deep learning toolkits that help dermatologists catch melanoma earlier, turning the noisy abundance of dermoscopic data from a liability into an asset.
Subject of Research: Optimization-driven feature selection with a hybrid ADAMW and SGD optimizer for deep learning-based skin cancer detection from dermoscopic images
Article Title: Optimization-driven feature selection for enhanced skin cancer detection using deep learning
Article References: Thorat, A., & Jadhav, C. R. (2026). Optimization-driven feature selection for enhanced skin cancer detection using deep learning. Cluster Computing, 29(12), Article 730. https://doi.org/10.1007/s10586-026-06575-y
Image Credits: AI Generated
DOI: 10.1007/s10586-026-06575-y
Keywords: skin cancer, deep learning, feature selection, hybrid optimizer, ADAMW, stochastic gradient descent, convolutional neural networks, dermoscopy, melanoma, Grad-CAM, explainability, ISIC-2020
Cite Scienmag News
Nathaniel Bowman. (October 11, 2026). Hybrid Optimizer Sharpens Deep Learning’s Eye for Skin Cancer. Scienmag. https://scienmag.com/hybrid-optimizer-sharpens-deep-learnings-eye-for-skin-cancer/
Nathaniel Bowman. "Hybrid Optimizer Sharpens Deep Learning’s Eye for Skin Cancer." Scienmag, 11 October 2026, https://scienmag.com/hybrid-optimizer-sharpens-deep-learnings-eye-for-skin-cancer/. Accessed 11 October 2026.
Nathaniel Bowman. "Hybrid Optimizer Sharpens Deep Learning’s Eye for Skin Cancer." Scienmag. October 11, 2026. https://scienmag.com/hybrid-optimizer-sharpens-deep-learnings-eye-for-skin-cancer/

