An artificial intelligence model that can recognize cancerous skin lesions with remarkable precision has been unveiled by a team of biomedical engineers working in Iraq and Oman, and the most striking thing about it is not the headline accuracy figure but where that accuracy holds up. The system, described in the journal Multimedia Tools and Applications, correctly classified dermoscopic images of pigmented skin lesions 98.47 percent of the time on a large public benchmark, and then, without a single additional round of training, achieved 95.17 percent accuracy on a completely separate dataset collected with different cameras, different patients, and different clinical protocols. In a field where deep learning models routinely crumble when moved from the data they were trained on to the messy reality of a hospital, that kind of cross-domain performance is the metric that matters.
The work was carried out by Zeyad Qasim Habeeb of the University of Technology in Baghdad, Branislav Vuksanovic of the Military Technological College in Muscat, and Imad Q. Alzaydi of the University of Information Technology and Communications in Baghdad. Their starting point was ConvNeXt V2-Large, a modern convolutional neural network architecture that has emerged as a serious rival to vision transformers for image recognition tasks. Rather than adopting the architecture off the shelf, the researchers surgically modified it for the specific visual challenges of dermoscopy, the technique in which dermatologists photograph skin lesions through a specialized magnifying lens that renders subsurface structures of the skin visible.
The first set of modifications concerns attention, the mechanism by which a neural network learns to prioritize the parts of an image that carry diagnostic weight. The team embedded Convolutional Block Attention Module, or CBAM, blocks into the network. CBAM operates along two complementary dimensions: a channel attention module that learns which feature maps, and therefore which kinds of visual patterns, deserve emphasis, and a spatial attention module that learns where in the image the network should focus. In a dermoscopic image, this distinction is crucial. The lesion itself may occupy a small fraction of the frame, surrounded by healthy skin, hair, oil bubbles, or the dark circular border that the dermatoscope itself imposes. By forcing the network to weigh both what it sees and where it sees it, the attention modules help the model lock onto clinically meaningful structures such as irregular pigment networks, blue-white veils, and atypical vascular patterns.
The second major change replaces the conventional fully connected classification head at the end of the network with a Global Average Pooling layer followed by batch normalization, a combination the authors refer to as GAP-BN. Global Average Pooling collapses each feature map into a single value by averaging across its spatial extent, which drastically reduces the number of parameters in the classifier and, with it, the risk of overfitting, the failure mode in which a model memorizes the training images instead of learning generalizable rules. Batch normalization then stabilizes and standardizes the activations feeding into the final decision layer. For a medical imaging task where training data, however large, is still finite and expensive to label, this leaner head is a pragmatic safeguard against a model that looks brilliant in the lab and fails in the clinic.
Training strategy received equal attention to architecture. The researchers applied a suite of domain-specific augmentation techniques designed to simulate the natural variability of dermoscopic images. Among them was CutMix, a regularization method that cuts a patch from one training image and pastes it onto another, forcing the classifier to learn from blended, partially occluded scenes and to localize its features rather than relying on global shortcuts. They also used random elastic deformations, which warp the tissue geometry in ways that mimic how lesions stretch and shift across different body sites and different patients, and adaptive color normalization, which addresses one of the most stubborn problems in dermoscopy: the wide variation in illumination, calibration, and color rendering across imaging devices. A lesion that appears deep brown under one dermatoscope may appear washed out under another, and a model that confuses color shift with pathological change will not survive contact with real-world data.
The primary training and evaluation platform was the ISIC 2019 dataset from the International Skin Imaging Collaboration, a widely used benchmark containing a highly diverse collection of dermoscopic images distributed across eight clinically significant diagnostic categories, including melanoma, basal cell carcinoma, squamous cell carcinoma, actinic keratosis, benign keratosis-like lesions, melanocytic nevi, vascular lesions, and dermatofibroma. The breadth of these categories matters, because a model trained to distinguish only melanoma from moles learns a far simpler task than one that must navigate the full differential diagnosis a dermatologist faces. The 98.47 percent accuracy the enhanced model achieved on this eight-way classification task places it ahead of the comparison architectures and prior studies the authors evaluated.
The more demanding test, however, came next. The researchers took the trained model and applied it, unchanged, to the Dermatology Seven-Point Checklist dataset, known as DERM7PT, an independent external test set that the model had never seen in any form. No retraining, no fine-tuning, no adjustment of a single weight. This is the evaluation that separates models that have genuinely learned the visual language of skin disease from models that have memorized the statistical quirks of one dataset. DERM7PT was assembled under different conditions, with different equipment and patient populations, so its images differ systematically from ISIC 2019 in color distribution, resolution, and framing. The model still classified 95.17 percent of these external images correctly, a result the authors describe as crucial evidence of effectiveness in real-world scenarios where image properties vary.
The significance of this work sits within a broader and rapidly accelerating effort to bring deep learning into dermatology. Skin cancer is among the most common cancers worldwide, and melanoma, while less frequent than other skin cancers, is by far the deadliest when caught late. Earlier and more accurate diagnosis saves lives, and dermoscopy, although powerful in trained hands, is highly operator-dependent: studies comparing AI-based image classification with expert and non-expert dermatologists have shown how much diagnostic agreement varies with experience. A reliable algorithmic second opinion could therefore be valuable not only in well-resourced clinics but in primary care settings and underserved regions where specialist dermatologists are scarce. The authors of the new study, who have previously applied similar architecture-modification strategies to lung cancer detection and to the early detection of Parkinson’s disease from medical scans, frame their approach as part of a pipeline that empowers patients through AI and explainable human-computer interaction in personalized healthcare.
The field the researchers are pushing against is crowded and competitive. Recent years have produced attention-guided dual autoencoders, hybrid models fusing Squeeze-Excitation and DenseNet architectures, Swin transformers with shifted-window self-attention, ConvNeXt variants with focal self-attention, ensemble methods with test-time augmentation, and optimization schemes borrowed from ant colony and Harris Hawk metaheuristics. Many of these report impressive numbers on ISIC benchmarks, but far fewer report rigorous external validation on an untouched second dataset. That gap between benchmark performance and deployable generalization has been a persistent criticism of medical deep learning, and it is precisely the gap that the DERM7PT result is meant to address. The datasets themselves, ISIC 2019 and DERM7PT, are publicly available, which means other groups can verify the claims independently.
There are, as always, caveats before any such system reaches a patient’s bedside. Accuracy figures on curated research datasets, even external ones, do not guarantee performance on the full spectrum of skin tones, lesion locations, and imaging conditions encountered in routine practice, and regulatory approval for diagnostic software demands extensive prospective clinical validation. The authors report no competing interests and received no external funding for the research. Still, the combination of a carefully modified modern convolutional backbone, dual-dimension attention, a parameter-lean classifier, and augmentation tailored to the physics of dermoscopic imaging offers a template for how medical AI might earn the right to generalize. If subsequent studies replicate the cross-domain results on additional cohorts, the humble dermoscope, paired with an algorithm that knows both what to look at and where, could become one of the most consequential diagnostic tools of the decade.
Subject of Research: Deep learning classification of dermoscopic skin lesion images with cross-domain generalization
Article Title: Enhanced ConvNeXt model for dermoscopic skin lesion classification with cross-domain generalization
Article References: Habeeb, Z. Q., Vuksanovic, B., & Alzaydi, I. Q. (2026). Enhanced ConvNeXt model for dermoscopic skin lesion classification with cross-domain generalization. Multimedia Tools and Applications, 85(10), Article 786. https://doi.org/10.1007/s11042-026-21927-x
Image Credits: AI Generated
DOI: 10.1007/s11042-026-21927-x
Keywords: skin cancer, dermoscopy, deep learning, ConvNeXt V2, attention mechanism, CBAM, ISIC 2019, DERM7PT, melanoma, image classification, cross-domain generalization, data augmentation
Cite Scienmag News
Nathaniel Bowman. (September 30, 2026). AI Spots Skin Cancer With Record Accuracy Across Different Hospitals’ Images. Scienmag. https://scienmag.com/ai-spots-skin-cancer-with-record-accuracy-across-different-hospitals-images/
Nathaniel Bowman. "AI Spots Skin Cancer With Record Accuracy Across Different Hospitals’ Images." Scienmag, 30 September 2026, https://scienmag.com/ai-spots-skin-cancer-with-record-accuracy-across-different-hospitals-images/. Accessed 30 September 2026.
Nathaniel Bowman. "AI Spots Skin Cancer With Record Accuracy Across Different Hospitals’ Images." Scienmag. September 30, 2026. https://scienmag.com/ai-spots-skin-cancer-with-record-accuracy-across-different-hospitals-images/

