Deep learning has transformed medical imaging, but its appetite for data comes at a steep price. Training modern neural networks on the vast archives of chest X-rays and computed tomography scans that hospitals accumulate can consume enormous computational resources, often requiring hundreds of thousands of patient images to reach clinically useful accuracy. A new study published in the International Journal of Data Science and Analytics suggests that much of that data may be redundant. Researchers led by Mert Sehri of the University of Ottawa, working with colleagues at VSB – Technical University of Ostrava, demonstrate that a carefully engineered data loading strategy called selective embedding can deliver competitive diagnostic classification performance while training on dramatically fewer images.
The core insight behind the work is deceptively simple: beyond a certain point, adding more samples to a training set yields diminishing returns. The authors show mathematically and empirically that once a dataset has captured the effective diversity needed for learning, additional images contribute little new information. This early saturation means that the brute-force approach of feeding networks ever-larger volumes of radiological data wastes both energy and money. Instead of changing the architecture of the model or inventing a new loss function, the team focused on the order and composition of the data presented to the network during training, an often-overlooked lever in the machine learning pipeline.
Selective embedding, first introduced by the same group in a previous paper in Knowledge-Based Systems, is a preprocessing and data loading method that controls how training samples reach the model. In the medical imaging experiments described in the new study, the researchers paired the technique with TransUNet, a hybrid architecture that combines convolutional encoders with transformer-based attention, originally designed for medical image segmentation. The network’s encoder consists of three convolutional blocks with batch normalization, ReLU activations and max pooling, followed by a bottleneck and a transformer stage with four layers, eight attention heads and a multilayer perceptron ratio of 4.0. A classification head built on global average pooling maps the learned features to disease labels, while an auxiliary segmentation head remains unused in the loss.
The experiments drew on two publicly available medical imaging resources: CheXpert, a large chest radiograph dataset with uncertainty labels spanning fourteen thoracic pathologies, and a Kaggle collection of chest CT scans covering COVID-19, normal and pneumonia cases. The selective embedding loader alternates strictly between X-ray and CT samples during training, oversampling the minority modality so that the two streams are balanced. With a batch size of eight, this means every training batch contains exactly four X-ray and four CT images, enforcing both per-batch modality balance and temporal alternation simultaneously. During validation, the majority modality is downsampled instead, avoiding artificial inflation of performance estimates.
The results are striking. On the fourteen-label classification task, the method achieved a Macro-AUROC of approximately 0.870, and the authors report competitive performance exceeding 87 percent Macro-AUROC across medical imaging tasks overall, all while using reduced training data. In the simpler five-label CheXpert task with patient-level labels, near-perfect scores were achieved on the primary metrics. Perhaps most telling is a control experiment the team designed to disentangle the two ingredients of their loader. A stratified balanced batching baseline matched the same four-to-four modality balance within each batch but without alternation, and it reached only a Macro-AUROC of 0.785, well short of the 0.889 achieved by the full selective embedding approach. This gap indicates that the strict alternation of modalities, not merely the balance, is the primary driver of the performance advantage.
To understand why alternating inputs helps, the researchers turned to diagnostic tools drawn from optimization theory. They tracked the norm of the input Jacobian, which measures how sensitive the model’s outputs are to small perturbations of the input, as well as the Hutchinson trace of the Hessian at the classification head, which probes the curvature of the loss surface. Throughout training, Jacobian regularity values remained low and stable, suggesting that the model was not overreacting to minor input variations, a property closely tied to good generalization. The Hessian trace stayed within moderate ranges, indicating that the optimizer avoided sharp minima, the narrow basins in the loss landscape that are notorious for poor performance on unseen data.
These observations support a coherent theoretical picture. By presenting inputs from different modalities in strict alternation, the loader prevents the optimizer from over-specializing to the features of any single imaging type. The gradient updates are forced to accommodate heterogeneous data at every step, which smooths the loss landscape and keeps representation drift under control. The authors supply a mathematical explanation for why this controlled diversity reduces the amount of data required while maintaining classification performance, grounding the empirical findings in established results from information theory and statistical learning, including classical work on empirical processes and PAC-Bayesian generalization bounds.
The training protocol itself was deliberately modest, which strengthens the case for practical adoption. Images were converted to grayscale, replicated to three channels and resized to 224 by 224 pixels, with only light augmentation in the form of random horizontal flipping and rotations of up to seven degrees. The model was optimized with AdamW at a learning rate of 1e-5 and a weight decay of 1e-4, using a cosine annealing schedule over twenty epochs with a batch size of eight. All reported Macro-AUROC figures are averaged over three random seeds, with standard deviations reported alongside the means, lending statistical credibility to the comparisons. The loss function was binary cross entropy with logits, with positive weights adjusted per label to account for class imbalance, a common challenge in medical datasets where diseases are unevenly represented.
The implications extend well beyond the specific datasets studied. Hospitals and research labs in low-resource settings often cannot assemble the massive corpora that headline-grabbing AI systems rely on, and the energy cost of training on hundreds of thousands of images sits uneasily with the growing scrutiny of artificial intelligence’s environmental footprint. A method that extracts more information from fewer samples, simply by reordering how the network sees the data, requires no new hardware, no exotic architecture and no proprietary ingredients. Because the approach operates at the data loading stage, it is in principle compatible with a wide range of models and could be combined with other efficiency techniques such as pruning, distillation or automated augmentation strategies like RandAugment and LayerMix.
There are, of course, caveats. The study relied exclusively on publicly available, de-identified datasets, and the authors note that no new data were generated during the work, so prospective validation on clinical data streams remains a task for future research. The alternating scheme as described is tailored to two-modality settings, and extending it to richer mixtures of imaging types, or to other domains such as the vibration and acoustic signal analysis the group has explored in industrial fault diagnosis, will require further engineering. Still, the central message is a provocative one for the field: the path to data-efficient medical AI may not run through bigger models or bigger datasets, but through a smarter conversation between the data and the learner, one carefully alternated batch at a time.
Subject of Research: Data-efficient deep learning for medical image classification using selective embedding data loading
Article Title: Selective embedding for data-efficient deep learning in medical image classification
Article References: Sehri, M., Cakir, E., Vashishtha, G., Chauhan, S., & Dumond, P. (2026). Selective embedding for data-efficient deep learning in medical image classification. International Journal of Data Science and Analytics, 22(1), Article 332. https://doi.org/10.1007/s41060-026-01314-3
Image Credits: AI Generated
DOI: 10.1007/s41060-026-01314-3
Keywords: selective embedding, deep learning, medical imaging, chest X-ray, CT scans, TransUNet, CheXpert, data efficiency, Macro-AUROC, generalization, data loading, machine learning
Cite Scienmag News
Blake Davidson. (October 8, 2026). Smarter Data Loading Cuts Medical AI Training Needs Without Sacrificing Accuracy. Scienmag. https://scienmag.com/smarter-data-loading-cuts-medical-ai-training-needs-without-sacrificing-accuracy/
Blake Davidson. "Smarter Data Loading Cuts Medical AI Training Needs Without Sacrificing Accuracy." Scienmag, 8 October 2026, https://scienmag.com/smarter-data-loading-cuts-medical-ai-training-needs-without-sacrificing-accuracy/. Accessed 8 October 2026.
Blake Davidson. "Smarter Data Loading Cuts Medical AI Training Needs Without Sacrificing Accuracy." Scienmag. October 8, 2026. https://scienmag.com/smarter-data-loading-cuts-medical-ai-training-needs-without-sacrificing-accuracy/

