Artificial intelligence is being promoted as a powerful bridge between medical images and the hidden genetics of cancer. But a new critique of a major review on lung cancer imaging warns that the bridge may not yet be as sturdy as the headlines suggest. Although AI systems can identify patterns in computed tomography scans that are invisible to the human eye, the evidence connecting those patterns to specific cancer-driving mutations remains uneven, difficult to compare and potentially more optimistic than the overall data justify.
The commentary, published in the Journal of Cancer Research and Clinical Oncology, examines a 2026 review by Xue and Chen that surveyed more than 150 studies on artificial intelligence, lung ground-glass nodules, adenocarcinoma and carcinogenic driver genes. Ground-glass nodules are hazy regions on CT scans that do not completely obscure the structures beneath them. They can represent early lung adenocarcinoma, a noncancerous lesion or an intermediate condition, making them clinically important but notoriously difficult to classify. The hope behind radiogenomics—the combination of medical imaging and molecular genetics—is that an algorithm might infer a tumor’s biological behavior, including its mutations, from its visual appearance.
This approach relies on the idea that genetic changes can alter the architecture and behavior of cancer cells in ways that eventually become visible in an image. Mutations in genes such as EGFR, for example, can influence cell growth, tissue organization, tumor density and the interaction between a lesion and its surrounding lung. CT scanners record differences in X-ray attenuation across tissue, producing measurements related to density, shape, margins, internal texture and spatial structure. Radiomics converts these images into large numerical datasets, while machine-learning models search for statistical associations between those features and molecular labels obtained from pathology or sequencing.
Deep-learning systems can extend this process by learning image representations directly from pixels or three-dimensional volumes. A model may be trained to distinguish nodules associated with one mutation from those associated with another, or to predict whether a tumor carries a potentially actionable alteration. Performance is often summarized using the area under the receiver operating characteristic curve, or AUC. An AUC of 0.5 indicates performance no better than random guessing, while an AUC of 1.0 represents perfect separation between categories. The studies discussed in the review reported values ranging from roughly 0.64 to above 0.95, depending on the gene, imaging method and patient cohort.
That wide range is precisely where the problem begins, according to V. P. Abel Jopaul and M. Lingaraj, the authors of the new commentary. The review did not explain how its literature was located and selected. It provided no stated databases, search dates, search terms or inclusion and exclusion criteria. Nor did it clarify whether the review was intended to be systematic or narrative. This omission matters because a synthesis containing more than 150 references can appear comprehensive even when the underlying search process is unknown. Without a transparent method, readers cannot determine whether studies with weak, negative or null results were actively sought or whether the review disproportionately captured striking positive findings.
The concern is especially important in a rapidly developing field where studies differ dramatically in design. Some investigations use a small, single-center retrospective cohort, while others combine scans from multiple hospitals. A model trained and tested on images from the same institution may learn local scanning protocols, reconstruction settings or patient-selection patterns rather than biological signals associated with a mutation. This phenomenon, known as dataset shift or shortcut learning, can produce impressive internal accuracy that falls sharply when the model encounters patients from another hospital, scanner manufacturer or demographic group.
Sample size and validation strategy also affect how much confidence can be placed in a reported AUC. In a small dataset, a handful of correctly or incorrectly classified cases can substantially change the estimate. If researchers repeatedly adjust a model after examining test results, the nominal test set may no longer provide an unbiased measure of performance. External validation, ideally using data collected independently at different institutions and at a later time, is therefore essential. Prospective studies are even more informative because they test whether a model can operate under the conditions of real clinical care rather than only within a carefully assembled retrospective database.
The commentary argues that the review presented individual headline numbers without a structured quality assessment. A result with an AUC above 0.95 may sound dramatically more convincing than one near 0.64, but the number alone does not reveal whether the studies used comparable endpoints, reference standards or validation procedures. Genetic status might be established through different sequencing methods, while the definition of a nodule or mutation-positive case may vary between cohorts. A high score from a small, homogeneous sample cannot automatically be treated as stronger evidence than a moderate score from a large, diverse and externally validated study.
The authors also identify a mismatch between the review’s language and the aggregate evidence it cites. In its framing sections, the review describes AI as transformative or revolutionary in connecting imaging phenotypes with oncogenic driver genes. Yet the review also references a systematic analysis reporting a mean AUC of approximately 0.64 for predicting driver mutations from histopathological images, with EGFR prediction reaching about 0.79. Those figures suggest that the field has detected meaningful signals but has not yet demonstrated consistently reliable performance suitable for routine clinical decision-making. An average AUC near 0.64 indicates modest discrimination, and even a value around 0.79, while potentially useful, does not by itself establish clinical utility.
This distinction between detecting a statistical association and delivering a clinically useful test is central. A model may identify mutation-related image patterns without being accurate enough to replace tissue sampling or guide treatment independently. Lung cancer therapy can depend on the presence of specific genomic alterations, and a false negative could deny a patient a targeted treatment, while a false positive could lead clinicians toward an inappropriate therapy. Before imaging-based predictions can influence such decisions, researchers must establish calibration, reproducibility, clinical benefit and safety, as well as performance across different populations and healthcare systems.
The critique does not dismiss AI radiogenomics or the review’s value. Xue and Chen’s article maps a broad research landscape that includes CT feature extraction, multimodal data fusion, prognostic modeling, federated learning and regulatory considerations. Such a map can help newcomers understand how imaging, pathology, sequencing and machine learning are being combined. But Jopaul and Lingaraj argue that its usefulness would increase if the evidence were presented with a declared search strategy, a structured comparison of studies and an appraisal of bias and validation quality. They also call for conclusions whose confidence matches the field’s average results rather than its most spectacular individual experiments. The message is likely to resonate beyond lung cancer: AI can find patterns at extraordinary scale, but trustworthy medical science depends on showing how those patterns were discovered, how often they hold up and whether they improve care for real patients.
Cite this news
SCIENMAG. (August 27, 2026). Commentary examines AI links between lung nodule CT features and cancer-driving genes. https://scienmag.com/commentary-examines-ai-links-between-lung-nodule-ct-features-and-cancer-driving-genes/
SCIENMAG. "Commentary examines AI links between lung nodule CT features and cancer-driving genes." Scienmag, 27 August 2026, https://scienmag.com/commentary-examines-ai-links-between-lung-nodule-ct-features-and-cancer-driving-genes/. Accessed 27 August 2026.
SCIENMAG. "Commentary examines AI links between lung nodule CT features and cancer-driving genes." Scienmag. August 27, 2026. https://scienmag.com/commentary-examines-ai-links-between-lung-nodule-ct-features-and-cancer-driving-genes/

