When surgeons remove a large lung tumor, the operation itself is only the beginning of a long and anxious vigil. For patients with locally advanced non-small cell lung cancer—specifically tumors classified as T3 or T4 under the tumor-node-metastasis staging system—even a technically complete resection with clear margins, known as R0 resection, does not guarantee that the disease is gone. A substantial fraction of these patients will experience a local recurrence at or near the surgical site within a few years, and clinicians currently have limited tools to identify, before it happens, which individuals are most at risk. A new study published in BMC Medical Imaging offers a data-driven approach to that problem, combining information from preoperative CT scans, postoperative pathology slides, and clinical variables into a single predictive model designed to estimate each patient’s two-year risk of local recurrence.
The research, led by Xinyu Li and Guangming Lu of Nanjing Medical University together with colleagues at Jinling Hospital, The Affiliated Changsha Central Hospital, and The First People’s Hospital of Chenzhou, is a two-center retrospective study built on patients who underwent complete resection of pT3–4N0–2M0 non-small cell lung cancer. The development cohort, drawn from a center treating consecutive patients between January 2015 and December 2021, comprised 135 patients, of whom 30 experienced a documented local recurrence as their first event. An entirely separate external center contributed 31 patients treated between January 2018 and December 2021, with 5 local-recurrence events, providing an independent test of the model’s generalizability. The team analyzed three complementary data types for each patient: preoperative contrast-enhanced computed tomography images, postoperative hematoxylin-and-eosin-stained whole-slide histopathology images, and routine clinicopathological variables such as tumor characteristics recorded after surgery.
What distinguishes the work from many prior artificial intelligence studies in oncology is its rigorous handling of a statistical subtlety that is often ignored: competing risks. In this patient population, some individuals die or develop distant metastases before a local recurrence is ever observed, which means those events preclude the outcome of interest. Standard survival models such as ordinary Cox regression can produce distorted risk estimates in this setting. The researchers instead used penalized Fine–Gray models, a framework specifically designed for competing-risk data, in which distant-first recurrence and death before local recurrence were treated as competing events. The primary outcome was defined precisely as the time from surgery to the first documented local recurrence, and the models were evaluated using time-dependent area under the curve values that account for these competing risks, along with Brier scores and calibration assessments at the two-year mark.
The imaging arm of the pipeline relied on radiomics, the high-throughput extraction of quantitative features from medical images. The team segmented both the intratumoral region—the tumor itself—and a peritumoral ring extending three millimeters beyond the tumor boundary on preoperative contrast-enhanced CT scans. This choice reflects a growing recognition in oncologic imaging that the tissue immediately surrounding a tumor, with its infiltrating immune cells, stromal changes, and early invasion, carries prognostic information that the tumor core alone does not. Radiomic features, quantifying properties such as texture heterogeneity, intensity distributions, and spatial patterns within each volume of interest, were filtered and selected using methods including the least absolute shrinkage and selection operator to prevent overfitting, and feature stability was assessed through intraclass correlation coefficients consistent with the Image Biomarker Standardization Initiative.
On the pathology side, the researchers trained a convolutional neural network on whole-slide images using a weakly supervised strategy. Rather than requiring pathologists to laboriously annotate which microscopic regions harbor prognostically important features—an expensive and inconsistent process—the network learned from slide-level labels alone, aggregating information across thousands of image patches to produce what the authors call pathomics features: numerical descriptors of the tissue’s cellular and architectural landscape. This pathology deep-learning model, referred to as Path-DL, was trained separately using a fixed 7:3 patient-level split and, critically, was not retrained within the cross-validation folds, a design decision that guards against the subtle information leakage that has undermined many published machine-learning studies in medicine. Gradient-weighted class activation mapping, or Grad-CAM, provided a way to visualize which regions of the slides the network attended to, offering pathologists a window into the model’s reasoning.
The final multimodal model fused the clinicopathological, intratumoral radiomics, peritumoral radiomics, and pathomics components at the score level, allowing each modality to contribute its own estimate of risk that could then be combined. The results, obtained through repeated nested five-fold cross-validation in the development cohort, tell a clear story about where predictive power resides. The clinicopathological model alone achieved a two-year area under the curve of just 0.492—essentially no better than a coin flip—underscoring how little conventional variables reveal about local recurrence risk in this population. Intratumoral and peritumoral radiomics performed meaningfully better, with AUCs of 0.689 and 0.694 respectively. The pathology deep-learning model reached 0.804, and the full multimodal fusion model topped the field at 0.827, with a 95 percent confidence interval of 0.689 to 0.936.
External validation, the true test of any predictive model, painted a more cautious but still encouraging picture. In the 31-patient external cohort, the fusion model achieved a two-year AUC of 0.683 (95 percent CI, 0.411–0.917), ahead of the clinicopathological model’s 0.310 and modestly above the radiomics models, which scored 0.605 and 0.616. The wide confidence intervals, an unavoidable consequence of the small external sample and its five events, temper any claim of proven clinical utility, but the direction of the results is consistent with the internal findings: the multimodal approach carries information that standard clinical assessment lacks. Formal paired comparisons quantified the incremental value of fusion over the pathology model alone. The paired difference in AUC was 0.023 (95 percent CI, −0.097 to 0.062) internally and 0.064 (−0.270 to 0.413) externally, while differences in Brier scores, which capture both discrimination and calibration, were 0.001 in both settings with confidence intervals straddling zero. In other words, the fusion model was not statistically superior to the pathology deep-learning model on its own—a finding the authors present transparently rather than overselling.
The implications for postoperative management of locally advanced lung cancer are nonetheless significant. Patients with resected pT3–4 non-small cell lung cancer currently receive relatively uniform surveillance and adjuvant treatment decisions, guided largely by stage and nodal status. A validated tool that stratifies local recurrence risk from data that already exist in the medical record—preoperative CT scans and the pathology slides prepared after surgery—could, if confirmed in larger prospective studies, allow clinicians to intensify follow-up imaging, consider adjuvant radiotherapy, or enroll high-risk patients in clinical trials while sparing lower-risk individuals unnecessary intervention. The inclusion of competing-risk methodology also means the model’s outputs are expressed as cumulative incidence of actual local recurrence, the quantity that matters in the clinic, rather than an inflated hazard-based surrogate.
The study’s transparency about its own limitations is notable and aligns with reporting standards such as the Checklist for Artificial Intelligence in Medical Imaging. Sample size is the most obvious constraint: 135 development patients with 30 events is modest by machine-learning standards, and 31 external patients cannot establish generalizability with statistical confidence. The retrospective design introduces the usual risks of selection bias, although the use of consecutive patients at both centers mitigates this. The models were locked before external evaluation, and the study was approved by the Institutional Review Board of Jinling Hospital with informed consent waived owing to the retrospective design, conducted in accordance with the Declaration of Helsinki. Shapley additive explanations, or SHAP, were used to interpret feature contributions, and decision curve analysis was employed to assess the clinical net benefit of the models across a range of risk thresholds—techniques that move the work beyond raw accuracy metrics toward genuine clinical decision support.
The research was supported by the Science and Technology Innovation 2030-Major Projects (grant 2020AAA0109500) and the General Program of the National Natural Science Foundation of China (grant 82371958), with technical support from the Deepwise multimodal research platform. The authors, including Yu Zong, Changsheng Zhou, Yang Cao, Zhen Zhou, Jianrui Li, Zhiyuan Sun, Xiaoqing Cheng, and Liying Wang, declare no competing interests. The article is published open access under a Creative Commons Attribution 4.0 license, with the accepted manuscript shared early under a citable permanent DOI ahead of the final version of record.
For the field of AI-assisted oncology, the study is a case study in methodological discipline: nested cross-validation for internal evaluation, a frozen upstream network to prevent data leakage, competing-risk statistics matched to the clinical question, external validation with locked models, and honest reporting of null findings in paired comparisons. It suggests that the microscopic world captured on a postoperative slide, read by a weakly supervised neural network, may hold more prognostic signal for local recurrence than anything clinicians currently measure—and that fusing it with imaging-based tumor and microenvironment features pushes performance modestly further. The next step, which the authors’ careful framing implicitly calls for, is prospective validation in larger, multi-institutional cohorts before such a model can guide real-world decisions for the thousands of patients who each year face the uncertainty of life after surgery for locally advanced lung cancer.
Cite Scienmag News
Nathaniel Bowman. (September 10, 2026). Radiomics-pathomics model predicts local recurrence in T3–4 lung cancer. Scienmag. https://scienmag.com/radiomics-pathomics-model-predicts-local-recurrence-in-t3-4-lung-cancer/
Nathaniel Bowman. "Radiomics-pathomics model predicts local recurrence in T3–4 lung cancer." Scienmag, 10 September 2026, https://scienmag.com/radiomics-pathomics-model-predicts-local-recurrence-in-t3-4-lung-cancer/. Accessed 10 September 2026.
Nathaniel Bowman. "Radiomics-pathomics model predicts local recurrence in T3–4 lung cancer." Scienmag. September 10, 2026. https://scienmag.com/radiomics-pathomics-model-predicts-local-recurrence-in-t3-4-lung-cancer/

