Immunotherapy has transformed the treatment landscape for advanced lung cancer, offering durable responses in patients who once had few options. Yet a stubborn problem has shadowed this progress: only a subset of patients actually benefit, and clinicians have had no reliable way to know in advance who those patients will be. A new study published in the Journal of Cancer Research and Clinical Oncology by XiaoFang Yuan, Jing Shu, ShangYao Mo, YaPing Li and colleagues at Beijing Anzhen Nanchong Hospital of Capital Medical University and Nanchong Central Hospital addresses this gap with a machine learning model built on a broad panel of clinical, inflammatory and molecular variables, achieving predictive accuracy above 83 percent in both training and validation cohorts.
The research team enrolled 342 consecutive patients with advanced lung cancer who were receiving immunotherapy, randomly assigning them to training and validation sets. From each patient, the investigators collected baseline measurements spanning three complementary domains: standard clinical parameters, systemic inflammatory markers, and multi-omics features that capture the biological state of the tumor and its surrounding immune environment. The logic behind this breadth is straightforward. Response to checkpoint inhibitors is not governed by a single molecule but by an intricate interplay between tumor genetics, immune cell infiltration, and the overall physiological condition of the patient. By pooling variables from each of these layers, the researchers aimed to build a predictor that reflects the full biological complexity of the disease.
The first analytical step was to identify which of the many collected variables actually separated responders from non-responders. Univariate analysis flagged six factors with statistically significant differences between the two groups, all at P values below 0.05: tumor mutational burden, known as TMB; a composite TIME score reflecting the tumor immune microenvironment; the neutrophil-to-lymphocyte ratio, or NLR; the platelet-to-lymphocyte ratio, or PLR; serum lactate dehydrogenase, or LDH; and albumin. Multivariate logistic regression then confirmed that all six remained independent predictors when their effects were adjusted against one another, a crucial check ensuring that none of the associations was merely an artifact of correlation with another variable.
Each of these six predictors tells a distinct biological story. Tumor mutational burden measures the number of mutations carried by the tumor, and a higher burden generally means the cancer cells produce more abnormal proteins that the immune system can recognize as foreign, making checkpoint blockade more likely to unleash an effective attack. The TIME score, by contrast, probes the local battlefield within and around the tumor, characterizing the density and composition of immune cells in the microenvironment. The NLR and PLR capture systemic inflammation from a routine blood count; elevated ratios often signal a pro-tumor inflammatory state in which neutrophils and platelets suppress the antitumor activity of lymphocytes. LDH is a marker of tumor burden and tissue breakdown, frequently elevated in aggressive disease, while low albumin reflects poor nutritional status and systemic illness, both of which can blunt the immune response that immunotherapy depends on.
With six independent predictors in hand, the team constructed three machine learning models to integrate them: random forest, support vector machine, and logistic regression. Each approach handles the data differently. Logistic regression fits a linear relationship between the predictors and the probability of response, offering transparency but limited flexibility. A support vector machine finds the boundary that best separates responders from non-responders in a high-dimensional feature space, while a random forest builds hundreds of decision trees on random subsets of the data and averages their votes, a strategy that captures nonlinear interactions and is inherently resistant to overfitting when properly tuned.
The random forest emerged as the clear winner. In the training set it achieved an area under the receiver operating characteristic curve, or AUC, of 0.835 with a 95 percent confidence interval of 0.776 to 0.895, and in the validation set it delivered a nearly identical AUC of 0.832 with a confidence interval of 0.740 to 0.925. That consistency between cohorts is the strongest evidence that the model has learned genuine biological signal rather than statistical noise. Logistic regression performed respectably, with AUCs of 0.822 in training but a drop to 0.757 in validation, while the support vector machine lagged behind at 0.773 and 0.749 respectively. The tight confidence intervals and the minimal degradation from training to validation together suggest a robust and generalizable tool.
Performance metrics alone, however, do not tell clinicians why a model makes a particular prediction, and black-box medicine has rightly drawn skepticism. To open the box, the researchers applied SHAP analysis, a technique derived from cooperative game theory that quantifies each feature’s contribution to individual predictions. The SHAP ranking placed the neutrophil-to-lymphocyte ratio first, followed by the TIME score, the platelet-to-lymphocyte ratio, tumor mutational burden, LDH, and albumin. This ordering is provocative. It suggests that systemic inflammatory markers, easily obtained from a simple blood draw, carry at least as much predictive weight as tumor mutational burden, which requires expensive genomic sequencing. If confirmed in larger studies, that finding could make sophisticated response prediction far more accessible in settings where advanced molecular diagnostics are unavailable.
Beyond discrimination, the team evaluated calibration and clinical utility. Calibration curves showed good agreement between predicted probabilities and observed outcomes, meaning that when the model says a patient has, for example, a 70 percent chance of responding, roughly 70 percent of such patients actually do. Decision curve analysis, which compares the net benefit of acting on the model’s predictions against default strategies such as treating everyone or treating no one, confirmed high clinical net benefit across a broad range of threshold probabilities. These are the tests that separate a scientifically interesting classifier from one that can safely inform real treatment decisions, and the model passed both.
The practical implications are considerable. Patients with advanced lung cancer typically face a choice between immunotherapy, chemotherapy, combinations, or other targeted approaches, and each option carries costs, toxicities and opportunity costs. A patient predicted to respond strongly could proceed to immunotherapy with confidence, while a patient predicted to be resistant might be steered toward alternatives or into combination trials designed to overcome primary resistance. The authors position their model as an objective, quantitative decision-support tool for individualized immunotherapy management, a supplement to clinical judgment rather than a replacement for it.
Caveats remain. The study was conducted at a single institution in China, and validation in independent, multiethnic cohorts will be needed before widespread adoption. The TIME score, while powerful, requires specialized assessment of tumor samples that not all centers can provide routinely. And as an early-access, peer-reviewed publication, the article may undergo minor editorial revisions before the final version of record. Even so, the study exemplifies a growing trend in oncology: rather than searching for one perfect biomarker, researchers are integrating many modest signals across biological scales with machine learning to produce predictions that outperform any single measure. For a disease that kills more people worldwide than any other cancer, a validated tool that helps put each patient on the right therapy the first time is a development worth watching closely.
Subject of Research: A multiomics machine learning model for predicting immunotherapy response in advanced lung cancer
Article Title: A multiomics model for predicting immunotherapy response in advanced lung cancer
Article References: A multiomics model for predicting immunotherapy response in advanced lung cancer. (n.d.). https://doi.org/10.1007/s00432-026-06603-9
Image Credits: AI Generated
DOI: 10.1007/s00432-026-06603-9
Keywords: advanced lung cancer, immunotherapy, multi-omics integration, machine learning, random forest, tumor mutational burden, neutrophil-to-lymphocyte ratio, tumor immune microenvironment, SHAP analysis, predictive model, biomarkers, cancer immunotherapy
Cite Scienmag News
Nathaniel Bowman. (September 23, 2026). Machine Learning Model Predicts Which Lung Cancer Patients Will Respond to Immunotherapy. Scienmag. https://scienmag.com/machine-learning-model-predicts-which-lung-cancer-patients-will-respond-to-immunotherapy/
Nathaniel Bowman. "Machine Learning Model Predicts Which Lung Cancer Patients Will Respond to Immunotherapy." Scienmag, 23 September 2026, https://scienmag.com/machine-learning-model-predicts-which-lung-cancer-patients-will-respond-to-immunotherapy/. Accessed 23 September 2026.
Nathaniel Bowman. "Machine Learning Model Predicts Which Lung Cancer Patients Will Respond to Immunotherapy." Scienmag. September 23, 2026. https://scienmag.com/machine-learning-model-predicts-which-lung-cancer-patients-will-respond-to-immunotherapy/

