Machine learning models that predict how long patients survive after stomach cancer surgery are showing remarkable accuracy, and one of the simplest statistical approaches is proving surprisingly hard to beat. A new study published in BMC Cancer has developed and internally validated three machine learning-based prognostic models for patients undergoing radical gastrectomy—the surgical removal of the stomach along with its draining lymph nodes—and found that these tools can sort postoperative patients into dramatically different survival groups, with five-year survival rates ranging from 91 percent down to just 22 percent.
The research, led by Jun Wu, Luqing Zhou, Haidi Chang and Xuhui Liao of the Department of Pathology at The First Affiliated Hospital of Lishui University, Lishui People’s Hospital, in collaboration with Tian Qin and Junxuan Zhu of Guangzhou LBP Medical Technology Co., Ltd., addresses one of the most persistent challenges in gastric cancer care. Gastric cancer remains a leading cause of cancer-related mortality worldwide, and outcomes among patients who undergo curative-intent surgery vary enormously. Two patients with seemingly similar tumors can face radically different futures, and clinicians have long lacked reliable tools to distinguish between them at the individual level. Accurate prognostic prediction is essential for individualizing treatment intensity and surveillance strategies, and it is this gap that the team set out to close with computational methods.
The study took a retrospective approach, collecting detailed clinical and pathological data from 200 patients who had undergone radical gastrectomy at the authors’ institution. The dataset was deliberately multidimensional. It incorporated variables including age, sex, TNM stage—the standard tumor-node-metastasis staging system—Lauren classification, which describes the histological growth pattern of the tumor, differentiation grade, tumor size, lymphovascular invasion, perineural invasion, and treatment information. Lymphovascular invasion, often abbreviated LVI, refers to tumor cells invading blood or lymphatic vessels, while perineural invasion, or PNI, describes cancer spreading along nerve sheaths. Both are recognized histological markers of aggressive tumor biology and were included as candidate predictors alongside traditional staging variables.
Three distinct machine learning models were constructed to predict survival outcomes. The first was the Cox proportional hazards model, or CoxPH, a statistical workhorse that estimates the effect of multiple variables on the hazard—or instantaneous risk—of death at any given time, assuming that these effects multiply the baseline risk in a proportional fashion. The second was the Random Survival Forest, or RSF, an ensemble method that grows large numbers of survival trees on randomly sampled subsets of the data and averages their predictions, an approach capable of capturing complex nonlinear relationships and interactions between variables without explicit specification. The third was Gradient Boosting Survival Analysis, or GBSA, which builds predictive power sequentially, with each new model trained to correct the residual errors of its predecessors, producing a strong composite learner from many weak ones.
Model performance was rigorously evaluated using 5-fold stratified cross-validation, a technique in which the data are divided into five parts, with each part serving once as a test set while the remaining four are used for training. Stratification ensures that each fold preserves the same proportion of events, which is critical when analyzing survival data. Two complementary metrics were used: the concordance index, or C-index, which measures how well a model ranks patients by risk—0.5 representing random guessing and 1.0 perfect discrimination—and the time-dependent area under the curve, or AUC, which quantifies predictive accuracy at specific time points after surgery.
The cohort itself underscored the severity of the disease. The 200 patients were 79.0 percent male, with a mean age of 65.4 years and a standard deviation of 10.7 years. Over a median follow-up of 64.2 months—more than five years—91 patients, or 45.5 percent of the cohort, died. This substantial event rate provided the statistical power needed to train and evaluate the models meaningfully.
Within the multivariate Cox regression framework, two factors emerged as independent prognostic determinants. Age carried a hazard ratio of 1.097 with a 95 percent confidence interval of 1.068 to 1.145 and a P value below 0.001, meaning each additional year of age was associated with roughly a 10 percent increase in the hazard of death. N stage—the extent of lymph node involvement—produced a hazard ratio of 1.498 with a 95 percent confidence interval of 1.100 to 2.322 and a P value of 0.035, confirming that nodal spread remains a powerful driver of postoperative mortality. These findings align with decades of clinical observation but also serve as the anchors of the predictive models.
When the three machine learning approaches were compared head-to-head, the results carried an instructive lesson. The Cox proportional hazards model achieved the highest cross-validated C-index, at 0.733 plus or minus 0.096, edging out its more elaborate competitors. This is a noteworthy outcome in a field often captivated by complex algorithms: a well-specified classical model, given a carefully curated set of clinically meaningful variables, can match or exceed the discrimination of ensemble methods—while remaining far more interpretable to practicing clinicians. The finding does not diminish the value of the other models, however. The Random Survival Forest demonstrated exceptional time-dependent AUC values of 0.911 at 12 months, 0.933 at 36 months, and 0.950 at 60 months, indicating that its ability to distinguish survivors from non-survivors actually improved the longer the follow-up extended, a property that could prove valuable for long-term surveillance planning.
Perhaps the most clinically striking result came from risk stratification based on the CoxPH model. Using tertiles of the model-derived risk score, the researchers divided patients into three distinct prognostic groups. The low-risk group, comprising 67 patients, achieved a five-year survival rate of 91.0 percent. The intermediate-risk group, with 66 patients, had a five-year survival of 57.6 percent. The high-risk group, also 67 patients, saw only 22.4 percent of its members alive at five years. The gulf between the top and bottom tiers—nearly 69 percentage points—demonstrates the real-world consequences of prognostic heterogeneity and the potential of algorithmic stratification to expose it. A patient placed in the high-risk group might reasonably be candidates for intensified adjuvant therapy and closer surveillance, while a low-risk patient could potentially be spared unnecessary treatment burden.
The authors emphasize that their models effectively predicted survival outcomes after radical gastrectomy and that the CoxPH model provided reliable risk stratification that may assist prognostic risk assessment in clinical practice. At the same time, they are careful to frame these findings as preliminary in an important respect: the study represents internal validation only. The models were developed and tested within the same single-institution cohort, albeit with cross-validation guarding against overfitting. Before any of these tools can inform bedside decisions, they must undergo external validation in independent cohorts from other institutions and populations—a standard requirement in the clinical prediction model literature, where optimism bias in internally validated models is well documented.
The study was approved by the Ethics Committee of Lishui People’s Hospital, and because of its retrospective design, the requirement for informed consent was waived, with all patient data de-identified prior to analysis. The research received no external funding, and the authors declare no competing interests. The work is published open access under a Creative Commons Attribution 4.0 International License.
The implications of this research extend beyond a single cancer type. As machine learning continues to permeate oncology, studies like this one offer a template for how to integrate it responsibly: define a clinically relevant question, assemble a multidimensional but interpretable variable set, compare classical and modern algorithms under honest cross-validation, and translate model outputs into clinically actionable risk categories. For gastric cancer patients, the promise is tangible—an algorithm that, at the time of surgery, can tell whether five years of life lie ahead with 91 percent confidence or 22 percent, and guide the intensity of follow-up and adjuvant treatment accordingly. Pending external validation, that promise is one step closer to reality.
Cite Scienmag News
Nathaniel Bowman. (September 10, 2026). Machine learning models predict outcomes after radical gastrectomy in multicenter study. Scienmag. https://scienmag.com/machine-learning-models-predict-outcomes-after-radical-gastrectomy-in-multicenter-study/
Nathaniel Bowman. "Machine learning models predict outcomes after radical gastrectomy in multicenter study." Scienmag, 10 September 2026, https://scienmag.com/machine-learning-models-predict-outcomes-after-radical-gastrectomy-in-multicenter-study/. Accessed 10 September 2026.
Nathaniel Bowman. "Machine learning models predict outcomes after radical gastrectomy in multicenter study." Scienmag. September 10, 2026. https://scienmag.com/machine-learning-models-predict-outcomes-after-radical-gastrectomy-in-multicenter-study/








