For millions of people living with diabetes, the most feared complication is not the disease itself but the quiet, progressive damage it inflicts on the retina. Proliferative diabetic retinopathy, the advanced stage of this damage, is a leading preventable cause of blindness in working-age adults between 20 and 74 years old. Worldwide, roughly 35.4 percent of diabetic patients show some degree of retinopathy, and about 7.5 percent have already progressed to the proliferative stage. In Asia, the picture is even starker: among people with type 2 diabetes, one in four has retinopathy, and proliferative disease accounts for 15 percent of cases. Once the disease reaches this stage, many patients require a delicate operation called pars plana vitrectomy to clear bleeding, remove scar tissue, and relieve traction on the retina.
The surgery can be sight-saving, but outcomes vary dramatically from patient to patient. Some regain useful vision and keep it for years, while others deteriorate despite technically successful operations. A new study from Shanxi Eye Hospital in China, published in Health Science Reports, tackles this uncertainty head-on by building a machine learning model that predicts, before and during surgery, which patients are likely to end up with poor long-term vision. The work, led by Xiaolu Wang and colleagues, draws on 609 eyes from 609 patients who underwent their first vitrectomy for proliferative diabetic retinopathy between January 2022 and January 2025, each followed for at least twelve months after the operation.
The research team defined a good visual prognosis as an improvement of at least 0.3 units on the logarithm of the minimum angle of resolution scale, a standard measure of best-corrected visual acuity, by the final follow-up visit. Patients whose vision worsened by more than 0.3 units, or who failed to improve by that margin, were classified as having a poor prognosis. This twelve-month endpoint matters because previous studies have shown that visual acuity tends to stabilize after the first year following vitrectomy, making it a meaningful window for judging the true success of the intervention. The team even assigned numerical values to eyes that could only count fingers, perceive hand motion, or detect light, ensuring that even the most severely affected patients could be graded consistently.
Before any modeling began, the researchers confronted a challenge familiar to anyone working with real-world medical records: missing data. Of the 52 candidate variables collected from the hospital’s electronic medical record system, body mass index had the highest proportion of missing values at 19.2 percent. Rather than discarding incomplete records, the team used multiple imputation, a statistical technique that fills in gaps by drawing on the relationships among the observed variables. They then split the cohort randomly into a training set of 487 patients and a validation set of 122, verifying through descriptive statistics that the two groups were comparable across demographic, surgical, and biochemical characteristics.
Variable selection proceeded in two stages. First, the team applied least absolute shrinkage and selection operator regression, a technique that penalizes model complexity and drives the coefficients of uninformative variables toward zero. As the penalty parameter converged to 0.02910, seven candidate predictors survived the cut: age, renal insufficiency, preoperative iris neovascularization, the type of tamponade used during surgery, serum alkaline phosphatase, indirect bilirubin, and serum gamma-glutamyl transferase. These candidates then entered a binary logistic regression, where four emerged as statistically significant. Renal insufficiency carried an odds ratio of 6.932, meaning patients with impaired kidney function faced nearly seven times the odds of a poor visual outcome. Preoperative iris neovascularization was even more ominous, with an odds ratio of 7.674. Silicone oil tamponade doubled the risk at an odds ratio of 2.799, while higher indirect bilirubin levels appeared protective, with each unit increase associated with roughly a 10 percent reduction in the odds of poor prognosis.
With these predictors in hand, the researchers trained six different machine learning algorithms: decision tree, random forest, support vector machine, multilayer perceptron, logistic regression, and Light Gradient Boosting Machine, known as LightGBM. Each model was tuned through grid search and ten-fold cross-validation, a procedure that repeatedly partitions the training data to ensure the model’s performance is not a fluke of any single split. On the held-out validation set, the differences between algorithms became clear. The support vector machine, despite respectable training performance, saw its area under the receiver operating characteristic curve collapse to 0.616 on unseen data, a classic signature of overfitting. The multilayer perceptron fared somewhat better at 0.753 but still showed a troubling gap between training and test performance. Random forest reached 0.767 but suffered from weak recall and accuracy, while decision tree and logistic regression posted areas under the curve of 0.717 and 0.783 respectively.
LightGBM emerged as the clear winner, achieving an area under the curve of 0.786 on the validation set along with the best recall score during cross-validation at 0.979. Calibration curves confirmed that the model’s predicted probabilities tracked observed outcomes well in both the training and validation cohorts. Decision curve analysis, which quantifies the clinical net benefit of using a model at various risk thresholds, showed that LightGBM outperformed both a strategy of treating all patients and one of treating none, at least across low-to-moderate risk thresholds. The authors candidly note that the model’s net benefit fluctuated in the moderate threshold range, approaching zero at certain points, which they attribute to sparse sample distribution in that interval or reduced calibration, a reminder that even the best models have limits.
Perhaps the most clinically valuable contribution is the study’s effort to open the black box. Machine learning models are often criticized for being inscrutable, making it hard for physicians to understand why a particular prediction was made. To address this, the team applied SHAP, or SHapley Additive exPlanations, a framework borrowed from game theory that assigns each feature a contribution value for every individual prediction. The SHAP analysis confirmed the regression findings: renal insufficiency, silicone oil tamponade, and preoperative iris neovascularization all pushed predictions toward poor prognosis, while higher indirect bilirubin pushed toward good prognosis. The researchers also built a web-based calculator that lets clinicians enter a patient’s values and receive an instant estimate of poor-prognosis risk, translating the algorithm into a practical bedside tool.
The biological stories behind the risk factors are compelling. Renal insufficiency reflects systemic vascular vulnerability: reduced kidney function impairs the clearance of uremic compounds that amplify inflammation and oxidative stress in the retina, raising levels of vascular endothelial growth factor, the very molecule that drives abnormal blood vessel growth in diabetic eye disease. Iris neovascularization, present in about 65 percent of proliferative diabetic retinopathy patients according to prior research, signals severe retinal ischemia; roughly 20 percent of these patients progress to neovascular glaucoma, a painful and often blinding condition. The association between silicone oil tamponade and poor outcomes requires careful interpretation, since surgeons reserve oil for the most complex cases such as tractional retinal detachment, meaning the oil may be a marker of disease severity rather than a direct cause of visual decline, though the authors note it can also dissolve lipophilic macular pigments and exert chronic mechanical pressure on the retina.
The protective role of indirect bilirubin is the study’s most novel finding. Bilirubin, long dismissed as merely a waste product of hemoglobin breakdown, is now recognized as one of the body’s most potent endogenous antioxidants, capable of neutralizing free radicals and suppressing oxidative reactions. Slightly elevated levels may reduce intracellular oxidative stress, improve insulin sensitivity, and regulate glucose metabolism, all of which could slow diabetic complications. The authors are appropriately cautious, calling for future analyses of nonlinear dose-response relationships, mediation by anti-inflammatory biomarkers, and sensitivity analyses to rule out confounding by outlier values. They also acknowledge the study’s main limitations: the data came from a single hospital’s retrospective records, and the model was validated only internally. Multi-center prospective studies will be needed before the tool can be widely deployed. Still, the work points toward a future in which ophthalmologists can identify high-risk patients before they reach the operating table, intervene earlier on kidney function and neovascular disease, and tailor follow-up care to those who need it most, potentially preserving sight for thousands of people who would otherwise face preventable blindness.
Subject of Research: A machine learning prediction model for long-term visual outcomes after vitrectomy in proliferative diabetic retinopathy
Article Title: Construction and Validation of a Long‐Term Visual Prognosis Prediction Model for Proliferative Diabetic Retinopathy After Vitrectomy Based on Machine Learning
Article References: Wang, X., Shi, J., Gao, Y., Li, T., Han, X., Guo, L., Jia, T., & Wang, Y. (2026). Construction and Validation of a Long‐Term Visual Prognosis Prediction Model for Proliferative Diabetic Retinopathy After Vitrectomy Based on Machine Learning. Endocrinology, Diabetes & Metabolism, 9(5), Article e70314. https://doi.org/10.1002/edm2.70314
Image Credits: AI Generated
DOI: 10.1002/edm2.70314
Keywords: diabetic retinopathy, vitrectomy, machine learning, LightGBM, visual prognosis, SHAP, iris neovascularization, renal insufficiency, indirect bilirubin, silicone oil tamponade, predictive modeling, ophthalmology
Cite Scienmag News
Ophelia Keating. (September 25, 2026). Machine Learning Model Predicts Long-Term Vision After Surgery for Advanced Diabetic Eye Disease. Scienmag. https://scienmag.com/machine-learning-model-predicts-long-term-vision-after-surgery-for-advanced-diabetic-eye-disease/
Ophelia Keating. "Machine Learning Model Predicts Long-Term Vision After Surgery for Advanced Diabetic Eye Disease." Scienmag, 25 September 2026, https://scienmag.com/machine-learning-model-predicts-long-term-vision-after-surgery-for-advanced-diabetic-eye-disease/. Accessed 25 September 2026.
Ophelia Keating. "Machine Learning Model Predicts Long-Term Vision After Surgery for Advanced Diabetic Eye Disease." Scienmag. September 25, 2026. https://scienmag.com/machine-learning-model-predicts-long-term-vision-after-surgery-for-advanced-diabetic-eye-disease/

