Every year, millions of patients undergo whole-body bone scintigraphy, a nuclear medicine scan that reveals hotspots of abnormal bone activity caused by cancer metastases, fractures, or benign disease. Deciding who truly needs the scan, and how urgently, has long depended on clinical intuition. Now a team at Dongzhimen Hospital of Beijing University of Traditional Chinese Medicine has built machine learning models that attempt to forecast the scan’s diagnostic outcome before a single image is taken, offering clinicians a statistical preview of what the camera will find. The study, published in BMC Medical Imaging, compares two very different modeling philosophies and delivers a sober lesson about how the passage of time can quietly undermine even well-built predictive tools.
The research team, led by Kuo Ma, Ziyang Du, and corresponding author Peilin Wu, assembled a retrospective cohort of 3,969 patients scheduled for whole-body bone scintigraphy. After excluding eleven cases with missing values, 3,958 patients remained. Each patient was assigned to one of three diagnostic categories: bone metastasis, fracture, or benign bone disease. The goal was ambitious in its simplicity: using only eleven clinical features available before the scan, could an algorithm reliably predict which of the three categories a patient would ultimately fall into? Such a tool would give physicians an adjunctive reference for risk stratification, potentially flagging high-risk patients for expedited imaging or additional workup.
The methodological design is where the study distinguishes itself from much of the medical machine learning literature. Rather than splitting patients randomly into training and test sets and calling it a day, the authors built a dual validation architecture. The entire dataset was first divided chronologically, by examination date, into a temporal training set and a temporal validation set at an 8:2 ratio. The temporal training set was then further partitioned randomly into a random training set and a random test set, again at 8:2. This two-tier structure allowed the team to measure not just how well the models performed on patients resembling their training data, but how they held up on patients examined later in time, a far more realistic simulation of clinical deployment.
Two algorithms were pitted against each other. The first, multinomial logistic regression, is a classical statistical workhorse that extends binary logistic regression to multiple classes, estimating the probability of each diagnostic category as a function of the input features through a set of linear equations. Its virtues are transparency and stability: every coefficient can be inspected, and the model’s behavior is fully determined by its parameters. The second, XGBoost, is a gradient-boosted decision tree ensemble that has dominated machine learning competitions for a decade. It builds hundreds of shallow trees sequentially, each one trained to correct the residual errors of its predecessors, capturing nonlinear relationships and feature interactions that a linear model cannot. XGBoost typically wins on raw predictive power but offers less interpretability and more opportunities for overfitting.
Hyperparameters for both models were optimized on the random training set using grid search combined with 5-fold cross-validation, a procedure that repeatedly divides the training data into five folds, trains on four, and validates on the fifth, averaging results to select the most robust parameter combination. With optimal parameters locked in, the models were retrained on both the random training set and the temporal training set, and final performance was evaluated on the random test set and the temporal validation set respectively. The evaluation went well beyond accuracy, incorporating the F1 score, the area under the receiver operating characteristic curve (AUC), the Brier score, calibration slope, and net benefit derived from decision curve analysis, a metric that asks whether acting on the model’s predictions would help patients more than it harms them.
On the random test set, both models delivered encouraging three-class performance, and, notably, no statistically significant differences emerged between the logistic regression and XGBoost across any evaluated metric, with all p-values exceeding 0.05. This parity is itself an interesting finding: the flexible, nonlinear XGBoost offered no measurable advantage over its transparent counterpart, suggesting that the eleven clinical features carry most of their predictive signal in relationships a linear model can capture. Feature importance analysis, supported by SHAP explanation plots in the supplementary material, identified history of cancer and history of trauma as the two most important predictors, which aligns with clinical expectation that prior malignancy and prior injury are the dominant forces shaping what a bone scan will reveal.
The story grew more complicated when the models faced the temporal validation set. Pre-calibration performance declined to some extent, a phenomenon the authors attribute to temporal drift, the gradual shift in patient characteristics, disease prevalence, and clinical practice that occurs as time passes. Decision curve analysis, which had provided exploratory evidence of potential net benefit on the random test set, showed that this potential benefit diminished in temporal validation. In other words, a model that looked clinically useful when tested on contemporaneous patients became less trustworthy when asked to make predictions for patients examined months later, precisely the situation any deployed model would face.
The team’s response to this drift is arguably the study’s most technically interesting contribution. They applied post-hoc calibration, a statistical correction that remaps the model’s raw probability outputs so they better reflect true observed frequencies. Platt scaling, which fits a simple logistic transformation to the model outputs, was used for the logistic regression, while isotonic regression, a more flexible non-parametric method that fits a monotonic step function, was applied to XGBoost. Crucially, the calibrators were fitted via 5-fold cross-validation on the temporal training set and then applied independently to the temporal validation set, avoiding the leakage that plagues many calibration studies. The results were clear: post-calibration probability reliability improved in the temporal cohort, even as raw discrimination did not recover. The models’ rankings of patients may have stayed roughly intact, but the probabilities attached to those rankings became honest again.
The authors are careful, almost unusually so, about what their findings do and do not establish. They state that calibration may be useful for improving probability reliability, but that the results do not establish temporal generalizability, discrimination improvement, or readiness for clinical deployment, and they emphasize that external validation is required before any real-world use. This restraint matters in a field where prediction models are frequently oversold. A model that predicts diagnostic categories before a scan could, in principle, help prioritize urgent cases, reduce unnecessary imaging, or guide the choice of additional tests, but only if its probabilities remain trustworthy across time and across hospitals, questions this single-center retrospective study cannot fully answer.
The study also carries practical lessons for anyone building clinical prediction tools. First, temporal validation should be treated as a standard requirement, not an optional extra, because random splits systematically flatter model performance. Second, calibration deserves the same attention as discrimination; a model with a respectable AUC can still produce probabilities so miscalibrated that they mislead clinical decisions. Third, the equivalence of logistic regression and XGBoost here is a reminder that simpler models remain competitive in many clinical prediction tasks, particularly when the feature set is small and largely categorical. The work, approved by the Institutional Ethics Committee of Dongzhimen Hospital and conducted under the Declaration of Helsinki with waived informed consent for anonymized retrospective data, received no specific funding and reports no competing interests. As hospitals increasingly experiment with pre-scan triage algorithms, this study offers both a template for rigorous evaluation and a warning: the future arrives one day at a time, and models must be recalibrated to meet it.
Subject of Research: Machine learning prediction of pre-scan diagnostic categories in whole-body bone scintigraphy
Article Title: Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost
Article References: Ma, K., Du, Z., & Wu, P. (2026). Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost. BMC Medical Imaging. https://doi.org/10.1186/s12880-026-02848-5
Image Credits: AI Generated
DOI: 10.1186/s12880-026-02848-5
Keywords: bone scintigraphy, machine learning, XGBoost, multinomial logistic regression, prediction model, calibration, temporal validation, bone metastasis, fracture, decision curve analysis, risk stratification, nuclear medicine
Cite Scienmag News
Ophelia Keating. (October 6, 2026). AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy. Scienmag. https://scienmag.com/ai-models-predict-bone-scan-outcomes-before-imaging-but-time-may-erode-their-accuracy/
Ophelia Keating. "AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy." Scienmag, 6 October 2026, https://scienmag.com/ai-models-predict-bone-scan-outcomes-before-imaging-but-time-may-erode-their-accuracy/. Accessed 6 October 2026.
Ophelia Keating. "AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy." Scienmag. October 6, 2026. https://scienmag.com/ai-models-predict-bone-scan-outcomes-before-imaging-but-time-may-erode-their-accuracy/

