Tuesday, October 6, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy

October 6, 2026
in Medicine
Ophelia Keating
By Ophelia Keating Scienmag Editorial Profile - Health Services Research
Reading Time: 5 mins read
0
AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy

AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every year, millions of patients undergo whole-body bone scintigraphy, a nuclear medicine scan that reveals hotspots of abnormal bone activity caused by cancer metastases, fractures, or benign disease. Deciding who truly needs the scan, and how urgently, has long depended on clinical intuition. Now a team at Dongzhimen Hospital of Beijing University of Traditional Chinese Medicine has built machine learning models that attempt to forecast the scan’s diagnostic outcome before a single image is taken, offering clinicians a statistical preview of what the camera will find. The study, published in BMC Medical Imaging, compares two very different modeling philosophies and delivers a sober lesson about how the passage of time can quietly undermine even well-built predictive tools.

The research team, led by Kuo Ma, Ziyang Du, and corresponding author Peilin Wu, assembled a retrospective cohort of 3,969 patients scheduled for whole-body bone scintigraphy. After excluding eleven cases with missing values, 3,958 patients remained. Each patient was assigned to one of three diagnostic categories: bone metastasis, fracture, or benign bone disease. The goal was ambitious in its simplicity: using only eleven clinical features available before the scan, could an algorithm reliably predict which of the three categories a patient would ultimately fall into? Such a tool would give physicians an adjunctive reference for risk stratification, potentially flagging high-risk patients for expedited imaging or additional workup.

The methodological design is where the study distinguishes itself from much of the medical machine learning literature. Rather than splitting patients randomly into training and test sets and calling it a day, the authors built a dual validation architecture. The entire dataset was first divided chronologically, by examination date, into a temporal training set and a temporal validation set at an 8:2 ratio. The temporal training set was then further partitioned randomly into a random training set and a random test set, again at 8:2. This two-tier structure allowed the team to measure not just how well the models performed on patients resembling their training data, but how they held up on patients examined later in time, a far more realistic simulation of clinical deployment.

Two algorithms were pitted against each other. The first, multinomial logistic regression, is a classical statistical workhorse that extends binary logistic regression to multiple classes, estimating the probability of each diagnostic category as a function of the input features through a set of linear equations. Its virtues are transparency and stability: every coefficient can be inspected, and the model’s behavior is fully determined by its parameters. The second, XGBoost, is a gradient-boosted decision tree ensemble that has dominated machine learning competitions for a decade. It builds hundreds of shallow trees sequentially, each one trained to correct the residual errors of its predecessors, capturing nonlinear relationships and feature interactions that a linear model cannot. XGBoost typically wins on raw predictive power but offers less interpretability and more opportunities for overfitting.

Hyperparameters for both models were optimized on the random training set using grid search combined with 5-fold cross-validation, a procedure that repeatedly divides the training data into five folds, trains on four, and validates on the fifth, averaging results to select the most robust parameter combination. With optimal parameters locked in, the models were retrained on both the random training set and the temporal training set, and final performance was evaluated on the random test set and the temporal validation set respectively. The evaluation went well beyond accuracy, incorporating the F1 score, the area under the receiver operating characteristic curve (AUC), the Brier score, calibration slope, and net benefit derived from decision curve analysis, a metric that asks whether acting on the model’s predictions would help patients more than it harms them.

On the random test set, both models delivered encouraging three-class performance, and, notably, no statistically significant differences emerged between the logistic regression and XGBoost across any evaluated metric, with all p-values exceeding 0.05. This parity is itself an interesting finding: the flexible, nonlinear XGBoost offered no measurable advantage over its transparent counterpart, suggesting that the eleven clinical features carry most of their predictive signal in relationships a linear model can capture. Feature importance analysis, supported by SHAP explanation plots in the supplementary material, identified history of cancer and history of trauma as the two most important predictors, which aligns with clinical expectation that prior malignancy and prior injury are the dominant forces shaping what a bone scan will reveal.

The story grew more complicated when the models faced the temporal validation set. Pre-calibration performance declined to some extent, a phenomenon the authors attribute to temporal drift, the gradual shift in patient characteristics, disease prevalence, and clinical practice that occurs as time passes. Decision curve analysis, which had provided exploratory evidence of potential net benefit on the random test set, showed that this potential benefit diminished in temporal validation. In other words, a model that looked clinically useful when tested on contemporaneous patients became less trustworthy when asked to make predictions for patients examined months later, precisely the situation any deployed model would face.

The team’s response to this drift is arguably the study’s most technically interesting contribution. They applied post-hoc calibration, a statistical correction that remaps the model’s raw probability outputs so they better reflect true observed frequencies. Platt scaling, which fits a simple logistic transformation to the model outputs, was used for the logistic regression, while isotonic regression, a more flexible non-parametric method that fits a monotonic step function, was applied to XGBoost. Crucially, the calibrators were fitted via 5-fold cross-validation on the temporal training set and then applied independently to the temporal validation set, avoiding the leakage that plagues many calibration studies. The results were clear: post-calibration probability reliability improved in the temporal cohort, even as raw discrimination did not recover. The models’ rankings of patients may have stayed roughly intact, but the probabilities attached to those rankings became honest again.

The authors are careful, almost unusually so, about what their findings do and do not establish. They state that calibration may be useful for improving probability reliability, but that the results do not establish temporal generalizability, discrimination improvement, or readiness for clinical deployment, and they emphasize that external validation is required before any real-world use. This restraint matters in a field where prediction models are frequently oversold. A model that predicts diagnostic categories before a scan could, in principle, help prioritize urgent cases, reduce unnecessary imaging, or guide the choice of additional tests, but only if its probabilities remain trustworthy across time and across hospitals, questions this single-center retrospective study cannot fully answer.

The study also carries practical lessons for anyone building clinical prediction tools. First, temporal validation should be treated as a standard requirement, not an optional extra, because random splits systematically flatter model performance. Second, calibration deserves the same attention as discrimination; a model with a respectable AUC can still produce probabilities so miscalibrated that they mislead clinical decisions. Third, the equivalence of logistic regression and XGBoost here is a reminder that simpler models remain competitive in many clinical prediction tasks, particularly when the feature set is small and largely categorical. The work, approved by the Institutional Ethics Committee of Dongzhimen Hospital and conducted under the Declaration of Helsinki with waived informed consent for anonymized retrospective data, received no specific funding and reports no competing interests. As hospitals increasingly experiment with pre-scan triage algorithms, this study offers both a template for rigorous evaluation and a warning: the future arrives one day at a time, and models must be recalibrated to meet it.

Subject of Research: Machine learning prediction of pre-scan diagnostic categories in whole-body bone scintigraphy

Article Title: Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost

Article References: Ma, K., Du, Z., & Wu, P. (2026). Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost. BMC Medical Imaging. https://doi.org/10.1186/s12880-026-02848-5

Image Credits: AI Generated

DOI: 10.1186/s12880-026-02848-5

Keywords: bone scintigraphy, machine learning, XGBoost, multinomial logistic regression, prediction model, calibration, temporal validation, bone metastasis, fracture, decision curve analysis, risk stratification, nuclear medicine

Cite Scienmag News

Ophelia Keating. (October 6, 2026). AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy. Scienmag. https://scienmag.com/ai-models-predict-bone-scan-outcomes-before-imaging-but-time-may-erode-their-accuracy/

Ophelia Keating. "AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy." Scienmag, 6 October 2026, https://scienmag.com/ai-models-predict-bone-scan-outcomes-before-imaging-but-time-may-erode-their-accuracy/. Accessed 6 October 2026.

Ophelia Keating. "AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy." Scienmag. October 6, 2026. https://scienmag.com/ai-models-predict-bone-scan-outcomes-before-imaging-but-time-may-erode-their-accuracy/

Tags: AI models for bone disease diagnosisbone metastasisbone scan outcome predictionbone scintigraphybone scintigraphy outcome predictioncalibrationchallenges in maintaining accuracy of predictive healthcare toolsclinical decision support for bone scanscomparison of modeling philosophies in medical AIdecision curve analysisearly prediction of bone abnormalitiesfractureimpact of time on predictive accuracyMachine learningmachine learning for nuclear medicinemultinomial logistic regressionnuclear medicinepre-imaging diagnostic modelsprediction modelpredictive modeling of cancer metastasisretrospective cohort analysis in medical imagingrisk stratificationtemporal validationXGBoost
Share26Tweet16
Previous Post

Inside South Korea’s Plan to Turn a Smart City Into an AI City

Next Post

One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns

Related Posts

One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns
Medicine

One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns

October 6, 2026
AI Reads Ultrasound to Sort Parotid Tumors With Striking Accuracy
Medicine

AI Reads Ultrasound to Sort Parotid Tumors With Striking Accuracy

October 6, 2026
Hidden Platelet Failure May Decide Who Survives Traumatic Brain Injury
Medicine

Hidden Platelet Failure May Decide Who Survives Traumatic Brain Injury

October 6, 2026
What Makes Dementia Family Caregivers Feel Empowered? New Study Maps the Components
Medicine

What Makes Dementia Family Caregivers Feel Empowered? New Study Maps the Components

October 6, 2026
Scientists Uncover a Fat-Metabolism Switch That Fuels Esophageal Cancer and Blunts Chemotherapy
Medicine

Scientists Uncover a Fat-Metabolism Switch That Fuels Esophageal Cancer and Blunts Chemotherapy

October 6, 2026
Aging Eye Study Uncovers Molecular Cascade Behind Blinding Retinal Scarring
Medicine

Aging Eye Study Uncovers Molecular Cascade Behind Blinding Retinal Scarring

October 6, 2026
Next Post
One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns

One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • One Short Night of Sleep Can Skew Concussion Test Scores, Study Warns
  • AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy
  • Inside South Korea’s Plan to Turn a Smart City Into an AI City
  • AI Reads Ultrasound to Sort Parotid Tumors With Striking Accuracy

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading