Thursday, September 3, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Medicine

Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis

September 3, 2026
in Medicine
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 7 mins read
0
Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis

Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

A large-scale study drawing on three seasons of data from all four Irish professional rugby union teams has found that machine learning models cannot reliably predict which elite male players are about to sustain a non-contact lower limb injury. The research, led by Chris Leckey of University College Dublin and the Irish Rugby Football Union, together with colleagues including Eamonn Delahunt and Nicol van Dyk, was published in Sports Medicine – Open and represents one of the most comprehensive attempts yet to apply artificial intelligence to injury forecasting in rugby union. Despite access to granular satellite tracking, wellness and musculoskeletal screening data from 129 professional players, the best-performing model achieved only modest discrimination and produced one correct injury prediction for every 77 false alarms.

Rugby union presents a formidable injury problem. The professional game carries an injury incidence of roughly 91 injuries per 1000 player hours, far exceeding elite soccer at 36 per 1000 hours and Gaelic football at 56 per 1000 hours, reflecting the sport's unique blend of collisions and intermittent high-speed running. Lower limb injuries are the most common, and previous research has shown that teams with higher injury burden perform worse in league competition. Both the International Olympic Committee and World Rugby have called for better injury surveillance and data analytics to protect athletes, fuelling interest in machine learning approaches that can detect subtle patterns in high-dimensional athlete data.

The appeal of machine learning in this context is easy to understand. Modern professional teams generate enormous volumes of information every day, from satellite-derived running loads to subjective wellness scores and clinical screening measurements. Traditional statistical approaches, which typically examine one or a handful of variables at a time, struggle to interrogate datasets of this size and complexity. Machine learning algorithms, by contrast, are designed to identify nonlinear interactions among hundreds of variables simultaneously, and they have shown promise in other domains of sports science, including performance forecasting and load management. It was against this backdrop of optimism that the Irish research team designed their study.

However, most machine learning studies in rugby have been confined to a single team or a single season, limiting generalisability and often relying on inconsistent injury definitions. Findings from one club's squad, trained on one season's worth of data, may simply reflect the idiosyncrasies of that playing style, that medical team's record-keeping, or that particular year's fixture congestion. The Irish research team set out to address both shortcomings by pooling longitudinal data from Connacht, Leinster, Munster and Ulster Rugby, along with the Ireland men's national team, across the 2021/22, 2022/23 and 2023/24 seasons. The study adopted a strict, consensus-based definition of injury: any non-contact, soft-tissue injury to the muscles, ligaments or tendons of the lower limb that caused the player to miss a subsequent training session or match.

This definitional rigour matters more than it might first appear. Much of the existing literature on injury prediction has used looser criteria, counting any complaint of pain or any medical attention event regardless of time loss. By restricting the analysis to injuries with genuine functional consequences, the researchers reduced the noise in their outcome variable, but they also drastically reduced the number of positive cases available for the algorithms to learn from, a trade-off that would prove consequential.

The datasets were substantial. Wearable Global Navigation Satellite System units manufactured by STATSports recorded external workload for 47,547 training sessions and 5,773 match exposures, capturing metrics such as total distance, maximal speed, metabolic load and acceleration-deceleration demands. Players also logged self-perceived wellness scores on a mobile application, rating sleep quality, appetite, fatigue, stress, muscle tightness and soreness on Likert scales from 0 to 10. Team medical staff administered musculoskeletal screening tests, including knee-to-wall ankle dorsiflexion, sit-and-reach flexibility and adductor squeeze strength measured with a sphygmomanometer. All injury records were centrally collated in the IRFU's athlete management system, and where the nature of an injury was unclear from clinical notes, the researchers reviewed match footage to determine whether the inciting event involved contact.

From an initial pool of 281 registered players, inclusion was restricted to the 129 athletes with complete GNSS, wellness and musculoskeletal data across all three seasons, a decision the authors acknowledge may have introduced selection bias toward players with greater availability. Applying the strict injury definition reduced 1,047 recorded time-loss injuries to 261 in-scope events. The median player sustained two such injuries over the study period, while 26 players escaped entirely. Back three players had the highest injury rate per player at 2.64, followed by half backs at 2.45. The median time loss per injury was 13 days, with the most severe case sidelining a player for 160 days. Injury incidence rates were 12.6 per 1000 hours in matches and 2.4 per 1000 hours in training.

Data preprocessing was extensive. GNSS features with more than 20% missing data were dropped, reducing 234 variables to 89, and collinearity checks using Pearson correlation and variance inflation factors removed redundant metrics. Outliers beyond three standard deviations of positional match demands were excluded as likely sensor malfunctions, accounting for 2% of records. Missing GNSS data were imputed using multiple imputation by chained equations, wellness data with k-nearest neighbours, and musculoskeletal values with last observation carried forward. The team then engineered temporal features, calculating exponentially weighted moving averages over 3-, 7-, 14-, 21- and 28-day windows, cumulative seven-day loads, and week-over-week percentage changes. Crucially, the models were never shown data from the day of injury, and predictions were only generated for days with training or match exposure, preventing the model from inflating its accuracy by correctly predicting "no injury" on rest days.

These design choices reflect hard-won lessons from earlier work in the field. Several high-profile injury prediction studies in other sports have been criticised for data leakage, in which information from the injury day itself, such as a truncated training session, inadvertently signals that something has gone wrong. Others have reported deceptively high accuracy simply because the models learned to predict the overwhelmingly common outcome of "no injury" on any given day. By withholding injury-day data and restricting predictions to exposure days, the Irish team built safeguards against both pitfalls, ensuring that whatever performance the models achieved would be an honest reflection of their predictive ability.

Five algorithms were compared: logistic regression and support vector machines as linear and kernel-based benchmarks, alongside Random Forest, XGBoost and CatBoost ensemble methods. Class imbalance was addressed through cost-sensitive learning with weighted loss functions. To avoid information leakage, the data were split temporally, with everything before 1 October 2023 used for training and the remainder reserved for testing, yielding a 75/25 split. Hyperparameter optimisation was performed with nested stratified three-fold cross-validation using the Optuna framework, and model explainability was assessed post hoc with SHapley Additive exPlanations (SHAP) analysis.

The results were sobering. All five models fell into the "poor" predictive category, defined as a receiver operating characteristic area under the curve between 0.5 and 0.7. CatBoost performed best, achieving a test ROC AUC of 0.66, an area under the precision-recall curve of 0.008 against an injury prevalence of just 0.39%, and a precision of 0.013. XGBoost was the only other model to exceed a test ROC AUC of 0.6, at 0.61. Logistic regression matched Random Forest on the test set but failed to beat chance during cross-validation, a fluctuation the authors attribute to high variance in performance estimation under severe class imbalance.

The precision-recall figure deserves particular attention, because it speaks directly to clinical usefulness. With injuries occurring on fewer than one in 250 exposure days, even a model with apparently respectable ROC AUC can be practically useless, flagging far too many false positives to be actionable. This is precisely what the Irish data showed.

The practical implications of these numbers are stark. Of the 54 injuries sustained during the nine-month testing period, the model correctly flagged only eight, involving six players. Across 13,833 observations, the model issued 626 high-risk alerts, roughly four per provincial team each week, but only one in 78 alerts was correct. The authors warn that such a high false-positive rate would likely induce "alarm fatigue" among medical practitioners, desensitising them to warnings and undermining any real-world deployment. Injury severity also showed no clear relationship with predicted risk, and while 52% of injuries occurred at risk scores of 0.5 or above, that region comprised 27% of all predictions, with injuries scattered across the entire risk profile.

Imagine, for a moment, the operational reality behind those statistics. A team doctor receiving four alerts per week, nearly all of them false, would face an impossible choice: either restrict or modify the training of a healthy player on the basis of a warning that is almost certainly wrong, or ignore the alerts altogether and lose whatever marginal benefit the system might have offered. Neither option is defensible in a professional environment where training decisions carry competitive and financial consequences.

The SHAP analysis offered a further cautionary insight: four of the five most influential features were non-modifiable risk factors. The strongest predictor was simply whether the next day's session was a match or a training session, reflecting the fivefold higher injury incidence in matches. Injury history to date and time since last injury ranked second and third, with age also prominent. The only workload metric in the top five was the 28-day exponentially weighted moving average of maximal speed, which appeared protective at higher values, consistent with prior evidence that chronic sprint exposure mitigates lower limb injury risk. The authors stress that SHAP values reflect contribution to predictions, not causal relationships, and should be interpreted cautiously given the model's weak performance.

That the most informative signals were things teams already know, that matches are riskier than training, that previously injured and older players are more vulnerable, is itself instructive. The models, given hundreds of sophisticated workload and screening variables, essentially rediscovered basic epidemiology without adding new actionable insight on top of it.

The study's strengths lie in its unprecedented scale and standardisation: identical GNSS hardware, wellness applications and screening protocols across all teams, with injuries logged in a single centralised athlete management system. Yet the authors acknowledge limitations, including potential differences in playing styles and surfaces between clubs, selection bias toward durable players, within-player correlation from including recurrent injuries, unthresholded last-observation-carried-forward imputation of musculoskeletal data, non-uniform screening frequency for players under clinical care, and the absence of calibration or decision-analytic evaluation. No sample size calculation was performed, and no observational machine learning finding should be read causally.

Subject of Research: Medicine

Subject of Research: Medicine

Article Title: Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis

Article References: Leckey, C., van Dyk, N., Berndsen, J., Doherty, C., Lawlor, A., & Delahunt, E. (2026). Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis. Sports Medicine - Open, 12(1), Article 123. https://doi.org/10.1186/s40798-026-01060-7

Image Credits: AI Generated

DOI: 10.1186/s40798-026-01060-7

Keywords: data-driven sports injury management, elite male rugby union injury prevention, injury prevention strategies in rugby, injury risk prediction algorithms, machine learning in sports medicine, machine learning injury prediction in rugby, machine learning model validation in sports, multi-season sports injury analysis, multi-team sports injury studies, non-contact lower limb injury risk assessment, rugby injury epidemiology, sports injury risk modeling

Cite Scienmag News

Blake Davidson. (August 31, 2026). Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis. Scienmag. https://scienmag.com/machine-learning-model-development-and-evaluation-for-non-contact-lower-limb-injury-risk-prediction-in-elite-male-rugby-union-a-multi-team-multi-season-analysis/

Blake Davidson. "Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis." Scienmag, 31 August 2026, https://scienmag.com/machine-learning-model-development-and-evaluation-for-non-contact-lower-limb-injury-risk-prediction-in-elite-male-rugby-union-a-multi-team-multi-season-analysis/. Accessed 3 September 2026.

Blake Davidson. "Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male Rugby Union: A Multi-team, Multi-season Analysis." Scienmag. August 31, 2026. https://scienmag.com/machine-learning-model-development-and-evaluation-for-non-contact-lower-limb-injury-risk-prediction-in-elite-male-rugby-union-a-multi-team-multi-season-analysis/

Tags: challenges in AI-based sports injury predictioncomparison of injury risks across contact sportscomprehensive multi-season rugby injury data analysisdata-driven sports injury managementelite male rugby union injury preventionimpact of injuries on team performance in rugbyinjury incidence rates in professional rugbyinjury prevention strategies in elite rugbyinjury prevention strategies in rugbyinjury risk prediction algorithmslimitations of machine learning in sports injury predictionmachine learning in sports medicinemachine learning injury prediction in rugbymachine learning injury prediction in rugby unionmachine learning model validation in sportsmulti-season sports injury analysismulti-team sports injury studiesnon-contact lower limb injury risk assessmentrugby injury epidemiologyrugby player wellness and musculoskeletal screeningsatellite tracking and injury risk modeling in rugbysports injury forecasting using artificial intelligencesports injury risk modelingsports medicine injury prediction
Share26Tweet16
Previous Post

Method for quantification of microplastic release from plastic-based materials during weathering

Next Post

Groundwater and hydrogeochemical patterns of Campi Flegrei active caldera (southern Italy) for volcanic hazard assessment

Related Posts

Living Near Green Spaces May Lower ALS Risk, Italian Study Finds
Medicine

Living Near Green Spaces May Lower ALS Risk, Italian Study Finds

September 3, 2026
Ketogenic Diet Shows Promise as Epigenetic Therapy for SETD1B Epilepsy
Medicine

Ketogenic Diet Shows Promise as Epigenetic Therapy for SETD1B Epilepsy

September 3, 2026
Adverse Drug Reactions Common and Often Preventable in Ethiopian Obstetric Patients
Medicine

Adverse Drug Reactions Common and Often Preventable in Ethiopian Obstetric Patients

September 3, 2026
Expert Insights on Cancer Care Leadership With Professor Timothy Eberlein
Medicine

Expert Insights on Cancer Care Leadership With Professor Timothy Eberlein

September 3, 2026
Recurrent Pulmonary Alveolar Proteinosis Challenges Care After Double Lung Transplant
Medicine

Recurrent Pulmonary Alveolar Proteinosis Challenges Care After Double Lung Transplant

September 3, 2026
Blue-Light Blocking Glasses and Morning Light Shift Young Adults’ Body Clocks Earlier
Medicine

Blue-Light Blocking Glasses and Morning Light Shift Young Adults’ Body Clocks Earlier

September 3, 2026
Next Post
Groundwater and hydrogeochemical patterns of Campi Flegrei active caldera (southern Italy) for volcanic hazard assessment

Groundwater and hydrogeochemical patterns of Campi Flegrei active caldera (southern Italy) for volcanic hazard assessment

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Amodal completion enables 3D wheat reconstruction from a single image
  • Comprehension is in the Eye of the Reader: An Eye Tracking Study of Children and Adults
  • Scientists reveal sugar transporter gene regulation in peony by ABA signals
  • New Mixed-Ligand Metal Complexes Show Promise as Antibiotics, Antioxidants and Corrosion Shields

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading