Road traffic crashes kill more than 1.3 million people every year, and the burden falls overwhelmingly on low- and middle-income countries, which account for over 90 percent of road traffic deaths according to the World Health Organization. Beyond the human toll, the societal costs of traffic accidents in developed economies are estimated to range between 0.5 and 6.0 percent of GDP, with human losses making up 50 to 75 percent of those costs in various countries. For decades, researchers have tried to predict how severe a crash will be, using everything from ordered logit and probit regressions to, more recently, sophisticated machine learning and deep learning models. Now, a sweeping systematic review published in Heliyon has mapped an entire decade of that research, examining 237 primary studies published between 2014 and 2024 to answer a deceptively simple question: when algorithms decide which crashes matter most, can we actually understand why they decided?
The review, led by Karim Asif Sattar and colleagues at Universiti Putra Malaysia, is the first comprehensive systematic literature review to focus specifically on explainable artificial intelligence, or XAI, in road traffic crash severity prediction. Previous surveys had examined machine learning models or neural network applications in isolation, but none had systematically synthesized the explainability and interpretability techniques layered on top of them. The researchers searched Web of Science and Scopus, retrieving 3,616 records, of which 2,768 remained after duplicate removal. Through a PRISMA-guided screening process involving two independent reviewers, they whittled the pool down to 237 peer-reviewed journal articles, each of which applied at least one machine learning or deep learning technique to crash severity prediction. The timing of the review window was deliberate: before 2014, the field was dominated by traditional statistical approaches, whereas the past decade witnessed the rapid adoption of ensemble learning, deep neural networks, and post hoc explanation methods such as SHAP and LIME.
The central tension the review exposes is the classic trade-off between accuracy and interpretability. Logistic regression and decision trees are transparent, but they struggle with high-dimensional data and nonlinear relationships. Artificial neural networks and ensemble models capture complex patterns with impressive accuracy, yet their inner workings are notoriously opaque, the so-called black box problem. In a domain where decisions affect human lives, from emergency dispatch to infrastructure investment, that opacity is not merely an academic inconvenience. The authors argue that both strong predictive performance and model interpretability are essential for deploying decision-support systems in emergency response applications, where the swift dispatch of medical personnel to accident locations can mean the difference between recovery and fatality.
On the modeling side, the review found that machine learning techniques were used more frequently than deep learning approaches, with roughly a third of studies employing both. Random forest was the single most popular algorithm, appearing in 109 applications, followed by decision tree variants and support vector machines. Boosting methods such as XGBoost, gradient boosting decision trees, and AdaBoost featured prominently, and newer frameworks like LightGBM and CatBoost showed steady growth in adoption over the review period. Among deep learning architectures, artificial neural networks led the pack, with convolutional neural networks and long short-term memory models gaining traction, the latter particularly for sequential and time-series crash data. Graph neural networks, which represent crash records as nodes in a network, showed limited but increasing adoption toward the end of the period.
Perhaps the most striking finding concerns feature selection, which the authors categorize as ablation methods: techniques used to identify, rank, or eliminate input variables during model development. Approximately 64 percent of the reviewed studies employed such methods. The toolbox is remarkably diverse, spanning wrapper-based approaches like Boruta and recursive feature elimination, regularization techniques such as LASSO and elastic net, information-theoretic measures including information gain and gain ratio, statistical tests like chi-square and Pearson correlation, and even nature-inspired optimization algorithms including genetic algorithms, coyote optimization, and multi-objective evolutionary frameworks like NSGA-II and PESA-II. In one illustrative case, the PESA-II algorithm reduced a feature set from 31 variables to 13 while achieving a prediction accuracy of 94.7 percent, demonstrating how aggressive dimensionality reduction can coexist with strong performance.
On the explainability side, SHAP, which stands for SHapley Additive exPlanations, emerged as the undisputed champion. Grounded in cooperative game theory, SHAP assigns each feature a contribution value for individual predictions, guaranteeing a unique attribution solution with properties including local accuracy, missingness, and consistency. The review documented a marked increase in SHAP usage from 2019 to 2024, applied alongside XGBoost, LightGBM, random forest, CatBoost, and decision trees, among others. Its appeal lies in providing both global and local explanations with intuitive visualizations that help communicate findings to policymakers. Yet the technique is not without drawbacks: computational costs can balloon on large datasets, and the additive decomposition assumption may not fully capture complex feature interactions, while strongly correlated features can produce unstable attribution values.
Other explainability techniques occupy important niches. LIME, or Local Interpretable Model-Agnostic Explanations, approximates a complex model with a simple surrogate around individual observations, making it valuable for case-level investigation. In one study, LIME revealed that clear weather conditions were positively associated with fatal accidents in three of four examined cases, suggesting drivers become careless in good conditions, a factor that global importance rankings might have obscured. Partial dependence plots and individual conditional expectation curves visualize how predictions shift as features change, with one analysis finding that driving speeds exceeding 60 kilometers per hour in residential zones increased injury accident likelihood by around 10 percent. Accumulated local effects plots offer an advantage over partial dependence by remaining robust to correlated features, while permutation feature importance and leave-one-covariate-out methods estimate importance by measuring performance degradation when features are shuffled or removed.
The practical payoff of this transparency is already visible in the literature. Studies on work zone crashes identified lane closures, work zone length, and annual average daily truck traffic as major severity contributors, informing safer work zone design and dynamic warning systems. Research on rural mountainous freeways linked steep grades, sharp curves, and driver unfamiliarity with severe outcomes, supporting interventions such as adaptive speed limits, geofencing-based driver warnings, and connected-vehicle technologies. Truck crash studies using SHAP with XGBoost revealed that crash configuration matters most for both passenger car and truck driver injuries, while freight-focused analyses tied population density and freeway lane mileage to injury severity, prompting calls for dedicated freight corridors. The review also synthesized policy recommendations spanning speed enforcement, mandatory restraint use, pedestrian fencing, rumble strips, and stricter driving-hour regulations for truck operators.
The authors are careful to note the review’s limitations: the search was restricted to two databases and English-language journal articles, conference papers were excluded, and no formal risk-of-bias assessment was conducted. Because the included studies differed substantially in datasets, severity definitions, and validation procedures, the reported frequencies of technique use should not be read as evidence of predictive superiority. Still, the trajectory is clear. The review points toward future directions including counterfactual explanations, integrated gradients, causal XAI approaches, and domain-adapted large language models, alongside persistent challenges in establishing causal rather than merely associative relationships between crash factors and injury severity. As transportation agencies increasingly consider algorithmic decision support for emergency response and road safety planning, this decade-long map suggests that the era of unexplainable black boxes in crash prediction is steadily, and necessarily, coming to a close.
Subject of Research: Explainable machine learning and deep learning approaches for predicting road traffic crash severity
Article Title: A systematic mapping review of explainable machine and deep learning approaches for road traffic crash severity prediction: A decade-long review (2014-2024)
Article References: Sattar, K. A., Ishak, I., Suriani, L., Rum, S. N. M., & Masiur Rahman, S. (2026). A systematic mapping review of explainable machine and deep learning approaches for road traffic crash severity prediction: A decade-long review (2014-2024). Heliyon, 12(15), Article e45513. https://doi.org/10.1016/j.heliyon.2026.e45513
Image Credits: AI Generated
DOI: 10.1016/j.heliyon.2026.e45513
Keywords: explainable AI, road traffic crashes, crash severity prediction, machine learning, deep learning, SHAP, LIME, feature selection, random forest, XGBoost, systematic review, road safety
Cite Scienmag News
Blake Davidson. (October 4, 2026). Decade of Data: How Explainable AI Is Cracking Open the Black Box of Crash Severity Prediction. Scienmag. https://scienmag.com/decade-of-data-how-explainable-ai-is-cracking-open-the-black-box-of-crash-severity-prediction/
Blake Davidson. "Decade of Data: How Explainable AI Is Cracking Open the Black Box of Crash Severity Prediction." Scienmag, 4 October 2026, https://scienmag.com/decade-of-data-how-explainable-ai-is-cracking-open-the-black-box-of-crash-severity-prediction/. Accessed 4 October 2026.
Blake Davidson. "Decade of Data: How Explainable AI Is Cracking Open the Black Box of Crash Severity Prediction." Scienmag. October 4, 2026. https://scienmag.com/decade-of-data-how-explainable-ai-is-cracking-open-the-black-box-of-crash-severity-prediction/

