Low Earth orbit is becoming a crowded place. Tens of thousands of tracked objects, from functioning satellites to fragments of decades-old rocket bodies, sweep around the planet at speeds approaching eight kilometers per second, and every close pass between two of them is a potential catastrophe. Each day, the United States Space Force’s 18th Space Control Squadron issues Conjunction Data Messages, or CDMs, warning satellite operators that two objects are predicted to come uncomfortably close. The overwhelming majority of these warnings lead to nothing: the objects miss by kilometers. But the handful that matter can destroy a spacecraft, trigger a cloud of debris, and set off a cascade that endangers everything in the same orbital shell. Deciding which alerts deserve a costly avoidance maneuver has long been one of the most consequential judgment calls in space operations, and a new study published in Applied Intelligence argues that a carefully engineered hybrid of Monte Carlo simulation and ensemble machine learning can make that call with unprecedented reliability.
The research, led by Meriem Ouari of the National School of Artificial Intelligence in Algiers together with colleagues Fella Manel Zerrouki, Seif Eddine Bouziane, and Amel Gacem of the Algerian Space Agency, tackles a problem that has quietly frustrated both statisticians and operators for years. Conjunction data is brutally imbalanced. Out of hundreds of thousands of recorded close approaches, only a tiny fraction ever evolve into genuinely dangerous encounters, and the probability values reported in CDMs cluster heavily at the numerical detection floor, a floor so low that it effectively means the screening system could not distinguish the risk from zero. Machine learning models trained naively on such data learn to predict the boring answer, that nothing will happen, and they achieve impressive-looking accuracy while missing precisely the events that justify firing thrusters. The Algerian team’s framework is designed from the ground up to defeat that failure mode.
At the heart of the approach lies a probabilistic modeling layer built on Monte Carlo methods, a family of techniques that resolve uncertainty by simulating thousands or millions of randomized scenarios and examining the distribution of outcomes. In the collision risk context, the uncertainties in the two objects’ positions and velocities, which propagate from imperfect tracking measurements and atmospheric drag variations, are sampled repeatedly to generate an ensemble of possible encounter geometries. Rather than collapsing this rich uncertainty into a single point estimate, the framework preserves the full spread of simulated outcomes and feeds distributional features derived from it into the learning pipeline. This is a deliberate marriage of classical astrodynamics and modern data science: the physics of orbital uncertainty is expressed in the language of probability, and the machine learning layer is then asked to learn how those probabilistic signatures map onto actual risk.
On top of the Monte Carlo foundation, the researchers constructed a stacked ensemble, an architecture in which multiple diverse models are trained on the same problem and a higher-level learner combines their predictions. The base layer integrates gradient boosting implementations, including the family of algorithms exemplified by XGBoost, LightGBM, and CatBoost, alongside deep neural networks designed for tabular data. The choice is well grounded in the recent machine learning literature, which has repeatedly shown that tree-based models still outperform deep networks on structured tabular datasets of the kind CDMs represent, while neural networks contribute complementary representations of the same underlying features. Stacking, a technique whose lineage traces back to David Wolpert’s work on stacked generalization in the early 1990s, allows a meta-model to learn when to trust each base model, extracting signal from their disagreements rather than discarding it.
The team went further, evaluating event-level sequence models that treat the successive CDMs issued for a single conjunction event as a temporal trajectory rather than as isolated snapshots. This matters because a close approach is not judged once; it is monitored over days as tracking data improves and the predicted miss distance is refined again and again. The pattern of how a risk estimate evolves across those updates carries information that no single message contains. A conjunction whose predicted probability creeps steadily upward across successive CDMs tells a different story from one whose estimate flickers around the detection floor, and models that can read those temporal patterns gain access to a richer evidentiary basis for classification.
Perhaps the most operationally significant contribution is the calibration strategy the authors introduce. In forecasting, a model is said to be calibrated when its stated probabilities mean what they say: if it assigns a ten percent risk to a hundred different events, roughly ten of those events should actually materialize. Uncalibrated models may rank risks correctly while systematically overstating or understating them, which is disastrous when a stated probability is the trigger for spending propellant and interrupting a mission. The study draws on established calibration theory, including the work of Kuleshov, Fenner, and Ermon on calibrated regression and the classical DeGroot and Fienberg framework for comparing forecasters, to ensure that the final probabilistic outputs are trustworthy, not merely discriminative. The distinction between ranking events well and assigning them honest probabilities is central to the paper’s philosophy.
The experimental stage of the work rests on a substantial and publicly available benchmark: the European Space Agency’s Spacecraft Collision Avoidance Challenge dataset, hosted on the ESA Kelvins platform and originally released to spur machine learning competition in this domain. The dataset comprises 199,082 Conjunction Data Messages spanning 15,321 distinct conjunction events, a scale that makes it one of the most demanding resources available for this task. It was also the proving ground for a 2021 machine learning competition whose design and results were later documented by Uriot, Izzo, and colleagues, giving the community a shared reference point. Against this benchmark, the hybrid framework delivered an F2 score of 0.808, a metric that deliberately weights recall more heavily than precision, reflecting the operational reality that missing a genuine collision threat is far worse than investigating a false alarm. The model also achieved an area under the receiver operating characteristic curve of 0.981, indicating near-perfect separation between dangerous and benign encounters across all decision thresholds.
Those numbers, strong as they are, only tell half the story. The authors emphasize that the framework maintains well-calibrated probabilistic forecasts alongside its classification performance, meaning the model does not have to sacrifice honest uncertainty for discriminative power. This dual achievement is rare. Many high-performing classifiers produce probability outputs that are badly distorted, particularly under severe class imbalance, and post-hoc fixes often trade one quality for the other. Demonstrating both simultaneously on a dataset of this size suggests that the Monte Carlo ensemble architecture is capturing genuine structure in the conjunction data rather than exploiting statistical shortcuts. For satellite operators, that combination translates directly into better decisions: fewer wasted maneuvers on phantom threats, fewer sleepless nights over ambiguous alerts, and a defensible quantitative basis for the threshold at which action is taken.
The broader context makes the timing of this work significant. The population of resident space objects in low Earth orbit has grown dramatically, driven by large commercial constellations, an expanding launch cadence, and the legacy debris documented in foundational studies such as Liou and Johnson’s 2006 assessment of orbital debris risks in Science. Every new satellite multiplies the number of potential conjunctions that screening systems must evaluate, and manual analysis cannot scale to that volume. Automated, AI-driven risk assessment is therefore not a luxury but a necessity, and the Algerian team’s results, building on prior benchmarking efforts at European space debris conferences and on their own earlier work with physics-informed generative adversarial networks, chart a credible path toward systems that screen conjunctions autonomously while reserving human judgment for the genuinely marginal cases.
There are, of course, limits to what any data-driven system can promise. Conjunction events that end in collision are so rare that even a dataset of nearly two hundred thousand messages contains few true positives, and no benchmark can fully reproduce the stakes of a real operational decision. The authors are candid that their framework is a step toward reliable AI-driven collision risk assessment, not a replacement for the full conjunction assessment pipelines maintained by major space agencies. Yet the study’s central lesson is likely to resonate well beyond orbital mechanics: when decisions are made under extreme imbalance and uncertainty, the winning recipe combines physically grounded probabilistic modeling, ensembles of complementary learners, temporal context, and rigorous calibration. As the sky grows busier, the algorithms that keep it safe will be the ones that not only know which encounters are dangerous, but can say exactly how confident they are.
Subject of Research: Hybrid Monte Carlo and ensemble machine learning for probabilistic satellite collision risk prediction from Conjunction Data Messages
Article Title: Probabilistic satellite collision risk prediction via Monte Carlo ensembles
Article References: Ouari, M., Zerrouki, F. M., Bouziane, S. E., & Gacem, A. (2026). Probabilistic satellite collision risk prediction via Monte Carlo ensembles. Applied Intelligence, 56(15), Article 443. https://doi.org/10.1007/s10489-026-07494-6
Image Credits: AI Generated
DOI: 10.1007/s10489-026-07494-6
Keywords: satellite collision risk, conjunction data messages, Monte Carlo simulation, ensemble learning, gradient boosting, deep neural networks, probabilistic forecasting, calibration, space debris, low Earth orbit, ESA collision avoidance challenge, stacked generalization
Cite Scienmag News
Blake Davidson. (October 4, 2026). AI Ensemble Learns to Predict Which Satellites Are Truly About to Collide. Scienmag. https://scienmag.com/ai-ensemble-learns-to-predict-which-satellites-are-truly-about-to-collide/
Blake Davidson. "AI Ensemble Learns to Predict Which Satellites Are Truly About to Collide." Scienmag, 4 October 2026, https://scienmag.com/ai-ensemble-learns-to-predict-which-satellites-are-truly-about-to-collide/. Accessed 4 October 2026.
Blake Davidson. "AI Ensemble Learns to Predict Which Satellites Are Truly About to Collide." Scienmag. October 4, 2026. https://scienmag.com/ai-ensemble-learns-to-predict-which-satellites-are-truly-about-to-collide/

