Some of the most consequential questions in climate science are also the hardest to answer with computers: how likely is a storm so severe it might strike only once a century, or a flood that defies every event in the historical record? Estimating the probabilities of such rare extremes by brute force is brutally expensive. To estimate the probability of a once-per-century event to within ten percent relative error, a climate model would need to run for roughly ten thousand simulated years, with the vast majority of that computing time spent simply waiting for the next disaster. A new study by Justin Finkel and Paul A. O’Gorman, published in Nonlinear Processes in Geophysics, tackles a deceptively simple question that sits at the heart of a cleverer approach: exactly when should you nudge a simulation to turn a moderate event into an extreme one?
The technique in question is known as ensemble boosting, a member of the rare event sampling family. Instead of waiting for extremes to appear, researchers identify moderately severe events in a relatively short simulation and re-run them with small perturbations to their initial conditions, generating ‘descendant’ simulations that are physically plausible but potentially far more severe. The method, which grew out of rare event sampling first developed for nuclear safety assessment in the 1950s, has been applied to heatwaves, tropical cyclones, and precipitation. But its success hinges on a parameter the authors call the advance split time, or AST: how far before the event’s peak one applies the perturbation. Split too late, and the ensemble members have no room to diversify before the event passes. Split too early, and they forget the special conditions that made their ancestor extreme in the first place, drifting back toward ordinary climatology.
Previous work by the same authors had found an empirical rule of thumb in a simple toy model: perturb at the moment when ensemble members have dispersed to about three-eighths of their eventual error saturation. But that rule was tuned to the Lorenz-96 system, a one-dimensional chain of variables with little geographic structure, and it underestimated the optimal timing when applied to temperature and precipitation extremes in a full general circulation model. The new study asks whether a more principled, generalizable criterion exists, one that behaves less like an arbitrary algorithmic dial and more like an intrinsic property of the physical system, analogous to how Lyapunov exponents encode the timescale on which small errors double.
To explore the question, the researchers chose a test system of intermediate complexity: a two-layer quasigeostrophic flow, a classic idealized model of baroclinic instability that captures jets, waves, and vortices reminiscent of midlatitude storm tracks, augmented with a passive tracer whose local concentration serves as the target variable. The tracer plays the role of a pollutant plume or localized heavy rainfall: it is advected by traveling waves, so its extremes are sudden, transient spikes rather than slowly building anomalies. This intermittency is precisely what makes the advance split time matter. The model, with roughly eleven thousand degrees of freedom and two eastward jets whose positions meander on hundred-day timescales, offers enough spatial structure to test how the optimal timing varies with location, something the one-dimensional toy model could never reveal.
The experimental design was exhaustive by necessity. The authors ran a short direct simulation of about eleven simulated years to harvest a pool of ‘ancestor’ extreme events, and a much longer control simulation of about forty-four years to serve as ground truth. For each ancestor, they launched ensembles of twenty-one descendants at twenty different advance split times ranging from two to forty days before the ancestral peak, perturbing the flow with a single impulsive kick drawn from a carefully designed two-dimensional space of amplitudes and phases. A quadratic regression model then mapped each perturbation to the resulting severity of the descendant’s tracer spike, allowing the researchers to estimate full conditional probability distributions rather than relying on scattered samples alone.
From these conditional distributions, the authors built estimates of the climatological tail, the probability distribution of the most severe events, using two related estimators. The first, which they introduce and call MoCTail, averages the conditional tail distributions across ancestors. The second, dubbed PoPTail and adapted from recent work by other researchers, pools the numerators and denominators of the probability ratios before dividing. Both performed similarly in practice, and both demonstrated that boosted ensembles, when launched from a well-chosen split time, can reconstruct the tail of the distribution far more accurately than counting events from the unboosted simulation alone.
The central contribution is a proposed criterion for choosing the split time without any knowledge of the ground truth. The authors evaluated several candidates, including thresholded entropy, a measure borrowed conceptually from reinforcement learning that rewards ensembles whose severity distributions are both high on average and widely spread. Short split times produce narrow, tightly clustered severities with low information content; long split times produce wide but unremarkable distributions that regress toward climatology. Thresholded entropy peaks precisely in the golden window between these failure modes, where descendants are both diverse and severe. The authors also tested expected improvement, another acquisition-function-style criterion, though it proved less reliable across the full range of target locations.
The results were striking. At a representative latitude, the optimal uniform split time was fourteen days, or roughly one to two eddy turnover timescales, and the entropy-based criterion selected split times for individual events that landed in nearly the same region of the optimization landscape. Allowing each event to choose its own timing, with no access to the ground truth, produced tail estimates nearly as accurate as exhaustively searching over all possible uniform timings. Across latitudes, the optimal timing varied systematically: shorter splits were favored near the northern edges of the westerly jets, where meridional wind shear is negative, and longer splits near the southern edges. The old three-eighths rule, while not optimal everywhere, ran through the middle of the meandering valley of good performance, retaining value as a starting guess.
Compared at equal cost with a plain simulation, the boosted approach delivered speedups of roughly one-and-a-half to three times, with the advantage growing as the target probabilities became smaller. That is modest compared to the orders-of-magnitude gains reported for full rare event algorithms, but the authors emphasize that speed was never the goal. The point was to establish, for the first time in a physically meaningful turbulent system, what actually needs to be optimized before efficiency can be pursued. The findings suggest that the optimal split time is a property of the dynamics and the target observable, not of the algorithm, which means it could in principle be estimated once and reused.
The road to applying these ideas in operational climate models remains long. The authors note that deploying the method at scale will require adaptive optimization strategies that home in on good split times during a run, rather than the exhaustive grid searches used here, and that the shape of the perturbations themselves is a largely unexplored lever. Still, the study offers something rare in this field: a concrete, entropy-based answer to the question of when to intervene in a chaotic system, and evidence that the answer generalizes across locations within the flow. For a science that must quantify the risk of disasters too rare to observe, knowing exactly when to plant the butterfly flap may prove to be an essential first step.
Subject of Research: Optimizing the advance split time in ensemble boosting for rare event sampling of extreme tracer fluctuations in a quasigeostrophic turbulent flow
Article Title: Boosting ensembles for statistics of tails at conditionally optimal advance split times
Article References: Finkel, J., & O'Gorman, P. A. (2026). Boosting ensembles for statistics of tails at conditionally optimal advance split times. Nonlinear Processes in Geophysics, 33(2), 233-265. https://doi.org/10.5194/npg-33-233-2026
Image Credits: AI Generated
Keywords: rare event sampling, ensemble boosting, extreme weather, advance split time, quasigeostrophic model, climate statistics, thresholded entropy, passive tracer, turbulence, probability tails, Monte Carlo methods, baroclinic instability
Cite Scienmag News
Violet Maxwell. (October 9, 2026). Scientists Find the Sweet Spot for Simulating Rare Weather Extremes. Scienmag. https://scienmag.com/scientists-find-the-sweet-spot-for-simulating-rare-weather-extremes/
Violet Maxwell. "Scientists Find the Sweet Spot for Simulating Rare Weather Extremes." Scienmag, 9 October 2026, https://scienmag.com/scientists-find-the-sweet-spot-for-simulating-rare-weather-extremes/. Accessed 9 October 2026.
Violet Maxwell. "Scientists Find the Sweet Spot for Simulating Rare Weather Extremes." Scienmag. October 9, 2026. https://scienmag.com/scientists-find-the-sweet-spot-for-simulating-rare-weather-extremes/

