Flood engineers and water managers spend their careers planning for rains they have never seen. Because the historical record is short and the future is uncertain, they rely on precipitation generators: statistical models that spin out thousands of years of plausible synthetic rainfall, day by day and gauge by gauge, which can then be fed into hydrological models to stress-test dams, drainage networks and flood defences. A new study published in Advances in Statistical Climatology, Meteorology and Oceanography by Jakob Benjamin Wessel of the University of Exeter and Richard E. Chandler of University College London now overhauls one of the most widely used families of these generators, extending it in ways that make its simulated storms more faithful to reality, particularly at the extremes where the stakes are highest.
The family in question is built on generalised linear models, or GLMs, a workhorse of modern statistics. In a GLM-based precipitation generator, each day at each site is treated in two stages. First, a logistic regression determines the probability that rain falls at all, using predictors such as the season, the location’s geography, the previous day’s weather and indices of large-scale atmospheric circulation. Second, on days that turn wet, the rainfall amount is drawn from a gamma distribution whose mean is linked to another set of covariates. The approach has earned its popularity honestly: it can represent systematic variation in space and time, generate rainfall at ungauged locations, handle missing data and unequal record lengths gracefully, and compete favourably with more elaborate alternatives such as hidden Markov models. Crucially, because the covariates can include large-scale atmospheric variables, the same machinery doubles as a statistical downscaling tool, allowing climate model projections to be translated into local rainfall scenarios.
Yet the standard GLM framework carries two structural handicaps. The first concerns the gamma distribution used for wet-day intensities. For mathematical convenience, the shape parameter of that distribution is held constant across all conditions, which forces the standard deviation of daily rainfall to be strictly proportional to its mean. There is no physical justification for this constraint; it is simply a legacy of computational simplicity. The second handicap is linearity: covariates are combined only as weighted sums, so any curved relationship, say between temperature and rainfall intensity, must be approximated by clunky polynomial terms chosen by hand. Wessel and Chandler suspected that these rigidities, rather than any fundamental flaw in the gamma family itself, were behind a well-documented shortcoming of GLM generators: their tendency to misrepresent how extreme precipitation varies with the seasons, overestimating winter tails and underestimating summer ones.
To test that hypothesis, the authors turned to generalised additive models for location, scale and shape, known as GAMLSS. This framework, developed by Rigby and Stasinopoulos in 2005, allows every parameter of the rainfall distribution, not just the mean, to respond to covariates, and it permits those responses to be smooth, flexible curves rather than straight lines. The curves are built from spline basis functions, with cyclic splines handling the repeating annual cycle and thin-plate bivariate splines capturing spatial variation across latitude and longitude. Because such flexible models can, in principle, contort themselves to fit almost any dataset perfectly, GAMLSS are fitted by maximising a penalised log-likelihood that rewards smoothness, with the degree of penalisation chosen automatically by the software. The result is a generator in which atmospheric conditions can influence both the average rainfall on a wet day and its variability, letting the data decide how those influences are shaped.
One immediate technical obstacle was model selection. Standard tools such as the Akaike Information Criterion assume that observations are independent, which is plainly false for a network of rain gauges swept by the same frontal systems. For parametric models the authors employed an adjusted version of the criterion that accounts for inter-site dependence, sometimes called the network information criterion. For semiparametric models, where the sheer number of spline basis functions makes that adjustment unreliable, they adopted a resampling-based alternative, the WIC, which compares penalised log-likelihoods across carefully constructed bootstrap samples. Because ordinary bootstrap resampling would destroy the spatial and temporal structure of daily rainfall, the authors resampled entire days, stratified so that each resampled dataset preserved the original numbers of observations, wet days and missing values.
The second major contribution addresses a subtler weakness of GLM generators: the way they handle spatial dependence. Traditionally, occurrence and intensity are treated separately, so when synthetic rainfall is generated across a network, the dry sites are simply ignored when intensities are drawn at wet neighbours. A site adjacent to a parched location is thus just as likely to record a downpour as one surrounded by rain. Wessel and Chandler adapted a transformed Gaussian fields approach, effectively a Gaussian copula, in which a vector of correlated standard normal variables underlies the whole network. Each latent value is mapped to a rainfall amount: if it falls below a site-specific threshold the site stays dry, and otherwise it is converted through the inverse gamma distribution into a wet-day intensity. Because one latent field drives both occurrence and intensity simultaneously, the model can naturally produce the phenomenon of spatial intermittence, where sites near the edge of a rain band record only light showers.
Estimating the correlation structure of that latent field posed its own challenge, since records rarely overlap perfectly across a century of observations. The authors estimated each pairwise correlation by maximum likelihood, treating dry-day values as censored, and then fitted a Matérn correlation function to the resulting estimates, weighted by the amount of data behind each pair. This guarantees a valid correlation matrix, smooths out estimation noise and even supplies correlations for pairs of stations with no overlapping record at all. In the study region, a modest 50 by 40 kilometre catchment of the river Blackwater in southern England, where the terrain rises from the Thames floodplain to escarpments near 300 metres and where around 80 percent of days see either nearly all or nearly none of the gauges wet, this unified treatment proved essential.
The case study itself drew on daily records from 51 quality-checked stations in the Met Office MIDAS archive spanning 1959 to 2022, conditioned on atmospheric covariates derived from the ERA5 reanalysis: two-metre temperature and dewpoint temperature, mean sea level pressure and ten-metre wind speed. Comparing four generators, a parametric GLM, a semiparametric GAM and their GAMLSS extensions, the authors found that most key monthly statistics, including means, wet-day proportions, standard deviations and autocorrelations, fell comfortably within the range spanned by repeated simulations. The semiparametric GAMLSS captured wet-day proportions and variability best, with the improvements traceable respectively to flexible covariate curves and to the liberated shape parameter. Quantile-quantile comparisons showed that relaxing the constant-shape assumption modestly reduced the seasonal overestimation of winter extremes, though it did not eliminate it, suggesting that fully taming seasonal tails may eventually require stepping outside the gamma family altogether, perhaps toward three- or four-parameter distributions or extreme value approaches.
Two further findings give the study practical bite well beyond model elegance. First, the authors discovered that a seemingly innocuous preprocessing step, rounding rainfall records to a common resolution such as 0.5 millimetres to reconcile changes in measurement practice, can substantially distort the fitted distribution, including its upper tail, because the shape of a gamma distribution is tightly coupled to its behaviour at small values where rounding bites hardest. Rather than rounding, they recommend handling resolution changes inside the model itself, via a binary covariate flagging the era of coarser measurement. Second, and reassuringly for practitioners, while seasonal tail biases appeared in the pooled single-site distributions, none appeared in the catchment-averaged rainfall, meaning that for most hydrological applications the generators already perform satisfactorily at the scales that matter. The authors caution that semiparametric models extrapolate poorly to atmospheric conditions far outside their training range, a caveat relevant to climate change downscaling, and that the computational cost of the most flexible models remains nontrivial. Still, with code released openly and the framework demonstrated on one of the most intensively studied catchments in the literature, the work offers hydrologists a sharper, more honest instrument for imagining the rains of the future.
Subject of Research: Statistical modelling of multisite daily precipitation using GAMLSS and transformed Gaussian fields for improved stochastic weather generation
Article Title: Improving multisite precipitation generators based on generalised linear models
Article References: Wessel, J. B., & Chandler, R. E. (2026). Improving multisite precipitation generators based on generalised linear models. Advances in Statistical Climatology, Meteorology and Oceanography, 12(1), 149-172. https://doi.org/10.5194/ascmo-12-149-2026
Image Credits: AI Generated
DOI: 10.5194/ascmo-12-149-2026
Keywords: precipitation generators, generalised linear models, GAMLSS, stochastic weather generation, spatial dependence, extreme precipitation, flood risk, statistical downscaling, Gaussian copulas, hydrology, ERA5 reanalysis, gamma distribution
Cite Scienmag News
Violet Maxwell. (October 8, 2026). Smarter Rain Simulators: New Statistical Upgrade Sharpens Extreme Precipitation Forecasts. Scienmag. https://scienmag.com/smarter-rain-simulators-new-statistical-upgrade-sharpens-extreme-precipitation-forecasts/
Violet Maxwell. "Smarter Rain Simulators: New Statistical Upgrade Sharpens Extreme Precipitation Forecasts." Scienmag, 8 October 2026, https://scienmag.com/smarter-rain-simulators-new-statistical-upgrade-sharpens-extreme-precipitation-forecasts/. Accessed 8 October 2026.
Violet Maxwell. "Smarter Rain Simulators: New Statistical Upgrade Sharpens Extreme Precipitation Forecasts." Scienmag. October 8, 2026. https://scienmag.com/smarter-rain-simulators-new-statistical-upgrade-sharpens-extreme-precipitation-forecasts/

