Artificial intelligence has stormed into weather forecasting, and few areas have moved faster than nowcasting, the short-range prediction of precipitation over the next couple of hours. Generative models trained on radar archives can paint strikingly realistic rainfields, complete with the crisp textures that classical statistical methods tend to blur away. Yet a new study from the Royal Meteorological Institute of Belgium suggests that when it comes to the one thing forecasters need most from an ensemble, knowing where the forecast is likely to be wrong, these dazzling models still have a fundamental blind spot. The research, published in Weather and Climate Dynamics, puts a generative diffusion model called LDCast head-to-head with the operational stalwart STEPS, and the verdict is nuanced: both systems size up their uncertainty well, but neither can tell you where that uncertainty lives.
The team, led by Martin Bonte with Lesley De Cruz, Fabian Debal and Stéphane Vannitsem, evaluated the two models over Belgium using the RADCLIM radar product, a quantitative precipitation estimate that merges radar measurements with rain gauge observations through Kriging with External Drift. Crucially, neither model was fine-tuned for the Belgian domain. LDCast was used with its original pre-trained weights, and STEPS was run through the open-source pysteps library exactly as a national weather service would deploy it. This choice matters, because in practice most forecast offices will not have the time, data or expertise to retrain a state-of-the-art AI model before using it. The evaluation therefore reflects a realistic operational scenario rather than a best-case laboratory demonstration.
The two models represent fundamentally different philosophies of nowcasting. STEPS, short for Short-Term Ensemble Prediction System, is built on Lagrangian persistence: it advects the observed rainfall field along a motion field estimated by optical flow, decomposes the field into a cascade of spatial scales, and lets each scale evolve through an auto-regressive process. Randomness enters through stochastic perturbations of both the motion field and the intensities. LDCast, by contrast, is a latent diffusion model. A variational autoencoder compresses sequences of rainfall fields into a compact latent space, a forecaster network predicts the latent representation of future fields, and a denoiser stack then generates different ensemble members conditioned on that output. The generative approach naturally produces diverse, nondeterministic forecasts, which makes ensemble construction almost effortless.
To compare the two fairly, the researchers selected ten convective and ten stratiform rainfall events, classified using ERA5 reanalysis data on convective and large-scale precipitation rates. For each event they generated 50-member ensembles and examined the spread-error relationship scale by scale, using Fourier power spectra. In a perfectly calibrated ensemble, the spread among members should match the error of the ensemble mean against observations. The spectral analysis showed that both models come remarkably close to this ideal for scales larger than about five kilometres. LDCast was slightly underdispersive overall, though somewhat overdispersive at large scales during convective events, while STEPS was underdispersive at short lead times for stratiform rain and became well calibrated as lead time grew.
One subtlety the authors had to untangle involved the smallest scales, below roughly five kilometres, where the power spectrum of radar observations flattens out. This flattening is not a physical feature of rainfall but an artefact of white noise contamination in the radar data and of the smoothing introduced by merging with rain gauges. The apparent underdispersion of both models at these scales therefore cannot be blamed on the ensembles themselves. Interestingly, the error at these smallest scales saturates almost immediately, within the first five minutes of the forecast, at scales between five and ten kilometres depending on the model and event type, meaning the issue is largely irrelevant for practical forecasting.
Beyond the raw size of uncertainty, the team probed the internal geometry of the ensembles using the covariance matrix of the members and its eigenvectors, essentially a principal component analysis of the forecast uncertainty. The results revealed a striking architectural difference. STEPS ensembles, especially in convective cases, are dominated by a handful of directions in phase space: convective ensembles concentrate their variability in roughly ten dominant modes, and stratiform ensembles in about five. The residual vectors of individual members align along these few leading eigenvectors, meaning the members are quite similar to one another. LDCast ensembles, in contrast, display much more homogeneous eigenvalue spectra, spreading their variability more evenly across many modes. Both models, however, adapt their ensembles to the meteorological situation, developing larger-scale perturbations for stratiform events and smaller-scale ones for convection, a sign that the models are responding to genuine dynamical differences between the two regimes.
The eigenvalue growth itself followed power laws in lead time, a behaviour consistent with the theoretical predictability limits of atmospheric flows whose energy spectra follow the classical minus five-thirds power law at nowcasting scales. Errors seeded at small scales cascade toward larger ones, and the perturbation modes gradually develop power at progressively larger scales as the forecast advances. This hierarchy, in which the largest-amplitude perturbations affect the largest scales, mirrors the way predictability is lost in the real atmosphere, and it is encouraging that both a physics-inspired statistical scheme and a purely data-driven generative model reproduce it.
The critical test came when the researchers asked whether the ensembles could localise the error in space, not just quantify it. They used two complementary metrics: the cosine of the angle between the actual forecast error and its projection onto the ensemble perturbations, a geometric measure related to the PECA score, and the probabilistic Fraction Skill Score, which assesses spatial accuracy while tolerating displacement errors. The clever twist was the construction of surrogate ensembles. MAAFT surrogates, inspired by the Iterated Amplitude-Adjusted Fourier Transform, preserve the statistical fingerprints of the original ensembles, namely the distribution of pixel values and the power spectra of the perturbations, but scramble the complex Fourier phases that encode spatial structure. If the real ensembles contain genuine spatial information, they should outscore these statistical ghosts.
They did not. The scores of the MAAFT surrogates were essentially indistinguishable from those of the actual STEPS and LDCast ensembles for both metrics, at all lead times and for both event types. In Fourier language, the complex phases of the ensemble perturbations, which carry all the spatial localisation information, are effectively random. The ensembles know how big the error should be at each scale, and they even reproduce the right distribution of rainfall intensities, but they have no dynamical insight into where the error will actually materialise. The skill they display is statistical, not dynamical. A similar picture emerged from the Fraction Skill Score analysis, where both models’ scores saturated with lead time toward values achievable by purely random predictions constrained to the same value distributions.
The authors are careful about what this means for operations. The finding does not invalidate generative nowcasting; LDCast’s spread still provides a useful, nearly unbiased estimate of error magnitude across most scales, even without regional retraining, and both models demonstrably adapt their uncertainty to the weather regime. But forecasters hoping that an ensemble of AI rainfields will point toward the specific locations where the forecast is going wrong will be disappointed. The study also notes a practical asymmetry: when STEPS runs out of radar data at the edge of its domain, it produces explicit missing values that honestly flag where information is lacking, whereas a generative model like LDCast simply invents plausible-looking values with no such transparency. Future work, the authors suggest, should test a version of LDCast fine-tuned for Belgium, evaluate other generative architectures such as DGMR, which produces less diverse members, and explore blending nowcasting models with numerical weather prediction, an approach already available in pysteps. For now, the message is clear: AI ensembles have learned the statistics of rainfall uncertainty, but the dynamics of where that uncertainty strikes remain, for the moment, beyond their grasp.
Subject of Research: Evaluation of ensemble spread-error relationships and spatial error representation in precipitation nowcasting using STEPS and the generative AI model LDCast over Belgium
Article Title: Spread/error relationship and spatial error representation in precipitation nowcasting: comparison of STEPS and generative AI
Article References: Bonte, M., De Cruz, L., Debal, F., & Vannitsem, S. (2026). Spread/error relationship and spatial error representation in precipitation nowcasting: comparison of STEPS and generative AI. Weather and Climate Dynamics, 7(3), 1837-1852. https://doi.org/10.5194/wcd-7-1837-2026
Image Credits: AI Generated
Keywords: precipitation nowcasting, LDCast, STEPS, generative AI, ensemble forecasting, spread-error relationship, latent diffusion model, radar, uncertainty quantification, Fraction Skill Score, convective precipitation, Weather and Climate Dynamics
Cite Scienmag News
Sloane Callahan. (October 8, 2026). AI Rain Forecasts Get the Size of Uncertainty Right but Not the Place. Scienmag. https://scienmag.com/ai-rain-forecasts-get-the-size-of-uncertainty-right-but-not-the-place/
Sloane Callahan. "AI Rain Forecasts Get the Size of Uncertainty Right but Not the Place." Scienmag, 8 October 2026, https://scienmag.com/ai-rain-forecasts-get-the-size-of-uncertainty-right-but-not-the-place/. Accessed 8 October 2026.
Sloane Callahan. "AI Rain Forecasts Get the Size of Uncertainty Right but Not the Place." Scienmag. October 8, 2026. https://scienmag.com/ai-rain-forecasts-get-the-size-of-uncertainty-right-but-not-the-place/

