Deep learning has already transformed river forecasting across the globe, but a lingering question has kept many hydrologists cautious: can an artificial neural network that has never seen a snowpack truly understand floods born from melting snow? A new study from Norway, published in Hydrology and Earth System Sciences, provides the most detailed answer yet, and the verdict is largely encouraging. Researchers at the Norwegian Water Resources and Energy Directorate evaluated a Long Short-Term Memory network, or LSTM, against the country’s operational HBV model across 103 catchments, and for the first time scored the two models separately on floods generated by snowmelt, by rainfall, and by a mixture of the two.
The stakes are high in snow-influenced regions. Floods there arise from two fundamentally different processes. Rainfall floods typically strike in summer or autumn, responding within hours or days to intense precipitation. Snowmelt floods, by contrast, unfold in spring or early summer as weeks of accumulated snowpack release water in a slow, sustained surge. A model that handles one flood type well may fail on the other, and most large-scale evaluations of neural networks have quietly averaged over this distinction, because snow-dominated catchments form a minority in global datasets. By splitting flood events by their generating process, the Norwegian team exposed performance differences that conventional evaluations would have masked.
The LSTM architecture is well suited in principle to the challenge. Unlike traditional recurrent neural networks, which suffer from vanishing gradients and forget information over long sequences, LSTMs use memory cells with gates that selectively retain or discard information across hundreds of time steps. In hydrological terms, this means the network can encode winter precipitation and temperature conditions, hold them in an internal state through months of cold, and then use that stored memory to anticipate the timing and volume of spring runoff. Earlier work has shown that LSTM cell states track snow accumulation and melt with correlations above 0.8 against snow depth reanalysis products, suggesting the network builds an internal analogue of a snowpack without ever being told one exists.
To test this rigorously, the researchers trained a single LSTM model on daily discharge records from 200 gauging stations across mainland Norway, a country spanning 58 to 71 degrees north with annual precipitation ranging from under 300 millimetres in the sheltered east to more than 6000 millimetres on the Atlantic-facing coast. The model received four dynamic forcing series per catchment, derived from the 1-kilometre SeNorge_2018 gridded dataset: catchment-averaged precipitation, catchment-averaged mean temperature, and the catchment-minimum and catchment-maximum daily temperatures. Twenty-one static attributes describing each catchment’s geometry, physiography and climate completed the input. Training covered the fifteen hydrological years from September 2009 to August 2024, with hyperparameters tuned through five-fold cross-validation and the catchment-averaged Nash-Sutcliffe efficiency used as the loss function.
The benchmark was no straw man. The HBV model, a semi-distributed conceptual rainfall-runoff model dividing each catchment into ten elevation zones with explicit snow, soil and runoff storages, has underpinned Norwegian flood forecasting for decades. Its twelve parameters were calibrated catchment by catchment over the same period using the same loss function, ensuring a fair comparison in which both models saw identical data. Evaluation was then performed over a completely unseen fifteen-year window from September 1994 to August 2009, a period deliberately chosen to precede training and to avoid artefacts from changes in the precipitation observing network.
On overall streamflow simulation, the LSTM was emphatically strong. It achieved an average Nash-Sutcliffe efficiency of 0.84 and an average Kling-Gupta efficiency of 0.86 across the 103 evaluated catchments, with NSE exceeding 0.7 in 100 of them. Compared with HBV, the LSTM posted higher NSE in 96 percent of catchments and higher KGE in 77 percent, with the largest gains concentrated where the operational model performed worst. Mean absolute errors were lower for the LSTM in all but three catchments. The only metric where HBV prevailed was the variability ratio component of KGE, a known consequence of training against NSE, which mathematically rewards a slight underestimation of flow variability.
The flood-specific results revealed a genuine asymmetry. The LSTM simulated the correct peak day for 77 percent of rainfall-generated floods but only 50 percent of snowmelt-generated floods, a 27 percentage point gap that mirrored a similar 29 point gap in HBV. The researchers attribute this not to any failure to learn snow processes, but to the physics of the events themselves: snowmelt floods last days to weeks and lack the sharp, unambiguous peak of a rainfall-driven flood, so discharge on adjacent days differs little and hitting the exact day is intrinsically harder. Supporting this interpretation, peak magnitude errors were nearly identical between flood types, with close to half of all events, whether snowmelt or rainfall driven, simulated within 20 percent of the observed maximum. False alarm rates and probabilities of detection were likewise similar for the two pure flood types and slightly worse for mixed events.
Against the operational benchmark, the deep learning model won broadly. LSTM outperformed HBV in 70 to 91 percent of catchments depending on the flood type and metric, with median timing improvements of 14 percentage points for snowmelt floods and 11 for rainfall floods, and improvements exceeding 30 points in several catchments. The largest gains in peak magnitude came for rainfall-generated events, particularly in catchments where HBV’s errors exceeded 40 percent. Notably, the LSTM achieved its improved bias ratios without the precipitation correction factors HBV requires to balance its water budget, suggesting that the data-driven model’s freedom from hard physical constraints can, in this case, yield a more faithful representation of average streamflow.
The study’s authors are careful about what the comparison does and does not prove. By constraining the LSTM’s forcing data and training period to match HBV’s, they guaranteed fairness but also left obvious headroom: longer training records, richer atmospheric inputs, and per-catchment fine-tuning could all push performance further. They also note that the preferred model sometimes depends on the application, since flood warning prizes timing while inundation mapping prizes magnitude, and in roughly a fifth of catchments the two metrics pointed toward different models. Still, the overall message is one of confidence. A neural network trained on data alone, with no snow routine written into its code, can match or beat a physically structured operational model on both of the distinct flood-generating processes that define hydrology in snowy latitudes. For national hydrological services weighing the future of flood forecasting, that is a result worth taking seriously.
Subject of Research: Evaluating LSTM deep learning models for simulating snowmelt- and rainfall-generated floods in Norwegian catchments
Article Title: The ability of LSTM to model snowmelt versus rainfall generated floods
Article References: The ability of LSTM to model snowmelt versus rainfall generated floods. (n.d.). https://doi.org/10.5194/hess-30-6207-2026
Image Credits: AI Generated
DOI: 10.5194/hess-30-6207-2026
Keywords: LSTM, deep learning, flood prediction, snowmelt, rainfall-runoff modelling, hydrology, HBV model, Norway, streamflow, Nash-Sutcliffe efficiency, catchments, flood forecasting
Cite Scienmag News
Violet Maxwell. (October 8, 2026). AI Flood Model Learns to Read Snow and Rain Signals Better Than Operational Tools. Scienmag. https://scienmag.com/ai-flood-model-learns-to-read-snow-and-rain-signals-better-than-operational-tools/
Violet Maxwell. "AI Flood Model Learns to Read Snow and Rain Signals Better Than Operational Tools." Scienmag, 8 October 2026, https://scienmag.com/ai-flood-model-learns-to-read-snow-and-rain-signals-better-than-operational-tools/. Accessed 8 October 2026.
Violet Maxwell. "AI Flood Model Learns to Read Snow and Rain Signals Better Than Operational Tools." Scienmag. October 8, 2026. https://scienmag.com/ai-flood-model-learns-to-read-snow-and-rain-signals-better-than-operational-tools/

