In the heart of central-western Ivory Coast, the towns of Bouaflé and Zuénoula live with an annual threat that arrives with the rains. Situated in the Marahoué River basin, where annual precipitation ranges from 1200 to 1600 millimeters and a dense hydrographic network feeds seasonal flooding, these communities have repeatedly watched torrential downpours overwhelm low-lying areas, damage crops of yam, cassava, rice and maize, and endanger lives. A new study published in Advances in Statistical Climatology, Meteorology and Oceanography offers them something they have never had before: a data-driven early warning capability built entirely from deep learning and freely available satellite observations, capable of predicting daily rainfall up to seven days in advance.
The research, led by Satti J. R. Kamenan of the National Polytechnic Institute Félix Houphouët-Boigny in Yamoussoukro, is the first attempt at near real-time rainfall prediction in any of Ivory Coast’s river basins. The team built a forecasting framework around Long Short-Term Memory neural networks, a deep learning architecture specifically designed to remember patterns across long sequences of data. LSTM networks, first proposed by Sepp Hochreiter and Jürgen Schmidhuber in 1997, use internal gates—forget, input and output gates—that regulate what information is stored, discarded or passed forward through the network’s memory cells. This structure allows them to capture the long-term temporal dependencies and nonlinear behavior that characterize daily rainfall, something conventional statistical models have long struggled with.
Feeding the neural network required an ambitious data pipeline. The researchers combined ground-based rainfall records from SODEXAM, the Ivorian meteorological agency, covering ten stations in the Marahoué watershed from 1961 to 2022, with two NASA satellite products: GPM IMERG, a high-resolution precipitation dataset with a spatial resolution of 0.1 degrees and roughly 30-minute temporal resolution, and MERRA-2, a reanalysis that blends satellite observations, ground measurements and numerical weather models into consistent atmospheric fields including temperature, wind, pressure and humidity. All of this was processed on Google Earth Engine, the cloud-based platform that has become a workhorse for large-scale environmental data analysis. Before the satellite data could be trusted, the team validated it against rain gauge observations from 2000 to 2017, finding correlation coefficients of 0.80 at Bouaflé and 0.83 at Zuénoula—meaning the satellite product captured more than 80 percent of the observed rainfall variability.
Quality control was a critical and methodical step. The historical gauge records contained gaps of 7.29 to 8.04 percent, and raw observations can harbor errors that would corrupt any machine learning model. The team applied a two-stage screening protocol: first, the interquartile range method flagged values that deviated unusually far from each station’s typical distribution; second, a spatial consistency check using the network-wide median and median absolute deviation determined whether a flagged value was physically plausible or simply wrong. Crucially, extreme rainfall events that were spatially coherent across the network were retained even when they exceeded statistical thresholds, preventing the accidental deletion of genuine meteorological extremes. Only 689 of 29,510 observations—2.34 percent—were ultimately identified as outliers, and erroneous values were corrected using the validated GPM IMERG estimates as an auxiliary source.
Selecting the right predictors was equally deliberate. From twenty-one candidate atmospheric variables, the researchers used a Pearson correlation matrix to eliminate multicollinearity, keeping only variables with pairwise correlations below 0.60. They then ran a Random Forest model to rank the remaining candidates by importance, retaining twelve predictors—including seven-day antecedent rainfall, vertical atmospheric motion at 500 hectopascals, specific humidity at multiple pressure levels, and wind components—that each contributed meaningfully to predictive accuracy. The final dataset spanned 15,094 daily time steps from January 1984 to April 2024, with data through January 2020 used for training and the final four years reserved entirely for independent validation, a rigorous test of whether the models had truly learned generalizable patterns rather than memorizing the past.
Three separate LSTM models were trained, one for each forecast horizon: one day ahead, three days ahead and seven days ahead. The architecture settled on three layers of 100 neurons each, a learning rate of 0.001 with automatic reduction, a batch size of 32, and an early stopping mechanism that halted training when validation loss stopped improving. The team then pitted the LSTM against three popular machine learning benchmarks—Random Forest, Extra Trees and XGBoost—using standard metrics including the coefficient of determination, Nash-Sutcliffe Efficiency, normalized root mean square error and mean absolute error. The results were striking. At the one-day horizon, the LSTM achieved R-squared and NSE values exceeding 95 percent at both stations, with a normalized RMSE below 10 percent and a mean absolute error under 1 millimeter, placing it firmly in the ‘excellent’ category by conventional hydrological standards.
The gap between the deep learning model and its competitors widened as the forecast horizon stretched. At three days ahead, the LSTM still maintained R-squared values between 88 and 90 percent with a normalized RMSE below 10 percent, while the tree-based models saw their accuracy slip, with R-squared often dropping below 65 percent. By seven days ahead, the LSTM was the only model still producing statistically reliable forecasts, with R-squared values between 70 and 81 percent and errors remaining controlled, whereas Random Forest and Extra Trees deteriorated dramatically—mean absolute errors exceeding 7 millimeters and NSE values sometimes falling below 40 percent—and XGBoost performed worst of all. The explanation lies in fundamental architecture: tree-based ensembles treat observations largely as static feature sets, while the LSTM’s recurrent memory cells explicitly model the sequential structure and long-range dependencies inherent in daily precipitation series.
Honest scrutiny of the model’s errors revealed both strengths and limits. Residual analysis showed that predictions were largely unbiased for light to moderate rainfall, which constitutes the majority of events in the Marahoué basin, with mean residuals near zero at both stations. But the Kolmogorov-Smirnov test confirmed the residuals were not normally distributed, and a class-by-class examination exposed an intensity-dependent bias: some low-intensity events were slightly underestimated while moderate and heavy rainfall tended to be overestimated, with error dispersion growing sharply for extreme events. This is a well-documented weakness of deep learning applied to precipitation—intense rainfall is rare, irregular and heavily skewed, so models see too few examples to learn its behavior precisely, and recent studies have shown LSTM networks tend to smooth high-intensity peaks rather than reproduce them faithfully.
For the people of Bouaflé and Zuénoula, however, the practical implications are significant. Early warning systems depend less on perfectly reproducing once-a-decade extremes than on reliably tracking the ordinary rainfall regime and flagging major events in time for communities to act—and the LSTM delivers exactly that at lead times of one to three days, with usable guidance even at seven. The authors point to clear paths forward: enriching the training data with additional atmospheric predictors such as convective available potential energy, total column water vapor and sea surface temperatures over the Gulf of Guinea and ENSO regions, applying post-prediction bias correction, and exploring hybrid models that couple neural networks with physical approaches for extreme events. As climate change intensifies rainfall variability across West Africa, this study demonstrates that a developing region can build world-class forecasting capability from open satellite data, cloud computing and a well-designed neural network—no expensive radar infrastructure required.
Subject of Research: Deep learning-based daily rainfall forecasting using satellite and reanalysis data for flood early warning in the Marahoué River basin, Ivory Coast
Article Title: Intelligent daily rainfall prediction for early warning using deep learning and satellite data: application to Bouaflé and Zuénoula stations, Ivory coast
Article References: Kamenan, S. J. R., Youan, T. M., Adja, M. G., Soro, S. I., & Kouassi, A. M. (2026). Intelligent daily rainfall prediction for early warning using deep learning and satellite data: application to Bouaflé and Zuénoula stations, Ivory coast. Advances in Statistical Climatology, Meteorology and Oceanography, 12(1), 173-193. https://doi.org/10.5194/ascmo-12-173-2026
Image Credits: AI Generated
DOI: 10.5194/ascmo-12-173-2026
Keywords: deep learning, LSTM neural networks, rainfall prediction, flood early warning, GPM IMERG, satellite data, Ivory Coast, Marahoué River basin, machine learning, hydrology, MERRA-2, extreme rainfall
Cite Scienmag News
Violet Maxwell. (October 8, 2026). Deep learning turns satellite data into week-long flood warnings for Ivory Coast. Scienmag. https://scienmag.com/deep-learning-turns-satellite-data-into-week-long-flood-warnings-for-ivory-coast/
Violet Maxwell. "Deep learning turns satellite data into week-long flood warnings for Ivory Coast." Scienmag, 8 October 2026, https://scienmag.com/deep-learning-turns-satellite-data-into-week-long-flood-warnings-for-ivory-coast/. Accessed 8 October 2026.
Violet Maxwell. "Deep learning turns satellite data into week-long flood warnings for Ivory Coast." Scienmag. October 8, 2026. https://scienmag.com/deep-learning-turns-satellite-data-into-week-long-flood-warnings-for-ivory-coast/

