In the narrow, tide-swept waters between the Indonesian islands of Java and Sumatra, surface currents can rip through the Sunda Strait at speeds exceeding 150 centimetres per second, posing a constant challenge to the ships, ferries and port operators that depend on this vital shipping corridor. A new study published in Advances in Statistical Climatology, Meteorology and Oceanography has now put a dozen forecasting methods head to head to determine which can best predict these dangerous flows, and the verdict is clear: deep learning models deliver the most accurate forecasts, particularly when looking several hours ahead. The work, led by Dhava Gautama and Alifficionaldo Agpri Putra of Indonesia’s Meteorology, Climatology, and Geophysical Agency (BMKG), draws on more than three years of hourly radar observations of the strait, one of the most energetically tidal environments ever subjected to such a systematic comparison.
The data underpinning the study come from a high-frequency (HF) radar system operated by BMKG, which measures the eastward and northward components of surface flow across 291 active grid points spaced roughly one kilometre apart on a 21-by-21 grid. After rigorous quality control, including speed thresholds, temporal consistency checks and spatial outlier filtering, the record spans December 2022 to February 2026, yielding 28,321 hourly timesteps with about 85 percent temporal coverage. Missing values were filled using tidal harmonic predictions from the UTide software, ensuring that every forecasting model received gap-free input sequences. The data were then split chronologically into training, validation and test periods, with the test window covering July 2025 through February 2026, so that all methods were judged on conditions they had never seen.
The first phase of the evaluation compared twelve methods for one-step-ahead prediction, spanning a deliberately wide ladder of sophistication: naive persistence, tidal harmonic analysis, classical time-series models such as ARIMA and exponential smoothing, a reduced-rank spatio-temporal statistical model known as EOF-VAR, shallow machine learning, and several deep learning architectures. Evaluated consistently across all 291 grid cells and the full test period, the deep learning models achieved the lowest root-mean-square errors for both current components. A convolutional neural network (CNN) reached 11.31 centimetres per second for the zonal component, a skill score of 0.50 relative to persistence, while a hybrid CNN-gated recurrent unit (CNN-GRU) achieved 15.44 centimetres per second for the meridional component, a skill score of 0.42. All three neural networks posted correlation coefficients above 0.97 for the zonal component.
Perhaps the most scientifically revealing result of the first phase was not that deep learning won, but why the classical baselines lost. The pointwise statistical models, which treat each grid cell in isolation, fell well behind, with ARIMA managing skill scores of only around 0.12 to 0.13. Yet when the researchers allowed a purely linear statistical model to share information across space, through the EOF-VAR approach that decomposes the current field into empirical orthogonal functions and models the leading modes with a vector autoregression, it recovered most of the gap to the neural networks, reaching skill scores of 0.39 and 0.35. The authors conclude that the principal limitation of the classical baselines is their neglect of spatial structure, not their lack of nonlinearity. The neural networks nonetheless retained a consistent edge, which the researchers attribute to their nonlinear representation of the coupled spatio-temporal field.
The second phase probed how much past information the models need, retraining the three best deep learning architectures with lookback windows of three, six and twelve hours. Because the dominant semidiurnal M2 tide in the Sunda Strait has a period of roughly 12.42 hours, a twelve-hour window provides approximately one full tidal cycle of phase information. Only the hybrid CNN-GRU improved monotonically with longer lookbacks, edging its errors down to 11.21 and 15.41 centimetres per second, while the standalone CNN and GRU showed no consistent benefit. The interpretation is architectural: the GRU component can integrate tidal phase information from ordered sequences, while the CNN supplies spatial context at each timestep, whereas the standalone CNN stacks frames along the channel dimension and loses temporal ordering, and the standalone GRU flattens the spatial fields and loses spatial structure.
The third phase extended the task to six-hour nowcasting, comparing direct multi-step prediction against autoregressive decoding. Three architectures were tested: a ConvLSTM encoder-decoder that generates each future frame conditioned on its previous prediction, a bidirectional encoder-forecaster (BiEF) that processes the input sequence in both directions, and a direct multi-step CNN-GRU that emits all six future frames in a single forward pass. Under a carefully controlled comparison, with all five trainable models retrained across five random seeds on identical hardware and training budgets, accuracy turned out to be governed primarily by model capacity, plateauing near one million parameters. The one-million-parameter CNN-GRU-MS-Small achieved the best results, 18.39 and 22.07 centimetres per second, with skill scores of 0.77 and 0.72, and at matched capacity the direct and autoregressive strategies were statistically indistinguishable.
Where the architectures did differ sharply was computational cost, and this proved decisive for operational deployment. The direct multi-step model trains about 2.5 times faster and runs roughly 4.5 times faster at inference, because it requires a single forward pass rather than sequential decoding of the six-hour horizon. A single direct model also covers all 291 grid cells, whereas the strongest classical baseline would require fitting and maintaining 291 separate per-cell ARIMA models. On this basis the researchers recommend the compact CNN-GRU-MS-Small for operational nowcasting, noting that while all models are fast enough in absolute terms for hourly updates, the direct approach scales better to larger grids, higher update frequencies and constrained hardware.
The study also uncovered a striking diurnal rhythm in forecast errors. All five neural network models showed elevated errors during afternoon and evening local hours, and concurrent wind measurements from three BMKG weather stations at Merak, Ciwandan and Bakauheni revealed a clear sea-breeze signal peaking near 15:00 local time. The zonal prediction errors correlated positively with this afternoon wind enhancement, with a Spearman rank correlation of 0.48 across the 24 diurnal hours, consistent with sea-breeze-driven currents injecting variability that current-only models cannot anticipate. The meridional errors, by contrast, followed a distinct diurnal pattern not explained by wind speed, and the evidence points instead to tidal-phase effects. Seasonal stratification added a further layer: errors rose by 20 to 26 percent during the northwest monsoon, when stronger and more variable winds introduce non-tidal dynamics that no model in the comparison fully captured.
Two further analyses strengthened confidence in the findings. A tidal-residual decomposition showed that the deep learning models reduce non-tidal, likely wind-driven variability by 16 to 20 percent relative to persistence, while confirming that ARIMA’s modest advantage over persistence derives entirely from tracking the smooth tidal oscillation rather than from any skill at predicting the sub-tidal residual that dominates total error. And when the researchers mapped errors across the strait, they found the same geography for every method, classical and neural alike: errors lowest in the central domain and highest near the northwest and southeast boundaries where tidal amplification and spatial gradients are steepest, with per-cell error maps correlating at 0.96 to 0.98 between ARIMA and the deep networks. This indicates that the error distribution is set by the underlying tidal physics rather than by model class.
Relative errors of 3 to 7 percent of the observed speed range are of the same order as those reported in far calmer open-sea environments such as the Gulf of Thailand and Monterey Bay, though the authors caution that this normalisation partly reflects the strait’s exceptionally large speed range and that cross-site comparisons are limited by differences in resolution and horizon. The message, they stress, is not that deep learning is superior everywhere, but that it does not break down in energetic strait settings. For the thousands of vessels transiting the Sunda Strait each year, and for the growing number of coastal regions worldwide where HF radar networks monitor fast-changing waters, the study offers a practical blueprint: hybrid architectures that exploit both space and time, lookback windows long enough to capture a full tidal cycle, and direct multi-step designs that trade nothing in accuracy while cutting computational cost dramatically. As machine learning continues its advance through the geosciences, this benchmark demonstrates that the winning formula lies as much in honest, capacity-matched comparison against strong baselines as in the sophistication of the models themselves.
Subject of Research: Comparative evaluation of statistical and deep learning methods for forecasting high-frequency radar surface currents in the tidally energetic Sunda Strait
Article Title: Comparative evaluation of statistical and deep learning methods for high-frequency radar surface current forecasting in a narrow tropical strait
Article References: Gautama, D., & Putra, A. A. (2026). Comparative evaluation of statistical and deep learning methods for high-frequency radar surface current forecasting in a narrow tropical strait. Advances in Statistical Climatology, Meteorology and Oceanography, 12(2), 221-242. https://doi.org/10.5194/ascmo-12-221-2026
Image Credits: AI Generated
DOI: 10.5194/ascmo-12-221-2026
Keywords: ocean surface currents, high-frequency radar, deep learning, Sunda Strait, tidal forecasting, CNN-GRU, ConvLSTM, nowcasting, maritime safety, operational oceanography, machine learning, Indonesia
Cite Scienmag News
Blake Davidson. (October 8, 2026). Deep learning outpaces classical statistics in forecasting fierce currents of a tropical strait. Scienmag. https://scienmag.com/deep-learning-outpaces-classical-statistics-in-forecasting-fierce-currents-of-a-tropical-strait/
Blake Davidson. "Deep learning outpaces classical statistics in forecasting fierce currents of a tropical strait." Scienmag, 8 October 2026, https://scienmag.com/deep-learning-outpaces-classical-statistics-in-forecasting-fierce-currents-of-a-tropical-strait/. Accessed 8 October 2026.
Blake Davidson. "Deep learning outpaces classical statistics in forecasting fierce currents of a tropical strait." Scienmag. October 8, 2026. https://scienmag.com/deep-learning-outpaces-classical-statistics-in-forecasting-fierce-currents-of-a-tropical-strait/

