Ocean waves could become a major clean-energy resource, but turning that promise into reliable power is still difficult. Wave energy converters (WECs) must convert rapidly changing ocean motion into electricity while staying within strict mechanical and electrical safety limits. Even advanced control strategies struggle when the device dynamics are hard to model or inherently nonlinear.
A central challenge is that the best-performing control is noncausal: to maximize energy capture, the controller needs information about the waves not just now, but a few seconds ahead. In conventional approaches such as model predictive control, this requires an accurate mathematical description of the converter’s hydrodynamics. That requirement becomes especially problematic for emerging “soft” and flexible WEC concepts, where deformation and fluid–structure interactions are complex.
To bypass the need for an accurate device model, researchers combined reinforcement learning (RL) with short-term wave forecasting. The study uses proximal policy optimization (PPO), a popular RL algorithm for continuous control, to learn control actions directly from interaction with the environment. Rather than relying on a detailed hydrodynamic simulator, the agent adjusts control gains in real time based on the converter’s measured state and predicted wave characteristics.
In the experiments, the PPO agent controlled a point-absorber WEC benchmark by tuning the parameters of a noncausal scheme. The controller’s inputs included wave forecasts one, three, or five prediction steps into the future, allowing it to anticipate how upcoming wave conditions would affect buoy motion and actuator forces. This anticipatory capability is what distinguishes noncausal wave control from standard damping-based, purely causal strategies.
When tested using 200 seconds of real irregular wave-height data collected off the coast of Cornwall, UK—data not seen during training—the approach showed a measurable advantage. With a five-step look-ahead, the PPO controller generated up to 13.9% more energy than a conventional baseline, while keeping device motion and actuation within safety constraints throughout the trials. The longer forecast horizon also improved training efficiency, reducing the number of episodes needed to reach effective behavior.
The researchers further stressed the system by introducing imperfect information: they added noise to the wave forecasts and created mismatches between training and testing conditions. Despite these disturbances, the controller continued to operate stably and produce energy, indicating robustness to prediction errors and modeling uncertainty—factors that typically limit real ocean deployments.
According to the team, this work is among the first to combine policy-gradient RL with explicit wave prediction for wave-energy control. The authors describe it as a promising, model-free route toward controlling next-generation wave devices that are difficult to model with classical methods.
Next steps include evaluating the controller on a hardware-in-the-loop dSPACE real-time platform to test computational speed and timing under realistic hardware constraints. The framework will also be extended beyond mechanical power to include the electrical power conversion stage of a complete WEC system.
Subject of Research: Noncausal wave energy control using PPO reinforcement learning with wave prediction
Article Title: Proximal policy optimization–based noncausal control for wave energy conversion systems
News Publication Date: 18-Jun-2026
Web References: http://dx.doi.org/10.26599/OCEAN.2026.9470015
References: Ocean (DOI: 10.26599/OCEAN.2026.9470015)
Image Credits: Ocean, Tsinghua University Press
Keywords: wave energy, noncausal control, reinforcement learning, proximal policy optimization (PPO), wave forecasting, point-absorber WEC, model-free control, robustness, hardware-in-the-loop

