Every day, thousands of vessels broadcast their position, speed, and heading through the Automatic Identification System, a radio-based tracking network that has become the backbone of maritime traffic monitoring. For port authorities, coast guards, and shipping companies, the ability to predict where a ship will be minutes or hours from now is enormously valuable, supporting collision avoidance, route optimization, anomaly detection, and search-and-rescue planning. Yet the raw AIS stream is notoriously messy. Messages arrive at irregular intervals, transponders drop out, coordinates jump, and human input errors corrupt the very kinematic information that prediction models depend on. A new study published in the Journal of Big Data by Marilena Sinni and Dimitris M. Kyriazanos of the Institute of Informatics and Telecommunications at the National Centre for Scientific Research Demokritos in Athens argues that the path to accurate vessel trajectory prediction begins not with a bigger neural network, but with far more disciplined treatment of the data feeding it.
The researchers propose a two-stage framework that separates the problem into rigorous preprocessing and a carefully chosen recurrent architecture. In the first stage, a comprehensive pipeline takes raw AIS messages through data cleaning, trajectory extraction, kinematic-aware anomaly correction, and temporal resampling. The phrase kinematic-aware is the crucial detail. Rather than filtering points using generic statistical rules, the pipeline exploits the physical relationships between a vessel’s reported position, Speed Over Ground, and Course Over Ground to decide whether a given message is plausible. A ship cannot teleport across a harbor or reverse course between two consecutive pings while maintaining a constant heading, and the preprocessing stage is designed to catch exactly these kinds of inconsistencies before they poison the training data.
Temporal resampling addresses another stubborn problem in AIS analytics. Because vessels transmit at uneven and often unpredictable intervals, the time gaps between consecutive position reports can range from seconds to many minutes. Most sequence models, including the recurrent networks used in this study, implicitly assume a regular cadence. By resampling trajectories onto a consistent temporal grid, the pipeline produces analysis-ready sequences in which each step represents the same elapsed time, allowing the network to learn motion patterns without having to simultaneously compensate for irregular spacing. The authors emphasize that this seemingly mundane step, together with cleaning and anomaly correction, contributes measurably to prediction accuracy, a claim they back up with a systematic ablation study rather than mere assertion.
The second stage of the framework is a Bidirectional Gated Recurrent Unit network, or BiGRU, that jointly predicts three interrelated outputs: the vessel’s future position, its Speed Over Ground, and its Course Over Ground. A standard recurrent network reads a sequence in one direction, from past to present, and compresses everything it has seen into a hidden state. A bidirectional variant processes the sequence in both directions, allowing the representation of each time step to incorporate context from what comes before and after it. For trajectory data, this means the model can use the shape of the entire observed motion history to characterize the vessel’s current maneuvering state, capturing patterns such as turns, accelerations, and course changes that a purely forward-looking encoder may blur.
One of the more elegant technical choices in the study concerns the treatment of Course Over Ground. Heading is a circular variable: a course of 359 degrees and a course of 1 degree are separated by only 2 degrees, yet a naive model computing ordinary numerical differences would judge them to be 358 degrees apart. Treating angles as ordinary regression targets can therefore produce systematically misleading errors, particularly for vessels sailing near the 0-degree meridian of the compass. The researchers solve this by employing a specialized cosine-based distance for the COG output, which respects the circular geometry of the variable and ensures that predictions are penalized in proportion to the true angular separation. This kind of domain-informed loss design reflects a broader lesson in applied machine learning: matching the mathematical structure of the loss function to the physics of the problem can matter as much as architectural novelty.
To demonstrate that the gains come from the method rather than from favorable tuning, the authors benchmarked the BiGRU against a demanding lineup of alternatives under identical preprocessing and training protocols. The comparison set included unidirectional Long Short-Term Memory and GRU networks, a bidirectional LSTM, and attention-augmented variants of the GRU, BiGRU, and BiLSTM encoders. Attention mechanisms, which allow models to weight the most informative parts of the input sequence, have become nearly ubiquitous in sequence modeling, so their inclusion makes the evaluation considerably more convincing. The BiGRU with the cosine-aware course loss emerged as the strongest configuration, suggesting that for this task a well-matched recurrent architecture can hold its own against more elaborate designs.
Equally important is the ablation study, in which the authors systematically removed or altered individual preprocessing stages to measure each one’s contribution. The results confirmed that data cleaning, kinematic-aware anomaly correction, and temporal resampling each add measurable accuracy, validating the central thesis that rigorous denoising prior to model training is not optional overhead but a genuine performance driver. In a field where headlines tend to celebrate novel architectures, this finding is a useful corrective. A model trained on noisy, irregularly sampled, physically implausible trajectories will learn noise, and no amount of downstream sophistication can fully recover information that was destroyed upstream.
Generalization was tested on two real-world AIS datasets drawn from distinctly different maritime environments: traffic around the port of Brest in France and data from the Danish Maritime Authority. These regions differ in traffic density, vessel populations, maneuvering patterns, and data characteristics, so consistent performance across both is meaningful evidence that the framework is not overfitted to one geographic setting. The authors report that the approach delivered accurate multi-output trajectory prediction across these diverse conditions, supporting its potential as a general-purpose tool for maritime surveillance rather than a bespoke solution for a single waterway.
The work was carried out within the FaRADAI project, funded under grant agreement number 101103386, and reflects a growing European research effort to bring artificial intelligence to maritime domain awareness. The practical implications extend well beyond academic benchmarks. Reliable short-term trajectory prediction can feed collision-avoidance alerts, help vessel traffic services anticipate congestion in crowded approaches, flag vessels whose predicted paths deviate from expected behavior, and support environmental monitoring by anticipating where ships will travel under given conditions. Because the framework operates on standard AIS inputs, it could in principle be integrated into existing monitoring infrastructure without requiring new sensor deployments.
For the maritime analytics community, the study offers a clear recipe: respect the physics of the data, normalize its timing, and choose an architecture whose inductive biases match the structure of vessel motion. The combination of kinematic-aware preprocessing, bidirectional recurrence, and circular-aware loss design proved sufficient to outperform stronger and more fashionable alternatives on two independent datasets. As global shipping volumes grow and autonomous navigation moves from concept to trial, the demand for trustworthy trajectory forecasts will only intensify. This research suggests that the smartest investment may lie in the unglamorous work of cleaning and structuring the data, done with genuine understanding of how ships actually move through the water.
Subject of Research: Machine learning-based vessel trajectory prediction from AIS data
Article Title: Kinematic-aware AIS preprocessing and bidirectional recurrent modeling for multi-output vessel trajectory prediction
Article References: Sinni, M., & Kyriazanos, D. M. (2026). Kinematic-aware AIS preprocessing and bidirectional recurrent modeling for multi-output vessel trajectory prediction. Journal of Big Data. https://doi.org/10.1186/s40537-026-01562-x
Image Credits: AI Generated
DOI: 10.1186/s40537-026-01562-x
Keywords: AIS, vessel trajectory prediction, BiGRU, maritime traffic monitoring, kinematic-aware preprocessing, recurrent neural networks, LSTM, GRU, data cleaning, course over ground, maritime surveillance, machine learning
Cite Scienmag News
Blake Davidson. (October 3, 2026). Smarter Data Cleaning and Bidirectional AI Sharpen Ship Trajectory Forecasts at Sea. Scienmag. https://scienmag.com/smarter-data-cleaning-and-bidirectional-ai-sharpen-ship-trajectory-forecasts-at-sea/
Blake Davidson. "Smarter Data Cleaning and Bidirectional AI Sharpen Ship Trajectory Forecasts at Sea." Scienmag, 3 October 2026, https://scienmag.com/smarter-data-cleaning-and-bidirectional-ai-sharpen-ship-trajectory-forecasts-at-sea/. Accessed 3 October 2026.
Blake Davidson. "Smarter Data Cleaning and Bidirectional AI Sharpen Ship Trajectory Forecasts at Sea." Scienmag. October 3, 2026. https://scienmag.com/smarter-data-cleaning-and-bidirectional-ai-sharpen-ship-trajectory-forecasts-at-sea/

