Traffic is one of the hardest signals a city produces. It surges and stalls in waves, obeys the slow rhythm of the morning commute one moment and collapses into a sudden jam the next, and it is constantly corrupted by sensor noise, weather, accidents and the sheer unpredictability of human behavior. For the engineers who build intelligent transportation systems, the ability to see even ten minutes into the future with confidence is the difference between smoothing a rush hour and being trapped inside one. A new study published in Applied Intelligence by Wenjing Wang and Deqian Fu of Linyi University in China argues that the missing ingredient is a way of listening to traffic at multiple scales at once, and their answer borrows two of the most powerful ideas in modern machine learning: the wavelet transform and the denoising diffusion model.
The core problem the researchers set out to solve is a familiar one in time series analysis. Traditional forecasting tools, from seasonal ARIMA models to Kalman filters, treat traffic data as a single stream and try to fit one model to all of its behavior. That works reasonably well when the road is quiet, but it breaks down precisely when it matters most, because the abrupt, short-lived fluctuations that signal an emerging bottleneck are statistically very different from the long, predictable swells of daily demand. Deep learning approaches, including convolutional and recurrent architectures, transformers and graph neural networks, have pushed accuracy forward, yet they still tend to blend these different temporal scales together rather than separating them. The result is a systematic compromise: models that capture the trend but miss the spike, or that react to noise as if it were a trend.
Wang and Fu’s framework, called the Multiscale Temporal-Spatial Diffusion Model informed by Wavelet Transform, or MTSDM/WT, attacks this compromise at its root. The first step is wavelet decomposition, a mathematical technique from signal processing that splits a time series into components of different frequency. In the traffic setting, the high-frequency components carry the short-term fluctuations, the sudden braking waves and transient congestion events that appear and vanish within minutes, while the low-frequency components encode the slow, recurring structure of the day, such as the predictable build-up and decay of rush-hour patterns. By separating these regimes before any prediction is attempted, the model can apply a forecasting strategy tailored to each one instead of forcing a single architecture to do both jobs at once.
The treatment of the two bands is deliberately asymmetric, and this is where the technical novelty of the paper lies. The high-frequency components are handed to a High-Frequency Refinement Module, which combines multiscale convolutional filters with attention mechanisms. The convolutions sweep across the signal at several receptive-field sizes, allowing the module to detect localized temporal patterns of varying duration, while the attention layers learn which of those patterns deserve weight at each moment. This design is aimed squarely at transient dynamics: the brief, sharp deviations from trend that conventional smoothers tend to average away. Because the module operates only on the high-frequency band, it does not have to disentangle these spikes from the underlying daily cycle, which makes the learning problem substantially cleaner.
The low-frequency components follow a completely different path through the architecture: they are processed by a denoising diffusion model. Diffusion models, the same family of generative models behind recent breakthroughs in image synthesis, work by progressively corrupting data with noise in a forward process and then training a network to reverse that corruption step by step. Applied to traffic trends, the forward diffusion gradually obscures the smooth low-frequency signal, and the reverse denoising process learns the distribution that the trends actually follow, reconstructing plausible future trajectories rather than producing a single brittle point estimate. This probabilistic character matters for planning, because a traffic management center that knows the range of likely outcomes can allocate resources more defensively than one that receives only a mean prediction.
Once the high-frequency band has been refined and the low-frequency band has been reconstructed, the two streams are recombined through the inverse wavelet transform, which is the exact mathematical inverse of the decomposition step. This guarantees that the fusion is not an ad hoc averaging of two separate forecasts but a principled reconstruction of a single coherent signal, with each band contributing the aspects of traffic it models best. The final output is a comprehensive prediction that preserves both the long-term structure and the fine-grained dynamics of the original series, and the authors argue that this separation-and-recombination principle is what gives the model its robustness under noisy, fluctuating conditions.
The empirical evaluation was carried out on real-world data, including private traffic flow records from four key urban intersections in Linyi City in Shandong Province, labeled GJ, GC, MJ and MC, as well as the publicly available PeMS04, PeMS07 and PeMS08 highway datasets from California’s Performance Measurement System. The authors assessed performance with the standard battery of forecasting metrics: Mean Absolute Error, Mean Absolute Percentage Error and Root Mean Squared Error, which measure the accuracy of point predictions, together with the Continuous Ranked Probability Score, which evaluates the quality of the full probabilistic forecast. Across these metrics, MTSDM/WT showed consistent gains over competing approaches, with a representative improvement of roughly 3 percent on the MJ dataset at a ten-minute prediction horizon.
Three percent may sound modest, but in the economics of traffic operations it is meaningful. Short-term forecasts feed directly into signal timing optimization, ramp metering, route guidance and incident response, and errors compound as prediction horizons lengthen. A model that is simultaneously more accurate and better calibrated reduces both the false alarms that erode operator trust and the missed surges that turn minor disruptions into gridlock. The paper also emphasizes resilience: because the diffusion component models the distribution of trends rather than a single path, and because the high-frequency module is trained specifically on noisy transient behavior, the framework degrades more gracefully than deterministic baselines when sensors drop out or conditions deviate from the historical norm.
Beyond its immediate practical value, the study contributes to a broader research current that is reshaping spatiotemporal machine learning. Diffusion models have recently spread from image generation into time series imputation, trajectory synthesis and urban flow inference, and the authors position their work as evidence that these generative tools are especially well suited to the multiscale structure of real-world signals. Pairing them with wavelet analysis offers a template that could transfer to other domains where fast events ride on slow cycles, from network traffic and electricity load to hydrology and financial markets. The work was supported by the Taishan Industrial Experts Program and the Shandong Provincial Natural Science Foundation, reflecting the provincial interest in deploying advanced artificial intelligence in regional infrastructure.
There remain, of course, the usual caveats that accompany any new forecasting architecture. The Linyi intersection data are private and cannot be independently examined, so external validation rests on the public PeMS benchmarks, and the reported gains, while consistent, will need to be reproduced by other groups on other networks before the approach becomes standard practice. Diffusion models also carry a computational cost at inference time, since generating a forecast requires running the reverse denoising chain, and deployment in latency-critical traffic control loops will depend on how efficiently that chain can be shortened. Still, the central insight of the paper is likely to endure: when a signal lives on many time scales at once, the smartest model is not one that averages across them, but one that decomposes the problem, predicts each scale with the tool it deserves, and then puts the pieces back together exactly.
Subject of Research: A wavelet-guided multiscale diffusion model for short-term traffic flow prediction
Article Title: Traffic flow prediction using multiscale diffusion model guided by wavelet transform
Article References: Wang, W., & Fu, D. (2026). Traffic flow prediction using multiscale diffusion model guided by wavelet transform. Applied Intelligence, 56(14), Article 402. https://doi.org/10.1007/s10489-026-07408-6
Image Credits: AI Generated
DOI: 10.1007/s10489-026-07408-6
Keywords: traffic flow prediction, diffusion model, wavelet transform, time series forecasting, intelligent transportation systems, deep learning, spatiotemporal modeling, denoising diffusion, signal processing, urban traffic management, probabilistic forecasting, Applied Intelligence
Cite Scienmag News
Blake Davidson. (October 8, 2026). Wavelet-Guided Diffusion Model Sharpens Urban Traffic Flow Forecasts. Scienmag. https://scienmag.com/wavelet-guided-diffusion-model-sharpens-urban-traffic-flow-forecasts/
Blake Davidson. "Wavelet-Guided Diffusion Model Sharpens Urban Traffic Flow Forecasts." Scienmag, 8 October 2026, https://scienmag.com/wavelet-guided-diffusion-model-sharpens-urban-traffic-flow-forecasts/. Accessed 8 October 2026.
Blake Davidson. "Wavelet-Guided Diffusion Model Sharpens Urban Traffic Flow Forecasts." Scienmag. October 8, 2026. https://scienmag.com/wavelet-guided-diffusion-model-sharpens-urban-traffic-flow-forecasts/

