Missing data is one of the most stubborn problems in modern data science. Sensors fail, communication links drop, medical monitors go offline for maintenance, and the streams of measurements that power everything from air quality forecasts to hospital early-warning systems come back riddled with holes. A new study published in Applied Intelligence by Siyuan Ma, Ziwei Xu, and Ryutaro Ichise of the Institute of Science Tokyo takes aim at this problem with a fresh twist on one of artificial intelligence’s hottest tools: the diffusion model. Their method, called the Implicit Trajectory-Constrained Diffusion network, or ITCD, delivers an average improvement of roughly 13 percent over existing imputation baselines across a wide range of scenarios, while also cutting inference time dramatically.
Time series data, sequences of measurements ordered in time, underpin countless real-world applications, including clinical diagnostics, air quality monitoring, electricity load forecasting, anomaly detection, and emergency event prediction. Because missing values are both common and inevitable, the quality of any downstream analysis depends heavily on how well those gaps can be filled. Traditional approaches relied on statistical techniques such as ARIMA or simple substitution with zeros or means, while later machine learning models like BRITS and GRU-D brought recurrent neural networks to the task. More recently, Transformer-based architectures such as SAITS and ImputeFormer have set strong benchmarks. Yet all of these methods share a common limitation: they treat imputation as a deterministic estimation problem, producing a single best guess while ignoring the genuine uncertainty that surrounds any missing measurement.
Diffusion models, the same generative technology behind today’s most striking image generators, offer a natural alternative. They work by gradually adding Gaussian noise to data in a forward process and then learning to reverse that process, step by step, transforming pure noise into realistic samples. Because they model entire probability distributions rather than point estimates, diffusion models can capture the uncertainty of missing values and even provide confidence bounds around their predictions. Models like CSDI, SSSD, and SPD have already applied this idea to time series imputation with promising results. But according to the Tokyo team, these existing approaches suffer from two fundamental weaknesses that undermine the consistency of their output.
The first weakness is what the researchers call padding-induced distribution shift. To train a neural network on incomplete data, missing entries must be filled with placeholder values, typically zeros, so that the data can flow through the network’s mathematical machinery. The trouble is that diffusion models learn to approximate the distribution of whatever they are fed during training. When a large fraction of the training inputs consists of artificial zeros rather than genuine measurements, the model’s learned target distribution drifts away from the true data distribution. The effect is far from trivial: the team’s exploratory analysis showed that simply swapping zero-padding for one-padding degraded performance by 54.7 percent on one benchmark, while switching from sample-mean padding to linear interpolation improved results by 62.8 percent. In other words, the seemingly mundane choice of how to fill in blanks before training can swing accuracy dramatically.
The second weakness lies in how diffusion models are optimized. Standard diffusion training is entirely noise-driven: the network learns to predict the Gaussian noise that was added at each step, following the distributional dynamics of the reverse process. But this scheme provides no direct supervision from the actual observed values in the data. The result, the researchers show, is an unstable denoising trajectory, the sequence of intermediate states the model passes through as it walks from pure noise back to a clean sample. Visualizations of the popular CSDI model reveal oscillatory, erratic trajectories with large errors persisting deep into the reverse process, which ultimately produces imputations that can be inconsistent with the observed data surrounding the gaps.
ITCD tackles both problems head-on with a two-part design. The first component is a two-stage diffusion architecture. In the Refinement Stage, the model trains on padding-corrupted samples but uses its own probabilistic generations to replace the artificial padding with distribution-consistent values. These refined samples then serve as training inputs for the Core Stage, a second diffusion model that learns on data far closer to the true underlying distribution. Because the Core Stage is the only component used at inference time, the Refinement Stage adds no computational cost when the trained model is deployed. Ablation experiments confirm the value of this design: adding the Refinement Stage reduced error by 18.6 percent on the PeMS traffic dataset and 65.4 percent on the Electricity dataset in one comparison, and by 14.3 percent on Beijing Air quality data in another.
The second component, Implicit Trajectory-Constrained Optimization, reshapes how the denoising process itself is guided. It introduces two complementary constraints. The Intermediate Constraint supervises the predicted noise at every step of the reverse process, and crucially, it does so across all positions in the data, not just the missing ones. Existing methods typically focus only on the missing regions, discarding the rich distributional information contained in the observed values. By minimizing the mean absolute error between predicted and added noise everywhere, the model is pushed to align its entire denoising trajectory with the true data distribution. The Terminal Constraint then supervises the fully reconstructed sample at the final reverse step, combining an Observed Reconstruction Task that learns directly from visible measurements with a Masked Imputation Task that forces accurate prediction of artificially hidden values. Together, these constraints implicitly steer the model along a coherent path from noise to data, without requiring any extra learnable parameters.
The architecture also includes a clever structural safeguard during sampling. Before each denoising step, the entries corresponding to observed positions are explicitly refreshed with the true conditional values, preventing inaccurate intermediate estimates from contaminating the missing regions and reducing error accumulation along the trajectory. A small amount of stochastic noise is deliberately retained at each step to keep the trajectory from collapsing into a rigid deterministic path, which strengthens the robustness of the imposed constraints. The denoising network itself uses attention mechanisms in two directions simultaneously, across time and across features, with a learned gating mechanism that blends the two perspectives and incorporates information about the current diffusion step.
The experimental evidence is striking. Tested on eight datasets spanning air quality, traffic, electricity, and healthcare, including Beijing Air, Italy Air, PeMS, Melbourne Pedestrian, ETTh1, Electricity, and two PhysioNet clinical benchmarks, ITCD outperformed a broad field of baselines built on Transformers, recurrent networks, convolutional networks, and generative models. Under 10 percent point missingness, it cut error by 35.6 percent on Italy Air, 22.7 percent on PeMS, and 12.5 percent on ETTh1 relative to the best competitors. Its advantage grew as missingness increased: while CSDI’s error on Beijing Air ballooned from 0.144 to 0.640 as missing ratios rose from low to 90 percent, ITCD’s curve stayed remarkably flat. The method also proved robust across point, subsequence, and block missing patterns, and across repeated random sampling runs it showed substantially lower variance and better distributional alignment with ground truth than CSDI, as measured by the Wasserstein distance.
Perhaps most surprising for a diffusion model, ITCD is fast. Because the trajectory constraints stabilize the reverse process, the model needs only about ten diffusion steps rather than the fifty typically used by competitors, and it still matches or exceeds their accuracy. That translates into an 89.1 percent reduction in inference time on the ETT dataset and a 52.8-fold acceleration on PhysioNet2012, and it even beats the recurrent BRITS model on most benchmarks. As a probabilistic model, ITCD also provides meaningful uncertainty bounds around its predictions, giving practitioners a more honest picture of what remains unknown. The researchers, whose work was supported by the Japan Science and Technology Agency, have released their code publicly and plan to explore how imputation quality affects downstream tasks and to investigate explicit trajectory search strategies. For a field drowning in incomplete data, a method that is simultaneously more accurate, more stable, and dramatically faster is likely to attract attention well beyond the machine learning community.
Subject of Research: Diffusion-based deep learning for consistent imputation of missing values in multivariate time series
Article Title: ITCD: implicit trajectory-constrained diffusion for consistent time series imputation
Article References: Ma, S., Xu, Z., & Ichise, R. (2026). ITCD: implicit trajectory-constrained diffusion for consistent time series imputation. Applied Intelligence, 56(14), Article 397. https://doi.org/10.1007/s10489-026-07427-3
Image Credits: AI Generated
DOI: 10.1007/s10489-026-07427-3
Keywords: time series imputation, diffusion models, missing data, deep learning, denoising trajectory, generative models, machine learning, uncertainty quantification, neural networks, Applied Intelligence, conditional diffusion, data science
Cite Scienmag News
Blake Davidson. (October 8, 2026). New Diffusion Model Fills the Gaps in Broken Time Series Data. Scienmag. https://scienmag.com/new-diffusion-model-fills-the-gaps-in-broken-time-series-data/
Blake Davidson. "New Diffusion Model Fills the Gaps in Broken Time Series Data." Scienmag, 8 October 2026, https://scienmag.com/new-diffusion-model-fills-the-gaps-in-broken-time-series-data/. Accessed 8 October 2026.
Blake Davidson. "New Diffusion Model Fills the Gaps in Broken Time Series Data." Scienmag. October 8, 2026. https://scienmag.com/new-diffusion-model-fills-the-gaps-in-broken-time-series-data/

