Every day, thousands of Swedes pick up the phone and call the national healthcare guide, describing fevers, coughs, rashes, and stomach complaints to telehealth nurses. Buried in those call records is a signal that public health officials desperately want to read earlier and more accurately: the first stirrings of a disease outbreak. A new study published in Machine Learning with Applications by Atiye Sadat Hashemi of Lund University and colleagues presents a carefully engineered framework that combines classical epidemiological modeling with modern deep learning to detect such outbreaks in real time, while also confronting an uncomfortable truth about how anomaly detection research is usually judged.
The core problem the researchers tackle is deceptively simple to state but notoriously hard to solve. Time-series anomaly detection aims to flag observations that deviate sharply from normal behavior, and in surveillance data a surge in symptom-related calls may herald the onset of an epidemic. Yet the field is plagued by flawed evaluation practices. The widely used point-adjust protocol, for example, credits an algorithm with correctly identifying an entire anomalous segment if it detects just a single point within it, a leniency so extreme that even randomly generated anomaly scores can appear highly accurate under it. Point-wise metrics such as precision, recall, and F1 score also treat a late detection as equivalent to a complete miss and struggle to capture event-like anomalies that unfold over days or weeks. As the authors and prior benchmarking studies note, these weaknesses have inflated claims of progress and produced misleading rankings of detection methods.
To build a fairer testbed, the team turned to one of epidemiology’s oldest and most trusted tools: the SEIR model, which divides a population into susceptible, exposed, infectious, and recovered compartments governed by differential equations. Rather than trying to label real historical outbreaks, an exercise fraught with subjectivity and bias, they simulated three synthetic disease outbreaks and injected them into genuine surveillance data. Disease A was designed to resemble measles, with high transmissibility and symptoms including fever, cough, conjunctivitis, and rash. Disease B mimicked COVID-19, with moderate transmissibility and respiratory symptoms, while Disease C modeled mpox, a lower-transmission illness marked by fever, rash, and localized skin effects. Each simulated outbreak was embedded into the call-record time series starting at day 1500 of the dataset and lasting 100 days, with the transmission rate tuned so that roughly 1, 15, or 50 percent of the population recovered by the end of the simulation.
The translation from epidemic model to surveillance signal is where the framework shows its technical sophistication. For each symptom variable, the expected outbreak-related increase in calls was computed by summing, across the three modeled infectious stages, the number of infectious individuals multiplied by the probability that a person in that stage exhibits the symptom, the probability that a symptomatic individual calls the hotline, and a demographic weighting factor of 0.8 for adult variables and 0.2 for child variables. Because the injected signals are deterministic given the chosen parameters, the resulting benchmark is fully reproducible, allowing different detection algorithms to be compared on identical, objectively labeled outbreak events, something real-world surveillance data can almost never provide.
Against this benchmark, the researchers evaluated five reconstruction-based anomaly detectors: principal component analysis, a vanilla autoencoder, a variational autoencoder, an LSTM autoencoder, and a Transformer autoencoder. All of these learn a representation of normal temporal behavior and flag anomalies when reconstruction error, the mismatch between input data and the model’s attempt to reproduce it, spikes above a threshold set at the 99th percentile of training scores. The models were tested in two modes. In the offline setting, the entire time series is analyzed retrospectively. In the online setting, which mirrors real deployment, the model starts with one year of training data and then processes incoming batches sequentially, incorporating each evaluated batch into its training set and recalculating its anomaly threshold, an incremental expanding-window strategy chosen because ground-truth labels are unavailable during real surveillance.
The results reveal a clear hierarchy among the architectures and an honest picture of the difficulty of the task. In offline testing, every model detected all three injected outbreaks at the event level, but the Transformer autoencoder led consistently, achieving the highest point recall for all three diseases and the fastest time to detection, flagging Disease A after 61 days, compared with 74 to 78 days for the other models. In the harder online setting, the Transformer autoencoder again performed best on Disease A, with a point recall of 0.395 and detection after 60 days. For the weaker signals of Diseases B and C, however, its advantage narrowed or vanished: the LSTM autoencoder matched it on Disease B, and only those two models managed to detect Disease C at all. The authors attribute the Transformer’s edge to its ability to capture complex temporal and cross-variable dependencies, while noting that the variational autoencoder’s probabilistic constraints can smooth out subtle anomalies and that PCA, limited to linear relationships, was least sensitive.
Sensitivity analyses added crucial nuance. Lower transmission rates degraded detection across the board, and at the lowest rate only the Transformer autoencoder still registered the outbreak. More strikingly, the timing of the outbreak within the surveillance record mattered enormously. When a measles-like outbreak was injected during the early pandemic phase of the data, the online Transformer detector achieved a point recall of 0.722 at baseline reporting probability, and up to 0.950 when the reporting probability was raised. The same outbreak injected in the post-pandemic period yielded a recall of only about 0.4 and a detection delay of 60 days, and at a reduced susceptible fraction of 50 percent it went entirely undetected. Outbreak detectability, in other words, depends not just on the pathogen but on the surrounding statistical background, the signal-to-noise ratio of the moment.
To test the framework against reality, the team also ran the detectors on the unmodified dataset, which contains four expert-identified COVID-19 waves and recurrent seasonal outbreaks of influenza and norovirus between 2019 and mid-2023. Using DBSCAN clustering to group point-wise anomaly detections into temporally coherent events, the two best models, the Transformer and LSTM autoencoders, overlapped nearly all known outbreak periods, including all four norovirus seasons and three influenza seasons, though the Transformer missed the second COVID-19 wave entirely and both models detected only fragments of the fourth. Detections that matched no known outbreak were reported as unmatched clusters rather than dismissed as false positives, a deliberate design choice reflecting the fact that real surveillance data may contain undocumented epidemiological events, seasonal behavioral shifts, or reporting artifacts that no complete ground truth can capture.
The authors are candid about the framework’s limitations. The SEIR-based injection assumes a fully mixed population and simplified symptom-reporting behavior, so it cannot reproduce the full complexity of real epidemics, and the ground-truth onset is defined as the true start of transmission in the simulation rather than the point where changes become visible in the data, which deliberately penalizes the detectors and keeps performance metrics conservative. Keeping potentially anomalous observations in the online training stream risks contamination, whereby a prolonged outbreak gradually becomes absorbed into the learned definition of normal. The comparison was also restricted to reconstruction-based methods, so no claim of superiority over classical statistical surveillance tools such as CUSUM and EWMA, or over forecasting-based approaches, is being made. Future work, the team writes, should compare these families of methods directly, explore spatiotemporal neural networks to identify at-risk municipalities, and add explainability techniques to make the detections interpretable for public health decision-makers.
What the study ultimately delivers is less a single winning algorithm than a discipline. By pairing epidemiologically grounded synthetic outbreaks with event-oriented evaluation metrics, point recall, event recall, time to detection, and an overlap duration coefficient analogous to the Intersection over Union measure, the framework offers the outbreak-detection field something it has badly needed: a benchmark in which performance claims can be trusted, timeliness is measured honestly, and the metric matches the public health reality that interventions target outbreak episodes, not isolated data points. As syndromic surveillance streams grow larger and faster, frameworks of this kind may determine whether the next emerging pathogen is noticed in days rather than weeks.
Subject of Research: Real-time detection of infectious disease outbreaks in symptom surveillance data using epidemiology-guided machine learning anomaly detection
Article Title: An epidemiology-guided machine learning framework for real-time anomaly detection in disease outbreak monitoring using symptom surveillance data
Article References: Hashemi, A. S., Dietler, D., Ohlsson, M., & Björk, J. (2026). An epidemiology-guided machine learning framework for real-time anomaly detection in disease outbreak monitoring using symptom surveillance data. Machine Learning with Applications, 26, Article 101015. https://doi.org/10.1016/j.mlwa.2026.101015
Image Credits: AI Generated
DOI: 10.1016/j.mlwa.2026.101015
Keywords: anomaly detection, disease surveillance, SEIR model, syndromic surveillance, machine learning, Transformer autoencoder, LSTM autoencoder, outbreak detection, time series analysis, evaluation metrics, DBSCAN clustering, public health
Cite Scienmag News
Phoebe Ingram. (October 1, 2026). Machine Learning Meets Epidemiology to Catch Disease Outbreaks in Real Time. Scienmag. https://scienmag.com/machine-learning-meets-epidemiology-to-catch-disease-outbreaks-in-real-time/
Phoebe Ingram. "Machine Learning Meets Epidemiology to Catch Disease Outbreaks in Real Time." Scienmag, 1 October 2026, https://scienmag.com/machine-learning-meets-epidemiology-to-catch-disease-outbreaks-in-real-time/. Accessed 1 October 2026.
Phoebe Ingram. "Machine Learning Meets Epidemiology to Catch Disease Outbreaks in Real Time." Scienmag. October 1, 2026. https://scienmag.com/machine-learning-meets-epidemiology-to-catch-disease-outbreaks-in-real-time/

