Researchers have developed a machine learning framework that could help water managers identify viral and bacterial contamination risks before conventional laboratory testing is complete. The approach combines environmental data with quantitative microbial risk assessment (QMRA), allowing routine measurements such as turbidity, temperature, dissolved oxygen and rainfall to be used to estimate pathogen concentrations and their potential effects on human health.
The study, published in Biocontaminant, focuses on a persistent challenge in drinking-water safety: the indicators traditionally used to track fecal contamination do not always accurately reflect the presence of viruses. Fecal coliforms, Escherichia coli and Enterococcus faecalis are widely monitored because they can signal contamination from sewage or animal waste. However, their concentrations may not rise and fall in parallel with viral pathogens, which can behave differently in aquatic environments and may remain infectious under conditions that affect bacteria in other ways.
To investigate whether routine water-quality information could improve pathogen surveillance, the researchers collected 95 surface-water samples from two drinking-water sources in a major city in eastern China. Sampling took place between May 2024 and December 2025. Each sample was analyzed for the three bacterial indicators as well as six pathogens: Pseudomonas aeruginosa, Salmonella species, Shigella species, adenovirus, norovirus and enterovirus. The inclusion of several viruses was particularly important because viral contamination can be difficult to infer from bacterial measurements alone and because some waterborne viruses can cause illness at relatively low exposure levels.
The researchers first examined the relationships among the organisms detected in the samples. The three fecal indicator bacteria were significantly correlated with one another, suggesting that they often reflected similar contamination patterns. Their relationships with the viral pathogens, however, were generally weak or inconsistent. This finding highlights a limitation in relying exclusively on bacterial indicators when assessing the safety of source water. A water sample with relatively modest bacterial counts could still warrant attention if other environmental conditions favor the persistence or transport of viruses.
The team compared six machine learning methods to determine how effectively physicochemical measurements could predict pathogen concentrations. The models included multiple linear regression, least-squares boosting, decision trees, support vector machines, random forests and multilayer perceptrons. These methods differ in how they identify relationships between variables: some assume relatively simple mathematical associations, while others can capture nonlinear interactions, thresholds and complex combinations of environmental conditions. Such flexibility is valuable in water systems, where rainfall, suspended particles, temperature and organic matter may influence different pathogens in different ways.
Random forest and decision tree models performed particularly well. After optimization, all of the models achieved coefficients of determination, or R² values, above 0.75, indicating that they explained a substantial proportion of the variation in the measured pathogen data. The decision tree model was especially effective for predicting P. aeruginosa, with an R² above 0.90. The researchers also tested the models against independent samples collected in January and February 2026. Most retained their predictive ability during this temporal validation, suggesting that the framework may be able to provide useful estimates beyond the period used for model training.
Prediction alone does not indicate whether a concentration represents a meaningful public-health threat, so the researchers connected the machine learning outputs to QMRA. This established risk-assessment approach estimates the probability that people will be exposed to a pathogen and the health burden that could result. The study expressed health impacts as disability-adjusted life years, or DALYs, a measure that combines years of life lost with years lived with illness or disability. Most estimated risks remained below the World Health Organization benchmark of 10⁻⁶ DALYs per person per year. Under unfavorable exposure conditions, however, Salmonella, Shigella and enterovirus showed probabilities of exceeding that benchmark.
The analysis identified disinfection efficiency as the most influential factor affecting estimated health risks. This result reinforces the importance of maintaining reliable treatment processes, particularly when source-water quality changes after heavy rainfall or other environmental disturbances. Effective disinfection can substantially reduce the concentration of infectious organisms reaching consumers, but its performance depends on factors such as disinfectant dose, contact time, water chemistry and the resistance of individual pathogens. Viruses and bacteria do not necessarily respond to treatment in the same way, making process control and pathogen-aware monitoring essential.
To make the models easier to interpret, the researchers used SHapley Additive exPlanations, known as SHAP. This technique estimates how much each input variable contributes to an individual prediction or to the model’s overall performance. Turbidity was one of the strongest predictors for the fecal indicator bacteria, accounting for between 41.6% and 62.1% of predictive importance in those models. Temperature, dissolved oxygen, rainfall and other measurements contributed differently depending on the pathogen being modeled. The authors emphasize that the framework is not intended to replace direct microbial testing, but to complement it by providing faster risk estimates from data that water utilities already collect. Wider validation across watersheds, seasons, treatment systems and land-use settings will be needed before the method can support routine operational decisions. If integrated with real-time sensors and early-warning systems, ML-QMRA could help identify periods of elevated viral risk sooner and guide more targeted interventions to protect drinking-water supplies.
Subject of Research: Machine learning and quantitative microbial risk assessment for predicting bacterial and viral pathogen risks in drinking-water sources
Article Title: A machine learning-quantitative microbial risk assessment (ML-QMRA) framework for predicting potential health risks from pathogens in drinking water sources
News Publication Date: 10-Jul-2026
Web References: https://doi.org/10.48130/biocontam-0026-0009
References: Guo B, Wang J, Yan C, Tang C, Yin K, et al. 2026. “A machine learning-quantitative microbial risk assessment (ML-QMRA) framework for predicting potential health risks from pathogens in drinking water sources.” Biocontaminant 2: e012. DOI: 10.48130/biocontam-0026-0009
Image Credits: Bingbing Guo, Jing Wang, Chicheng Yan, Chao Tang, Kun Yin, Lei Jiang and Changzheng Cui
Keywords: drinking water, waterborne viruses, viral pathogens, enterovirus, norovirus, adenovirus, machine learning, quantitative microbial risk assessment, QMRA, microbial contamination, water quality monitoring, public health, artificial intelligence, turbidity, disinfection

