<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI outperforming GANs in microbial prediction &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-outperforming-gans-in-microbial-prediction/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 06:47:10 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI outperforming GANs in microbial prediction &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Simulator Predicts Bacteria in Wastewater From Simple Measurements, Beating GANs by 35 Percent</title>
		<link>https://scienmag.com/ai-simulator-predicts-bacteria-in-wastewater-from-simple-measurements-beating-gans-by-35-percent/</link>
		
		<dc:creator><![CDATA[Morgan Morrow]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 06:47:10 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI outperforming GANs in microbial prediction]]></category>
		<category><![CDATA[AI wastewater bacteria detection]]></category>
		<category><![CDATA[AI-based sewage microbiome analysis]]></category>
		<category><![CDATA[bacterial concentration prediction]]></category>
		<category><![CDATA[bioreactor bacteria estimation software]]></category>
		<category><![CDATA[environmental health monitoring with AI]]></category>
		<category><![CDATA[Environmental Monitoring]]></category>
		<category><![CDATA[generative adversarial networks]]></category>
		<category><![CDATA[lifelong learning]]></category>
		<category><![CDATA[LSTM networks]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for microbial concentration]]></category>
		<category><![CDATA[Markov Chains]]></category>
		<category><![CDATA[membrane bioreactors]]></category>
		<category><![CDATA[open-source water treatment tools]]></category>
		<category><![CDATA[predictive modeling for water quality]]></category>
		<category><![CDATA[real-time water quality prediction]]></category>
		<category><![CDATA[routine physicochemical data in water quality assessment]]></category>
		<category><![CDATA[sensor data-driven wastewater analysis]]></category>
		<category><![CDATA[soft sensor in bioreactor monitoring]]></category>
		<category><![CDATA[soft sensors]]></category>
		<category><![CDATA[synthetic data generation]]></category>
		<category><![CDATA[wastewater treatment]]></category>
		<category><![CDATA[water-based epidemiology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233958</guid>

					<description><![CDATA[Researchers at KAUST have developed SALS-BioC, an open-source AI simulator that predicts bacterial concentrations in wastewater treatment plants from routine physicochemical measurements, outperforming GAN-based approaches by 35 percent through a novel synthetic data generator and lifelong learning framework.]]></description>
										<content:encoded><![CDATA[<p>Every day, wastewater treatment plants around the world process billions of liters of sewage, and hidden in that water are bacteria whose concentrations tell a critical story about public health, treatment efficiency, and environmental risk. The problem is that counting those microbes has always been slow, expensive, and laborious. Culture-based assays and flow cytometry require trained technicians, specialized equipment, and days of waiting, which means that by the time a contamination event is detected, the water has long since moved on. A team of researchers at King Abdullah University of Science and Technology (KAUST) in Saudi Arabia now believes artificial intelligence can close that gap, and they have released an open-source software tool designed to prove it.</p>
<p>The tool, called SALS-BioC, is described in the journal SoftwareX as a soft-sensor based adaptation learning simulator for predicting bacterial concentrations in membrane bioreactors. A soft sensor, in the engineering sense, is a machine learning model that estimates hard-to-measure quality variables from variables that are easy to measure continuously. In this case, the system takes routine physicochemical water quality readings such as pH, electrical conductivity, total suspended solids, biochemical oxygen demand, chemical oxygen demand, nitrate nitrogen, turbidity, and chlorine levels, and uses them to predict how many bacterial cells are present in the treated effluent. The idea is to replace days of laboratory work with an instant, data-driven estimate that plant operators can act on in real time.</p>
<p>The research team, composed of H. Bagci, I. N&#8217;Doye, F. Almulhim, and P.-Y. Hong, confronted a familiar obstacle in environmental machine learning: there is simply not enough data. Wastewater treatment plants face cost and confidentiality constraints that limit access to large datasets from water-based epidemiology studies, and microbial measurements are inherently scarce because they depend on weekly sampling campaigns. Machine learning models trained on small, static datasets tend to perform well in the laboratory but fail when confronted with new conditions, a problem known as poor generalization. When the researchers tested a conventional long short-term memory (LSTM) neural network on data from a biologically independent sampling period, the model achieved an impressive coefficient of determination of 0.96 on its own training and testing data, but collapsed to a negative value of minus 0.33 on the unseen replicate, meaning its predictions were worse than a naive average.</p>
<p>To solve the data scarcity problem, the team developed a novel synthetic data generation algorithm called EMCM-PS, short for an ensemble Markov chain model with a joint state representation and probabilistic sampling rule. The approach is a creative departure from mainstream generative techniques. Instead of using generative adversarial networks, which pit two neural networks against each other in a training game and often suffer from a failure mode known as mode collapse, EMCM-PS builds on Markov chains and Gaussian copulas. Each multivariate observation in the wastewater dataset is encoded as a single joint state, and new synthetic samples are drawn directly from the empirical probability distribution of those observed states rather than from a sequential transition matrix. The researchers argue this design choice matters because wastewater samples are collected at irregular intervals and do not form a strict time series, so forcing a random-walk structure onto the data would introduce spurious dependencies and trap the generator in a limited set of states.</p>
<p>The generation workflow begins by augmenting the original dataset with correlation-aware multiplicative noise, then discretizing the samples into quantile-based joint states. Synthetic states are sampled independently from the global frequency distribution of observed states and decoded back into continuous measurements using either a local mode, which samples from the observations associated with each state, or a uniform mode, which draws values within the bounds of the corresponding quantile bucket. The software includes a suite of visual diagnostics to verify that the synthetic data faithfully reproduce the statistical structure of the real measurements, including marginal distribution comparisons, Pearson correlation heatmaps, and dimensionality-reduction plots based on principal component analysis and t-distributed stochastic neighbor embedding. A fidelity scorecard summarizes these metrics so users can judge at a glance whether the generated data preserve the variability and dependency patterns of the original samples.</p>
<p>On top of this synthetic data engine sits the second key innovation: a lifelong learning framework built around the LSTM predictor. Unlike traditional machine learning models that are trained once on a static historical dataset and then frozen, the lifelong learning approach continuously updates the model as new batches of data arrive from the target environment, while preserving previously acquired knowledge through a dictionary learning mechanism. In the SALS-BioC workflow, a source-domain model is first trained on historical data, then adapted to a target-domain dataset processed in sequential batches of thirty samples. For each batch, the model predicts first and is updated with the true values only afterward, mimicking the realistic operational scenario in which laboratory confirmation lags behind the need for a prediction. The software even provides an animated visualization of how the root mean square error evolves batch by batch during adaptation, showing how quickly the model recovers its accuracy as it encounters the new domain.</p>
<p>The validation experiment drew on real data from the KAUST aerobic membrane bioreactor wastewater treatment plant, which treats a mix of municipal wastewater. Two biologically independent replicates were used, one collected weekly from July to October 2023 for model development and another from February to March 2024 to assess generalization. Fourteen physicochemical parameters were measured alongside bacterial abundances quantified by flow cytometry. When the lifelong learning model, trained on EMCM-PS-generated synthetic data, was adapted to the unseen replicate, it achieved a cross-validation coefficient of determination of 0.8723. When the same experiment was run using synthetic data from a Wasserstein generative adversarial network, the best-performing member of the GAN family, the score reached only 0.6450. That difference of 0.2273 translates into a 35.2 percent improvement for the new probabilistic approach, a striking margin in a field where incremental gains are the norm.</p>
<p>The software itself is designed to be accessible rather than the exclusive province of machine learning specialists. It is implemented as a modular Python application with a browser-based interface built on Flask, HTML, CSS, and JavaScript, using standard scientific libraries including NumPy, pandas, SciPy, scikit-learn, and PyTorch. Users upload CSV datasets, select target variables, configure parameters, and run simulations through three dashboard sections covering synthetic data generation, LSTM-based prediction, and lifelong learning adaptation. Trained models can be downloaded as serialized PyTorch files and reloaded later for validation without retraining, and Bayesian hyperparameter optimization is built in with configurable search iterations. The code is released under the MIT license on GitHub, and the authors note that the architecture deliberately separates the interface, the server orchestration, and the algorithms, so that new generators, predictive models, or adaptation strategies can be added through dedicated API routes without redesigning the application. Extensions to viral contaminant prediction are described as a natural next step.</p>
<p>The implications reach beyond one treatment plant in Saudi Arabia. Water-based epidemiology has surged in prominence since the COVID-19 pandemic demonstrated that sewage can serve as an early-warning system for disease outbreaks in entire communities, but the field remains bottlenecked by the pace of laboratory analysis. A reliable soft sensor that predicts bacterial concentrations from measurements plants already collect could enable continuous biological monitoring, early diagnosis of operational faults, and rapid response to contamination events, all without waiting for culture results. The authors also emphasize the educational value of the tool, arguing that its guided interface lowers the expertise required to design data-driven soft sensor systems and makes the technology accessible to interdisciplinary researchers and students.</p>
<p>The researchers are candid about limitations. EMCM-PS currently uses a fixed number of quantile-based buckets for all features, but different water quality variables have different distributions and may require different discretization resolutions, particularly under operational variations such as fluctuations in flow rate or for highly skewed measurements like total cell counts. Future work, they suggest, could develop an adaptive discretization strategy that tunes the bucket count for each feature dynamically. They also plan to validate the simulator across international wastewater datasets to establish its transferability. For now, SALS-BioC stands as a concrete demonstration that thoughtful statistical modeling, rather than ever-larger neural networks, can sometimes deliver the biggest wins when data are scarce, and that open, reproducible software may be the fastest route to putting adaptive artificial intelligence into the pipes and pumps of the world&#8217;s water infrastructure.</p>
<p><strong>Subject of Research:</strong> Machine learning-based soft sensor prediction of bacterial concentrations in membrane bioreactor wastewater treatment</p>
<p><strong>Article Title:</strong> SALS-BioC: A soft-sensor based adaptation learning simulator for predicting bacterial concentrations in membrane bioreactors</p>
<p><strong>Article References:</strong> Bagci, H., N’Doye, I., Almulhim, F., &amp; Hong, P.-Y. (2026). SALS-BioC: A soft-sensor based adaptation learning simulator for predicting bacterial concentrations in membrane bioreactors. <em>SoftwareX, 36</em>, Article 103084. <a href="https://doi.org/10.1016/j.softx.2026.103084" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103084</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103084" rel="noopener noreferrer">10.1016/j.softx.2026.103084</a></p>
<p><strong>Keywords:</strong> soft sensors, machine learning, wastewater treatment, membrane bioreactors, synthetic data generation, Markov chains, LSTM networks, lifelong learning, water-based epidemiology, bacterial concentration prediction, generative adversarial networks, environmental monitoring</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233958</post-id>	</item>
	</channel>
</rss>
