<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>BiLSTM &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/bilstm/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 21:57:59 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>BiLSTM &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Deep Learning Map Gives Greenhouses a Live 3D View of Heat and Humidity</title>
		<link>https://scienmag.com/deep-learning-map-gives-greenhouses-a-live-3d-view-of-heat-and-humidity/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 21:57:59 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[3D heat and humidity visualization in solar greenhouses]]></category>
		<category><![CDATA[advanced sensor networks for greenhouse climate monitoring]]></category>
		<category><![CDATA[AI-driven analysis of greenhouse microclimates]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[continuous 3D climate mapping in protected agriculture]]></category>
		<category><![CDATA[convolutional neural network]]></category>
		<category><![CDATA[data-driven greenhouse climate optimization]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning frameworks for greenhouse environmental control]]></category>
		<category><![CDATA[Deep learning greenhouse microclimate mapping]]></category>
		<category><![CDATA[digital twin]]></category>
		<category><![CDATA[humidity prediction]]></category>
		<category><![CDATA[innovative climate management in large-scale greenhouses]]></category>
		<category><![CDATA[microclimate]]></category>
		<category><![CDATA[precision agriculture]]></category>
		<category><![CDATA[real-time heat and humidity mapping in solar greenhouses]]></category>
		<category><![CDATA[sensor networks]]></category>
		<category><![CDATA[sensor-based microclimate monitoring in Chinese solar greenhouses]]></category>
		<category><![CDATA[solar greenhouse]]></category>
		<category><![CDATA[solar greenhouse architecture and microclimate]]></category>
		<category><![CDATA[spatial heterogeneity]]></category>
		<category><![CDATA[temperature prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=219366</guid>

					<description><![CDATA[Researchers in Beijing built a 27-sensor deep learning system that predicts and reconstructs three-dimensional temperature and humidity fields inside solar greenhouses, exposing hidden microclimates that single-point sensors miss.]]></description>
										<content:encoded><![CDATA[<p>Inside a Chinese solar greenhouse, the air is not one climate but many. On a sunny winter afternoon, the temperature near the translucent south-facing roof can be several degrees warmer than the air at the base of the crop canopy, while pockets of stagnant, humid air gather in the dense foliage where ventilation never quite reaches. For decades, growers have managed this invisible landscape with a single sensor hanging near the center of the structure, an approach that works well enough when plants are small but becomes increasingly misleading as the canopy fills the growing volume. A new study published in Smart Agricultural Technology argues that this one-point monitoring strategy is no longer good enough, and it offers a data-driven alternative: a deep learning framework that turns dozens of cheap sensors into a continuous, three-dimensional, forward-looking map of the entire greenhouse microclimate.</p>
<p>The research, led by Dongyu Wang and colleagues, was conducted in a commercial-style solar greenhouse at Baiwang Plantation in Beijing&#8217;s Haidian District. Solar greenhouses are a distinctive form of protected agriculture widely adopted in northern China, built around an asymmetric single-slope design: a light-transmitting south-facing roof, a massive heat-storing north wall, and insulated sidewalls. This architecture allows vegetables to be grown through cold seasons with minimal external energy input, but it comes with a thermodynamic price. The enclosure&#8217;s uneven heating produces strong nonlinear fluctuations, pronounced time-lag effects, and sharp three-dimensional spatial heterogeneity in temperature, humidity, and radiation, all of which complicate any attempt to characterize the environment from a single measurement point.</p>
<p>To capture that heterogeneity directly, the team instrumented the greenhouse with 27 temperature and humidity sensors arranged in three experimental zones along the structure&#8217;s 80-meter length. In each zone, sensors formed a 3-by-3 horizontal grid with 2-meter spacing, mounted at three heights: 0.5, 1.0, and 1.5 meters above the ground. An outdoor weather station recorded air temperature, humidity, pressure, solar radiation, wind, and rainfall. Everything was logged at 15-minute intervals from April through July 2024, spanning a full cucumber growth cycle from transplanting to maturity, and yielding more than 386,000 valid observation records. The crop was grown in north-south double rows on large ridges with drip irrigation under plastic sheeting, a standard commercial configuration that makes the findings relevant to real production settings.</p>
<p>The monitoring data revealed just how dramatically the microclimate changes as the crop develops. During the seedling stage, when plants were short, temperature curves at different sensor nodes were nearly synchronized, with daytime peak differences of only about 1.4 to 1.8 degrees Celsius. At that point, a single central sensor was still a reasonably fair representative of the whole space. But as the plants grew to roughly 1.22 meters, vertical stratification emerged between the lower and middle layers, and daytime temperature differences among nodes widened to about 4 degrees. By the maturity stage, with plants reaching approximately 1.85 meters, the maximum instantaneous vertical temperature difference hit 8 degrees Celsius, while relative humidity at the canopy bottom frequently stayed above 85 percent. The lower canopy had effectively become a separate, humid, disease-friendly microclimate that a central sensor could not see.</p>
<p>That observation motivated the core of the study: a neural architecture the authors call CNN-BiLSTM-SA, which combines three complementary components. A multi-scale convolutional module first scans the raw sensor sequences with parallel one-dimensional kernels of different sizes. Small kernels capture high-frequency pulses caused by instantaneous ventilation events, while larger kernels track slow trends dominated by the diurnal cycle. A neighborhood enhancement step at the input stage lets each sensor&#8217;s signal incorporate information from its spatial neighbors, mitigating the context loss that comes from treating sensors as isolated points. The convolutional features are then passed to a bidirectional long short-term memory network, or BiLSTM, whose forward chain learns historical dynamics such as cooling after ventilation and whose backward chain incorporates future influences such as heat released by the north wall, addressing the thermal inertia and time lags inherent in greenhouse physics.</p>
<p>The most distinctive element is the sensor-level gated attention mechanism. Instead of averaging all measurement points equally, the model computes a dynamic weight for each sensor at each time step, based on both that sensor&#8217;s temporal feature vector and global meteorological context such as solar radiation and outdoor conditions. A Softplus mapping generates the weights, which are normalized and used to blend all sensor features into a single representation of the greenhouse&#8217;s instantaneous physical state. In effect, the network learns when and where to listen: during periods of strong environmental fluctuation, it can up-weight sensors in high-gradient boundary zones or humid canopy cores, and relax that focus when conditions are uniform. The authors also added physics-informed weak constraints to the training loss, penalizing implausible spatial jumps between adjacent sensors and unrealistic temperature inversions between vertical layers, so that predictions remain consistent with fluid-dynamical continuity rather than merely fitting the numbers.</p>
<p>The performance results are striking. On the held-out test set, CNN-BiLSTM-SA achieved a mean absolute error of 0.886 degrees Celsius, a root mean square error of 1.236 degrees, and a coefficient of determination of 0.92 for temperature prediction, along with a mean absolute error of 2.619 percent, a root mean square error of 3.735 percent, and an R-squared of 0.93 for humidity. Removing the attention mechanism degraded performance substantially: compared with an otherwise identical CNN-BiLSTM model, the full architecture cut temperature root mean square error by 26.6 percent and humidity error by 37.9 percent. The model also degraded gracefully with longer forecast horizons, with temperature error rising only from 1.236 to 1.439 degrees between 1-hour and 24-hour predictions, and humidity error from 3.735 to 4.548 percent, indicating stable day-ahead forecasting without sudden failure. Notably, an ablation with an overly strong physics constraint, at a regularization weight of one, caused humidity predictions to collapse, a cautionary demonstration that physical priors must be balanced against data fitting rather than imposed at full strength.</p>
<p>Beyond point predictions, the framework reconstructs continuous three-dimensional fields by combining inverse distance weighting interpolation with parabolic geometric masking. The resulting visualizations recovered physically meaningful structures: an east-west temperature gradient of up to 6.4 degrees aligned with solar radiation incidence and heat accumulation, a north-south contrast driven by differential light transmission through the covering materials, and a vertical stratification pattern with maximum differences of 7 degrees reflecting buoyancy-driven warm-air ascent. On the humidity side, the model correctly located concentrated high-humidity cores in the middle and lower canopy, where crop transpiration and restricted airflow trap water vapor, and identified drier zones near ventilation boundaries. Two-dimensional slice comparisons showed that adding the attention mechanism sharply reduced residual errors, which without it frequently exceeded 2 degrees and 6 percentage points near the southern film and in densely planted central regions. The authors acknowledge some remaining smoothing of local humidity extremes, a reminder that fields governed by tightly coupled transpiration, airflow, and vapor retention are harder to resolve than temperature alone.</p>
<p>The practical implications extend to sensor economics and the emerging concept of greenhouse digital twins. The full 27-sensor network served as a dense reference for model calibration, but the team tested reduced configurations: keeping only the nine lower-layer sensors increased temperature and humidity errors by 78.8 and 45.5 percent respectively, proving that vertical gradients cannot be captured from below the canopy alone. A optimized 24-sensor arrangement, however, raised errors by only 1.38 and 3.59 percent relative to the full network, suggesting that operational deployments can trim costs after site- and season-specific calibration. Looking forward, the authors propose integrating LiDAR-derived canopy structure and traits such as leaf area index, coupling the framework with radiation and carbon dioxide variables through physics-informed neural networks, and feeding the reconstructed fields into crop growth and disease-risk models. In such a system, ventilation and heating would no longer respond to an average number from a central sensor, but to the actual geography of risk inside the greenhouse, targeting the humid canopy cores where fungal pathogens take hold before any single-point alarm would ever sound.</p>
<p><strong>Subject of Research:</strong> Deep learning-based spatiotemporal prediction and 3D reconstruction of temperature and humidity fields in solar greenhouses</p>
<p><strong>Article Title:</strong> A multi-sensor deep learning framework for spatiotemporal prediction and three-dimensional reconstruction of temperature and humidity fields in solar greenhouses</p>
<p><strong>Article References:</strong> A multi-sensor deep learning framework for spatiotemporal prediction and three-dimensional reconstruction of temperature and humidity fields in solar greenhouses. (n.d.). <a href="https://doi.org/10.1016/j.atech.2026.102517" rel="noopener noreferrer">https://doi.org/10.1016/j.atech.2026.102517</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.atech.2026.102517" rel="noopener noreferrer">10.1016/j.atech.2026.102517</a></p>
<p><strong>Keywords:</strong> solar greenhouse, deep learning, microclimate, temperature prediction, humidity prediction, attention mechanism, BiLSTM, convolutional neural network, precision agriculture, sensor networks, digital twin, spatial heterogeneity</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">219366</post-id>	</item>
		<item>
		<title>Federated AI Predicts Sepsis in ICUs Without Sharing Patient Data</title>
		<link>https://scienmag.com/federated-ai-predicts-sepsis-in-icus-without-sharing-patient-data/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 18:15:35 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[AI-based sepsis prediction without patient data sharing]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[clinical decision support]]></category>
		<category><![CDATA[cross-institutional healthcare data collaboration]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for early sepsis detection]]></category>
		<category><![CDATA[electronic health records]]></category>
		<category><![CDATA[electronic health records for predictive analytics]]></category>
		<category><![CDATA[FedAvg]]></category>
		<category><![CDATA[federated learning]]></category>
		<category><![CDATA[federated learning in healthcare]]></category>
		<category><![CDATA[health informatics]]></category>
		<category><![CDATA[healthcare data privacy regulations and AI solutions]]></category>
		<category><![CDATA[ICU]]></category>
		<category><![CDATA[ICU patient monitoring with AI]]></category>
		<category><![CDATA[improving sepsis outcomes with federated AI]]></category>
		<category><![CDATA[innovative approaches to healthcare data governance]]></category>
		<category><![CDATA[machine learning models respecting patient privacy]]></category>
		<category><![CDATA[privacy-preserving AI]]></category>
		<category><![CDATA[privacy-preserving machine learning for ICU patients]]></category>
		<category><![CDATA[sepsis prediction]]></category>
		<category><![CDATA[SHAP explainability]]></category>
		<category><![CDATA[simulation of hospital data for AI training]]></category>
		<category><![CDATA[Transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217930</guid>

					<description><![CDATA[Researchers have built a federated deep learning system that predicts sepsis across simulated hospitals with 93.6 percent accuracy while keeping all patient data local.]]></description>
										<content:encoded><![CDATA[<p>Sepsis, the body&#8217;s runaway immune response to infection, remains one of the leading causes of death in intensive care units across the United States, killing patients not because clinicians lack treatments but because the window for intervention is brutally narrow. Every hour that passes without recognition of the syndrome measurably worsens survival odds, which is why hospitals have long sought computational tools that can flag deteriorating patients before the classic signs become unmistakable. Electronic health records hold the raw material for such early warnings: heart rates, laboratory values, medication records, and scores of other variables streaming in from every bedside. The obstacle has never been the data itself but the walls around it. Privacy regulations and institutional governance policies make it extraordinarily difficult for hospitals to pool patient records, which means most predictive models are trained on a single institution&#8217;s population and often fail to generalize elsewhere. A new study published in Discover Social Science and Health proposes a way around this impasse, using federated learning to train a powerful hybrid deep learning model across simulated hospitals without any patient data ever leaving its home institution.</p>
<p>The research team, led by Miad Islam of Saint Leo University together with collaborators from institutions in Bangladesh, the United Kingdom, and the United States, built a framework that combines three complementary neural network architectures into a single sepsis prediction engine. The first component is a one-dimensional convolutional neural network, or 1D-CNN, which excels at scanning sequences of clinical measurements and picking out local patterns, much as a radiologist might scan an image for telling features. Layered on top of that is a bidirectional long short-term memory network, or BiLSTM, a recurrent architecture that reads the patient&#8217;s clinical timeline in both directions, capturing how early vital-sign changes foreshadow later laboratory abnormalities and how recent developments reframe earlier ambiguity. The final piece is a Transformer attention block, the same family of architecture that powers modern language models, which allows the network to weigh the relative importance of different clinical variables and time points dynamically rather than treating every input as equally significant. Together, these components form a model capable of learning the subtle, multi-scale signatures that precede sepsis onset.</p>
<p>The privacy mechanism at the heart of the study is federated learning, a training paradigm in which the data never moves. Instead of shipping records to a central server, each participating hospital trains the model locally on its own patients. Only the resulting model parameters, essentially long lists of numerical weights, are transmitted to a coordinating server, which averages them using an algorithm known as Federated Averaging, or FedAvg. The updated global model is then sent back to each site for another round of local training. Over many such rounds, the shared model gradually absorbs the statistical diversity of all participating institutions without any single record, diagnosis, or lab value ever being exchanged. For the purposes of this study, the researchers simulated this arrangement by partitioning a MIMIC-IV-style ICU dataset across three virtual hospital clients, allowing them to evaluate how the federated approach behaves under realistic distributed conditions while working with data that posed no privacy risk.</p>
<p>The results are striking for a model trained under such constraints. The federated hybrid framework achieved an accuracy of 93.6 percent and an area under the receiver operating characteristic curve, or AUROC, of 0.959, a standard measure of a classifier&#8217;s ability to distinguish septic from non-septic patients across all possible thresholds. Precision came in at 0.871, meaning that when the model raised an alarm it was usually justified, and specificity reached 98.24 percent, indicating it rarely flagged healthy patients falsely. The F1-score, which balances precision against recall, was 0.759, a figure that reflects the inherent difficulty of detecting a condition that is rare relative to the ICU population. For comparison, the researchers also trained a centralized baseline model on all the data pooled together, and it reached an AUROC of 0.997. That gap is real, but the federated model&#8217;s performance remains firmly in the range clinicians would consider useful, and it was achieved without the data centralization that privacy law makes impractical.</p>
<p>Perhaps even more consequential for real-world deployment is the framework&#8217;s communication efficiency. In federated systems, the cost of repeatedly shipping model weights between hospitals and the coordinating server can become prohibitive, particularly for institutions with limited bandwidth. Across all federated training rounds in this study, the total communication cost was just 36.09 megabytes, a figure small enough to travel over ordinary hospital networks in seconds. That efficiency matters because it suggests the approach could scale to many more participating sites without the coordination overhead becoming a bottleneck. In distributed healthcare environments, where connectivity is often uneven and IT infrastructure varies widely between a major academic medical center and a community hospital, keeping the communication footprint light is not a luxury but a prerequisite.</p>
<p>A model that clinicians cannot understand is a model they will not trust, so the researchers turned to SHAP, or SHapley Additive exPlanations, a technique borrowed from cooperative game theory that assigns each input variable a quantified contribution to every individual prediction. The SHAP analysis identified the Sequential Organ Failure Assessment score, a composite measure of dysfunction across six organ systems, as the most influential predictor of sepsis risk, followed by lactate level, white blood cell count, creatinine, and procalcitonin. Each of these is a familiar marker to any intensivist: lactate rises when tissues are starved of oxygen, white blood cells surge or crash during systemic infection, creatinine signals kidney injury, and procalcitonin is a well-established biomarker of bacterial sepsis. The fact that the model&#8217;s attention converged on variables that already carry clinical weight is reassuring, because it suggests the network learned genuine physiology rather than exploiting statistical artifacts in the data.</p>
<p>The authors are candid about the limits of their work, and their honesty is itself instructive. The study used a MIMIC-IV-style dataset rather than live records from multiple hospitals, and the federated setup was a simulation rather than a deployment across real institutions with genuinely heterogeneous patient populations, equipment, and documentation practices. The researchers also note that they did not formally define an adversary or threat model. They did not assess whether a malicious server or a colluding client could infer patient-level information from the exchanged model updates, a known vulnerability in federated systems where gradient information can sometimes leak details about training data. Techniques such as differential privacy, secure aggregation, and homomorphic encryption exist to close these gaps, and the authors flag rigorous validation on real multi-institutional ICU records as a necessary next step before any claim of scalable, secure deployment can be made.</p>
<p>Those caveats notwithstanding, the study lands at a moment when the tension between data hunger and data privacy has become the defining challenge of clinical artificial intelligence. The most accurate models are typically those trained on the largest and most diverse datasets, yet the most diverse datasets are precisely the ones that privacy law keeps fragmented across institutions. Federated learning offers a mathematically principled compromise, and pairing it with architectures that capture both local temporal patterns and long-range dependencies, as this hybrid CNN-BiLSTM-Transformer design does, may prove to be a template for predictive medicine well beyond sepsis. Acute kidney injury, respiratory failure, and cardiac deterioration are all conditions where early, distributed, privacy-preserving prediction could change outcomes.</p>
<p>What makes this work compelling as a piece of the larger puzzle is its demonstration that privacy and performance need not be framed as a zero-sum trade. A model trained without ever seeing another hospital&#8217;s records came within a few percentage points of the centralized ideal, and it did so with a communication budget measured in tens of megabytes and with explanations that map onto established clinical intuition. The road from a three-client simulation to a network of hospitals sharing a living model is long, and it will require threat modeling, regulatory negotiation, and prospective clinical validation. But the direction of travel is clear. If the next generation of ICU decision support can be trained collectively while respecting the sovereignty of every patient record, the hours that matter most in sepsis may finally be spent on treatment rather than detection.</p>
<p><strong>Subject of Research:</strong> Privacy-preserving federated deep learning for sepsis prediction in intensive care units</p>
<p><strong>Article Title:</strong> A hybrid deep learning framework for privacy-preserving sepsis prediction in distributed ICU environments using federated learning simulation</p>
<p><strong>Article References:</strong> Islam, M., Mohiuddin, T., Rahman, M. A., Sharfuddin, M., Islam, M. S., Sunny, S. R., &amp; Lokesh, E. (2026). A hybrid deep learning framework for privacy-preserving sepsis prediction in distributed ICU environments using federated learning simulation. <em>Discover Social Science and Health</em>. <a href="https://doi.org/10.1007/s44155-026-00489-1" rel="noopener noreferrer">https://doi.org/10.1007/s44155-026-00489-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44155-026-00489-1" rel="noopener noreferrer">10.1007/s44155-026-00489-1</a></p>
<p><strong>Keywords:</strong> federated learning, sepsis prediction, deep learning, ICU, electronic health records, privacy-preserving AI, Transformer, BiLSTM, SHAP explainability, clinical decision support, health informatics, FedAvg</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217930</post-id>	</item>
		<item>
		<title>Spectrum-Guided AI Method Predicts System Faults Before They Strike</title>
		<link>https://scienmag.com/spectrum-guided-ai-method-predicts-system-faults-before-they-strike/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 21:08:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[autoencoder]]></category>
		<category><![CDATA[Bayesian optimization]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[complex systems reliability]]></category>
		<category><![CDATA[discrete Fourier transform]]></category>
		<category><![CDATA[early system fault detection]]></category>
		<category><![CDATA[fault prediction]]></category>
		<category><![CDATA[fault prediction in large computing systems]]></category>
		<category><![CDATA[fault prediction without labeled data]]></category>
		<category><![CDATA[Gaussian Mixture Model]]></category>
		<category><![CDATA[intelligent system fault forecasting]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[log analysis]]></category>
		<category><![CDATA[machine learning for fault detection]]></category>
		<category><![CDATA[predictive maintenance in data centers]]></category>
		<category><![CDATA[real-time system health monitoring]]></category>
		<category><![CDATA[self-supervised learning]]></category>
		<category><![CDATA[spectrum-guided AI modeling]]></category>
		<category><![CDATA[subtle system behavior analysis]]></category>
		<category><![CDATA[system telemetry analysis]]></category>
		<category><![CDATA[teacher-student learning]]></category>
		<category><![CDATA[time-series analysis]]></category>
		<category><![CDATA[Unsupervised anomaly detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216365</guid>

					<description><![CDATA[Researchers have developed a spectrum-guided, teacher–student AI framework that adapts its time windows to the rhythms of unlabeled system logs and achieves an average F1-score of 88.49 percent in predicting faults before they occur.]]></description>
										<content:encoded><![CDATA[<p>Every large computing system, from an electric grid control room to a cloud data center, hums with a hidden language of logs. These timestamped records of internal events carry the earliest whispers of trouble: a memory leak that grows by a few megabytes an hour, a disk whose response times drift upward, a network link that begins dropping packets in a subtle rhythm. Catching those whispers early enough to act is one of the hardest problems in modern operations, and a team of researchers in China now reports a new approach that could make fault prediction substantially more reliable. Writing in the journal Complex &amp; Intelligent Systems, Lingyan Que and colleagues describe a framework that reads the rhythms of system telemetry, learns what normal behavior looks like without any labeled failure examples, and then teaches a fast prediction model to spot trouble before it fully develops.</p>
<p>The core challenge the researchers set out to solve is deceptively simple to state but notoriously difficult in practice. Real systems rarely fail in ways that come neatly labeled; operators almost never have a large library of annotated anomalous samples to train on, because genuine faults are rare, diverse, and often undocumented. At the same time, the signals that precede a fault unfold across multiple time scales simultaneously. A slow degradation may take hours to manifest, while a sudden spike can develop in seconds. Most existing methods force a single fixed observation window onto this data, which means they either blur out fast dynamics or miss slow ones. The new framework, which the authors call a spectrum-guided asymmetric window approach, attacks both problems at once by letting the data itself decide how the system should look at time.</p>
<p>The first step is a piece of signal processing that has been around for two centuries but proves remarkably effective here: the discrete Fourier transform, or DFT. By converting sequences of log-derived metrics from the time domain into the frequency domain, the method identifies the dominant periodic characteristics of the system&#8217;s behavior. In other words, it asks which cycles, whether daily load patterns, hourly batch jobs, or minute-level oscillations, actually dominate the telemetry. This spectral fingerprint then feeds into a Bayesian optimization strategy, a technique that efficiently searches a parameter space by balancing exploration against exploitation and updating a probabilistic model of where good solutions lie. Instead of a human engineer guessing window sizes by trial and error, the optimization procedure adaptively determines the appropriate window lengths for the two models at the heart of the architecture.</p>
<p>Those two models form a teacher–student pair, an arrangement borrowed from a family of techniques known as knowledge distillation. The teacher is a slower, richer model whose job is to understand normal behavior deeply; the student is a leaner network that must learn to make the same judgments quickly and, crucially, over a different stretch of time. In this framework, the teacher combines an autoencoder, a neural network trained to compress and then reconstruct its inputs, with a Gaussian mixture model, a statistical tool that models data as a weighted sum of several bell-shaped distributions. When the autoencoder reconstructs a normal pattern well, the corresponding mixture density is high; when it encounters something unfamiliar, reconstruction degrades and the density drops. The result is a continuous anomaly intensity, a soft label that expresses not just whether something looks wrong but how wrong it looks, graded from zero to full alarm.</p>
<p>That grading matters more than it might first appear. Binary labels, anomalous or not, throw away a great deal of information near the boundary of a developing fault, precisely where early prediction is most valuable. A soft label lets the student model learn the gradual onset of trouble, the way a vibration signature intensifies before a bearing fails or a queue length creeps up before a service collapses. The student, for its part, is built on a bidirectional long short-term memory network, or BiLSTM, a recurrent architecture that reads sequences in both forward and reverse directions and uses gated memory cells to preserve information over long spans. This gives it the capacity to capture long-range temporal dependencies, the kind of extended context that separates a genuine precursor from momentary noise.</p>
<p>But a teacher and a student looking through windows of different sizes cannot simply exchange knowledge, because their views of the data do not line up point for point. The researchers address this with what they call a temporal boundary alignment mechanism, a design element that maps representations across the mismatched time scales so that distillation remains meaningful. Conceptually, it ensures that what the teacher has learned about a stretch of system behavior can be transferred to the student even though the student is watching a longer or shorter slice of the same timeline. This cross-scale transfer is what allows the framework to be asymmetric in a productive way: the teacher can specialize in fine-grained, short-horizon anomaly scoring while the student generalizes over longer horizons suited to prediction rather than mere detection.</p>
<p>To test the approach, the team ran extensive experiments on four widely used benchmark datasets drawn from the anomaly detection literature: SMD, a server machine dataset collected from a large internet company; PSM, a pool server dataset capturing internal metrics; and SMAP and MSL, two datasets derived from telemetry of NASA Mars-orbiting and planetary rovers and lander systems. These benchmarks are standard proving grounds because they contain real, unlabeled time series with genuine anomalies and diverse periodic structures. Across all four, the proposed method achieved an average F1-score of 88.49 percent, a combined measure of precision and recall that penalizes both false alarms and missed detections. The authors report that this consistently outperformed several state-of-the-art approaches, suggesting that the combination of spectral guidance, adaptive windows, and soft-label distillation delivers a genuine advance rather than an incremental tweak.</p>
<p>The practical implications extend well beyond academic benchmarks. The research team is affiliated with State Grid Zhejiang Electric Power and associated grid automation companies in Hangzhou, which hints at the intended operating environment: power grid dispatching systems, where a failure to anticipate a fault can cascade into outages affecting millions of people. In such settings, unlabeled log data is abundant while confirmed fault labels are scarce, exactly the regime the new framework is built for. Because the method identifies periodic structure automatically and tunes its own windows through Bayesian optimization, it could in principle be redeployed to a new system without a lengthy manual re-engineering cycle, a property that matters enormously when operators must monitor heterogeneous fleets of servers, controllers, and communication links.</p>
<p>There are, of course, caveats worth keeping in view. The reported results come from retrospective benchmark datasets, and translating laboratory performance into live operational reliability typically demands further validation against drift, adversarial conditions, and the messy realities of production telemetry. The framework also inherits the computational costs of its components: Fourier analysis, Bayesian optimization, and bidirectional recurrent networks are all more demanding than simpler baselines, although the teacher–student design means the expensive teacher is needed only during training, while the deployed student remains comparatively lightweight. The study, which was conducted without external funding and is published open access, was received in May 2026, accepted in late August, and published on 26 September 2026, placing it among the most recent entries in a fast-moving field.</p>
<p>Even so, the conceptual contribution is likely to outlast any single benchmark score. By treating window length not as a nuisance parameter but as a quantity to be derived from the spectral structure of the data, and by replacing brittle binary labels with graded anomaly intensities distilled across time scales, the work reframes what fault prediction systems can learn from the logs they are given. As critical infrastructure grows more software-defined and more interdependent, the ability to read the rhythms of a system and act on their dissonance before failure arrives is shifting from a luxury to a necessity. Methods like this one, which let machines discover those rhythms on their own, point toward operations centers where the earliest whispers of trouble are not just heard but understood in time to answer.</p>
<p><strong>Subject of Research:</strong> Machine learning methods for time series fault prediction from unlabeled system logs</p>
<p><strong>Article Title:</strong> A spectrum-guided asymmetric window-based fault prediction method</p>
<p><strong>Article References:</strong> Que, L., Qian, J., Sun, Z., &amp; Huang, Y. (2026). A spectrum-guided asymmetric window-based fault prediction method. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02509-8" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02509-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02509-8" rel="noopener noreferrer">10.1007/s40747-026-02509-8</a></p>
<p><strong>Keywords:</strong> fault prediction, anomaly detection, time series analysis, knowledge distillation, teacher-student learning, Bayesian optimization, discrete Fourier transform, BiLSTM, autoencoder, Gaussian mixture model, self-supervised learning, log analysis</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216365</post-id>	</item>
		<item>
		<title>Hybrid AI Model Tames the Seasons to Predict Wind Power More Accurately</title>
		<link>https://scienmag.com/hybrid-ai-model-tames-the-seasons-to-predict-wind-power-more-accurately/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:19:27 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced wind speed prediction techniques]]></category>
		<category><![CDATA[atmospheric condition impact on wind data]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[computational efficiency in AI models]]></category>
		<category><![CDATA[computational intelligence]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning architecture for energy forecasting]]></category>
		<category><![CDATA[grid management and wind power]]></category>
		<category><![CDATA[grid stability]]></category>
		<category><![CDATA[hybrid deep learning models]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[machine learning for wind energy]]></category>
		<category><![CDATA[Renewable Energy]]></category>
		<category><![CDATA[renewable energy prediction]]></category>
		<category><![CDATA[seasonal modeling in renewable energy]]></category>
		<category><![CDATA[seasonal variability]]></category>
		<category><![CDATA[seasonal wind speed variation]]></category>
		<category><![CDATA[Sotavento wind farm]]></category>
		<category><![CDATA[time-series forecasting]]></category>
		<category><![CDATA[transformer models]]></category>
		<category><![CDATA[transformer vs recurrent neural networks]]></category>
		<category><![CDATA[wind power forecasting]]></category>
		<category><![CDATA[Yalova wind farm]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213163</guid>

					<description><![CDATA[Researchers have built a hybrid LSTM-BiLSTM deep learning model that adapts to seasonal wind patterns and outperforms transformers while using a fraction of the computing power.]]></description>
										<content:encoded><![CDATA[<p>Wind power is one of the cleanest and fastest-growing sources of electricity on the planet, but it carries an awkward secret: the wind does not behave the same way all year round. Monsoon and winter seasons tend to deliver strong, steady, and comparatively predictable winds, while summer and post-monsoon months bring weaker, more intermittent flows that leave turbines idling and grid operators guessing. A new study published in the journal Results in Engineering argues that this seasonal split is precisely where most forecasting models fall down, and it proposes a hybrid deep learning architecture designed specifically to handle the problem. The result is a model that outperforms not only standard recurrent networks but also fashionable transformer-based approaches, while using a fraction of the computational resources.</p>
<p>The research team, led by Manisha Galphade and colleagues, started from a deceptively simple observation: the statistical character of wind speed data changes with the seasons. The mean, variance, distribution shape, autocorrelation, and extreme values of wind records all shift as atmospheric conditions move from one regime to another. A single general-purpose model, trained on a full year of data, tends to overfit the turbulent high-wind seasons and underfit the calmer ones. Classical statistical tools such as ARIMA and seasonal ARIMA assume linearity and stationarity, assumptions that wind power data routinely violates. Machine learning methods like random forests and support vector machines capture nonlinear relationships but struggle with long-term temporal dependencies. Even the deep learning tools that have transformed other fields face difficulties when the underlying signal changes its personality every few months.</p>
<p>Transformers, the architecture behind the current artificial intelligence boom, have been proposed as a solution. Models such as the Informer, the Temporal Fusion Transformer, and the Autoformer bring powerful attention mechanisms to time-series forecasting, and they have shown impressive results on long sequences. But the authors point out a practical catch: wind farms rarely produce the enormous, clean datasets that transformers crave. Sensor coverage is limited, measurements are noisy, and self-attention scales quadratically with sequence length, driving up training time and hardware demands. With many hyperparameters to tune and a tendency to overfit small, messy datasets, transformers can be an expensive and fragile choice for operational wind forecasting, particularly at installations that lack cloud-scale computing.</p>
<p>The team&#8217;s answer is a serial hybrid that chains two recurrent architectures together. A long short-term memory network, or LSTM, first processes the raw wind power sequence. LSTMs are built around memory cells governed by three gates—an input gate, a forget gate, and an output gate—that control what information is stored, discarded, and passed forward, allowing the network to learn long-range dependencies without suffering from vanishing gradients. The output of this LSTM is then fed into a bidirectional LSTM, or BiLSTM, which reads the refined representations in both forward and backward directions. Because the BiLSTM operates on already-filtered, higher-level features rather than noisy raw signals, the combination performs a kind of progressive feature abstraction: the LSTM extracts stable sequential patterns, and the BiLSTM layers contextual understanding on top of them.</p>
<p>Before any training begins, the raw data undergoes careful preprocessing. Missing values are filled using imputation methods, with a random forest regression imputer chosen for one dataset because it captures nonlinear relationships without assuming a particular data distribution. Outliers are detected by standardizing each data point and flagging values that deviate too far from the mean, and the cleaned features are then rescaled to the range between zero and one using min-max normalization. These steps matter more than they might appear: wind farm records are riddled with sensor dropouts, negative power readings caused by measurement artifacts, and spikes from unusual atmospheric events, and a forecasting model fed uncleaned data will happily learn the noise.</p>
<p>The researchers tested their approach on two real wind farms with very different characters. The first is the Yalova wind farm in western Turkey, a 54,000-kilowatt installation of 36 turbines whose supervisory control and data acquisition system recorded wind speed, direction, generated power, and theoretical power at ten-minute intervals throughout 2018, yielding more than 46,000 records. The second is the Sotavento wind farm in Galicia, Spain, with 24 onshore turbines and a capacity of 17,560 kilowatts, providing hourly meteorological and generation data for 2014. In both cases the data was split by season, with thirty-day windows drawn from winter, spring, summer, and autumn, and the models were evaluated using root mean squared error, mean absolute error, and the coefficient of determination.</p>
<p>The results reveal a striking seasonal fingerprint. At Yalova, spring and summer proved the easiest to forecast, with the hybrid model achieving coefficients of determination as high as 0.99 and its best summer performance at a lookback window of six time steps. Autumn, a transitional season mixing summer and winter behavior, produced moderate errors, while winter was hardest of all: volatility, sudden spikes and drops, and non-stationary behavior pushed the error metrics up and the explanatory power down to roughly 0.80. At Sotavento the same pattern emerged, with the optimal lookback window shifting from five steps in spring to three in summer and just one in autumn. The authors identify this systematic analysis of lookback window optimization as a central contribution, showing that no single window length serves all seasons and that adaptive, season-specific temporal context can substantially improve accuracy.</p>
<p>The hybrid model did not just beat its individual components. Against standalone LSTM and BiLSTM baselines, it reduced mean absolute error by 1.38 percent in spring, 12.1 percent in summer, about 4.67 percent in autumn, and 5.94 percent in winter. It also outperformed temporal convolutional networks, attention-based LSTMs, and transformer models across both datasets, and a Diebold-Mariano statistical test confirmed that most of these improvements were significant at the five percent level, with the gaps against TCN and transformer models highly significant. An ablation study reinforced the design choices: removing either component, reversing the layer order, changing the unit counts, dropping regularization, or altering the learning rate all degraded performance, sometimes dramatically, confirming that the specific architecture and its hyperparameters are genuinely well balanced rather than accidentally lucky.</p>
<p>Perhaps the most persuasive numbers concern efficiency. The proposed model contains just 10,913 parameters, occupies 171 kilobytes, and trained in about 43 seconds—faster than every competing model tested, including the much larger TCN with its 89,473 parameters. Its normalized accuracy of 96.34 percent topped the field, edging out the LSTM at 95.11 percent and the transformer at 92.03 percent. The authors attribute this to a favorable bias-variance trade-off: transformers and attention models carry representational capacity that moderate-sized wind datasets cannot exploit, so their extra parameters mostly buy overfitting risk, while the LSTM&#8217;s gating mechanism acts as an inherent noise filter that attention mechanisms sometimes lack. With linear computational complexity in sequence length rather than quadratic, the architecture is well suited to real-time forecasting systems and edge deployments where memory and compute are scarce.</p>
<p>The implications reach well beyond two Spanish and Turkish wind farms. Accurate seasonal forecasting underpins grid planning, reserve operation, and maintenance scheduling, and as wind penetration grows, the cost of forecast error grows with it. The study is candid about its limits: winter remains difficult for every model tested, and the authors suggest that incorporating external meteorological variables, attention mechanisms, and ensemble methods could push performance further. But the core message is a useful corrective to the prevailing enthusiasm for ever-bigger architectures. For seasonal wind power prediction on realistic, medium-sized datasets, a thoughtfully composed pair of recurrent networks—one reading time forward, one reading it both ways—can beat the giants while running on hardware that fits in a wind farm&#8217;s back pocket.</p>
<p><strong>Subject of Research:</strong> Seasonal wind power forecasting using a hybrid LSTM-BiLSTM deep learning model tested on two wind farms</p>
<p><strong>Article Title:</strong> Seasonal wind power forecasting using data-driven and computational intelligence techniques</p>
<p><strong>Article References:</strong> Galphade, M., Dande, A., More, N., Nikam, V., &amp; Hatkar, V. (2026). Seasonal wind power forecasting using data-driven and computational intelligence techniques. <em>Results in Engineering, 32</em>, Article 113122. <a href="https://doi.org/10.1016/j.rineng.2026.113122" rel="noopener noreferrer">https://doi.org/10.1016/j.rineng.2026.113122</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.rineng.2026.113122" rel="noopener noreferrer">10.1016/j.rineng.2026.113122</a></p>
<p><strong>Keywords:</strong> wind power forecasting, LSTM, BiLSTM, deep learning, seasonal variability, renewable energy, time series forecasting, transformer models, grid stability, Yalova wind farm, Sotavento wind farm, computational intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213163</post-id>	</item>
		<item>
		<title>Entropy and AI join forces to catch costly serverless wallet attacks</title>
		<link>https://scienmag.com/entropy-and-ai-join-forces-to-catch-costly-serverless-wallet-attacks/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 08:57:48 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based cybersecurity]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[Azure Functions]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[cloud billing]]></category>
		<category><![CDATA[cloud billing exploitation]]></category>
		<category><![CDATA[cloud function abuse]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Denial of Wallet]]></category>
		<category><![CDATA[Denial of Wallet attacks]]></category>
		<category><![CDATA[entropy analysis in cybersecurity]]></category>
		<category><![CDATA[FaaS]]></category>
		<category><![CDATA[GAN data augmentation]]></category>
		<category><![CDATA[hybrid detection models]]></category>
		<category><![CDATA[innovative cybersecurity research]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[preventing costly serverless attacks]]></category>
		<category><![CDATA[scalable attack detection]]></category>
		<category><![CDATA[serverless application security]]></category>
		<category><![CDATA[serverless computing]]></category>
		<category><![CDATA[Serverless computing vulnerabilities]]></category>
		<category><![CDATA[serverless scalability risks]]></category>
		<category><![CDATA[Shannon entropy]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212290</guid>

					<description><![CDATA[Researchers at the University of Alicante have developed a two-stage detection system combining Shannon entropy analysis with deep learning to catch Denial of Wallet attacks that exploit serverless pay-per-use billing.]]></description>
										<content:encoded><![CDATA[<p>Serverless computing has become one of the most popular ways to build modern applications. Instead of renting servers around the clock, developers upload individual functions that a cloud provider instantiates only when an event triggers them, and bills only for the milliseconds of execution consumed. This Function-as-a-Service model eliminates infrastructure management and promises near-infinite scalability, but it also creates an entirely new kind of vulnerability. Researchers at the University of Alicante in Spain have now unveiled a hybrid detection model designed to catch a threat that traditional network defenses are poorly equipped to see: the Denial of Wallet attack, in which adversaries do not try to crash a service but simply to bankrupt its owner.</p>
<p>The study, published open access in the journal Cybersecurity by José Manuel Ortega Candel, Francisco José Mora Gimeno and Higinio Mora Mora, makes a sharp distinction between two attack classes that are often conflated. Distributed Denial of Service attacks aim to exhaust computational resources and render a system unavailable to legitimate users. Denial of Wallet attacks, by contrast, exploit the pay-per-use billing model itself. An attacker floods a serverless application with invocation requests, forcing the platform to auto-scale and execute thousands of billable function instances. The service may remain perfectly accessible throughout the assault; the damage appears instead on the monthly invoice, in the form of invocation charges and compute time measured in gigabyte-seconds. The authors argue that treating this as an availability problem, as conventional DDoS detection does, misses the point: the harm is economic, and every undetected attack invocation carries a direct, quantifiable monetary cost.</p>
<p>Because the monetary loss is only the visible consequence of a technical exploitation, the team&#8217;s primary objective is not to measure financial damage after the fact but to intercept the abnormal invocation behavior that triggers it. Their threat model assumes an attacker with black-box knowledge of publicly accessible HTTP trigger endpoints, no access to internal monitoring or billing systems, and full control over invocation rates and payloads. Crucially, the detection system operates only on execution metadata that platforms such as AWS CloudWatch and Azure Monitor already expose in real time: function name, trigger type, invocation timestamp and execution duration. No request content inspection or deep packet analysis is required, which keeps the approach lightweight and privacy-preserving.</p>
<p>The heart of the proposed architecture is a two-stage pipeline. The first stage computes Shannon entropy, a classical information-theoretic measure of randomness, over serverless execution features within sliding time windows. Rather than measuring entropy over network-layer features like IP addresses or ports, as DDoS detectors do, the model calculates it over the distribution of function invocation types, meaning pairs of function names and trigger categories, together with execution-duration buckets. This choice is deliberate: each invocation type carries a distinct cost profile, and an attacker seeking to maximize billing will inevitably alter this distribution, either by flooding a single expensive function or by spreading calls across many functions to exploit auto-scaling. A deviation in the entropy of this distribution therefore directly signals a change in billing-relevant behavior, even when the underlying traffic looks entirely legitimate at the network layer.</p>
<p>The entropy stage acts as a rapid triage filter with three tiers. Traffic whose entropy falls below a lower threshold is classified as clearly benign and never reaches the computationally expensive AI models. Values above an upper threshold are flagged immediately as suspicious with near-zero false negatives. Ambiguous transactions in the intermediate range are forwarded to the second stage, where supervised machine learning and deep learning classifiers perform precise analysis. Because a missed attack costs money while a false alarm does not, the threshold calibration is explicitly optimized to minimize the false negative rate, an inversion of the availability-focused logic of conventional DDoS systems. Dynamic thresholds computed from the mean and standard deviation of a rolling baseline of the five most recent windows allow the system to track gradual workload shifts such as daily traffic cycles without mislabeling them as attacks. A sensitivity factor of 2.5 was selected empirically, covering roughly 98.76 percent of normal traffic under a Gaussian assumption, and a persistence counter requires anomalies to survive across multiple consecutive windows before a formal alert is raised.</p>
<p>To train and evaluate the second-stage classifiers, the researchers had to overcome a fundamental obstacle: no publicly available labeled datasets for serverless Denial of Wallet attacks existed. They built one by combining a purpose-built DoW traffic simulator with fourteen days of real telemetry from the Microsoft Azure Functions dataset, incorporating call time series, execution durations and memory usage, and then augmenting the result with synthetic transactions generated by Generative Adversarial Networks. The final dataset contains 187,087 transactions. All experiments used Python on Google Colab with GPU and TPU acceleration, and the code and models have been released for independent verification.</p>
<p>The evaluation reveals a nuanced picture of when the hybrid architecture pays off. Under the optimal first-stage configuration, a 480-minute window with a low threshold, the entropy filter captured 83.08 percent of transactions as suspicious with a false negative rate of just 0.44 percent. Among the second-stage models, the bidirectional long short-term memory network, or BiLSTM, emerged as the recommended classifier, achieving a precision of 0.9810 and test accuracy of 0.9890, thanks to its ability to model invocation sequences in both forward and backward directions and capture temporal dependencies that static classifiers miss. For resource-constrained deployments, the simpler GRU architecture offers a more efficient alternative.</p>
<p>The most striking results concern computational efficiency under realistic attack densities. Because the original dataset contained an unrealistically high proportion of malicious traffic, roughly 70 percent, the entropy stage could filter out at most the legitimate minority. So the researchers expanded the dataset to 200,000 transactions and progressively reduced the attack density from 60 percent down to 1 percent, mirroring production conditions where attacks are rare. As the attack ratio fell, the efficiency gain of the entropy pre-filter soared. At a 1 percent attack transaction rate, the LSTM network achieved a processing gain of 91.95 percent, BiLSTM reached 91.11 percent, and the K-Neighbors classifier led traditional machine learning methods at 80.05 percent, all while maintaining F1-scores above 0.97 and ROC-AUC above 0.98 for the deep learning models. In other words, the hybrid design delivers its greatest savings precisely in the low-attack-density conditions that real serverless deployments actually experience, cutting the compute burden of security monitoring by more than nine tenths without degrading detection quality.</p>
<p>The authors are candid about the limits of their work. The evaluation was conducted offline on batch data, and the synthetic dataset, however carefully constructed, cannot fully replicate cold-start latency distributions, provider-specific concurrency throttling, billing rounding granularities and the workload heterogeneity of live production environments. They also note that under their specific test configuration, where attack traffic dominated, the two-stage pipeline actually imposed more total overhead than running the AI classifier alone; the entropy stage functioned primarily as an early-warning triage rather than a bulk traffic reducer in that regime. Future work will focus on large-scale deployment on OpenFaaS, an open-source platform that avoids commercial billing costs during attack experiments, and integration with cloud cost management APIs such as AWS Cost Explorer and Azure Cost Management, enabling alerts the moment projected spending exceeds user-defined thresholds.</p>
<p>For an industry betting its economics on serverless computing, the study arrives at an opportune moment. Major cloud providers handle ever-growing transaction volumes, and the same auto-scaling that makes functions resilient to downtime makes them exquisitely vulnerable to financial exhaustion. By re-deriving a decades-old information-theoretic concept for the peculiar anatomy of function invocations, and pairing it with deep learning classifiers that remember how workloads unfold over time, the Alicante team has produced the first hybrid methodology built specifically around the economics of Denial of Wallet attacks rather than adapted from DDoS playbooks. The message for cloud architects is clear: in the serverless era, the firewall that matters most may be the one guarding your budget.</p>
<p><strong>Subject of Research:</strong> Hybrid entropy and AI-based detection of Denial of Wallet attacks in serverless computing</p>
<p><strong>Article Title:</strong> Hybrid model for detecting Denial of Wallet attacks in serverless architectures</p>
<p><strong>Article References:</strong> Candel, J. M. O., Gimeno, F. J. M., &amp; Mora Mora, H. (2026). Hybrid model for detecting Denial of Wallet attacks in serverless architectures. <em>Cybersecurity, 9</em>(1), Article 221. <a href="https://doi.org/10.1186/s42400-026-00663-7" rel="noopener noreferrer">https://doi.org/10.1186/s42400-026-00663-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s42400-026-00663-7" rel="noopener noreferrer">10.1186/s42400-026-00663-7</a></p>
<p><strong>Keywords:</strong> serverless computing, Denial of Wallet, cybersecurity, Shannon entropy, deep learning, BiLSTM, FaaS, cloud billing, anomaly detection, machine learning, Azure Functions, GAN data augmentation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212290</post-id>	</item>
		<item>
		<title>Hybrid AI Detector Spots Machine-Written Text With Near-Perfect Accuracy</title>
		<link>https://scienmag.com/hybrid-ai-detector-spots-machine-written-text-with-near-perfect-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:23:26 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[1D CNN]]></category>
		<category><![CDATA[academic integrity]]></category>
		<category><![CDATA[advanced AI text discrimination techniques]]></category>
		<category><![CDATA[AI-generated text detection]]></category>
		<category><![CDATA[applications of hybrid AI detectors in education and journalism]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[challenges of differentiating human vs. AI-generated content]]></category>
		<category><![CDATA[combating academic dishonesty with AI detectors]]></category>
		<category><![CDATA[Complex & Intelligent Systems]]></category>
		<category><![CDATA[DAIGT dataset]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[HC3 dataset]]></category>
		<category><![CDATA[high-accuracy machine-written content classifiers]]></category>
		<category><![CDATA[hybrid deep learning models for AI text identification]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[misinformation]]></category>
		<category><![CDATA[multi-architecture neural network frameworks for detecting machine writing]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[near-perfect accuracy in AI text classification]]></category>
		<category><![CDATA[neural network fusion for AI text detection]]></category>
		<category><![CDATA[recent advancements in AI-generated content detection]]></category>
		<category><![CDATA[reliability of neural network-based AI text detectors]]></category>
		<category><![CDATA[Transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204556</guid>

					<description><![CDATA[Researchers have built a hybrid deep learning model combining BiLSTM, Transformer and 1D CNN components that detects AI-generated text with up to 99.47 percent accuracy.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence can now write essays, news stories, product reviews and exam answers that are often indistinguishable from human prose, and that capability has created an urgent problem: how do you tell machine-generated text apart from the real thing? A team of researchers from Zhengzhou University in China, the University of Okara in Pakistan and Taiz University in Yemen believes it has found a substantially better answer. In a study published in the journal Complex &amp; Intelligent Systems, the researchers describe a hybrid deep learning framework that fuses three complementary neural architectures into a single detector, achieving test accuracies of 99.47 percent on one benchmark dataset and 97.09 percent on another, results that place the model among the most reliable AI-text detectors reported to date.</p>
<p>The work was led by Muhammad Sohail, Zan Hongying and Muhammad Abdullah of Zhengzhou University&#8217;s School of Computer Science and Artificial Intelligence, together with Niu Guiling, Javed Rashid, Ghulam Ali, Muhammad Irfan and AbdulGuddoos S. A. Gaid, all of whom contributed equally to the research. Their motivation is straightforward. As large language models have grown more fluent, the risks they pose have grown with them. Academic dishonesty, in which students submit machine-written assignments as their own work, and the industrial-scale spread of misinformation on social media are the two threats the authors single out as most pressing. Detection tools built on a single type of neural network, they argue, tend to miss the subtle statistical fingerprints that separate synthetic prose from human writing, so the team set out to combine several kinds of pattern recognition into one system.</p>
<p>The architecture at the heart of the study weaves together three distinct deep learning components, each of which reads text in a different way. The first is a Bidirectional Long Short-Term Memory network, or BiLSTM, a recurrent architecture that processes a sequence of words in both forward and reverse order. Because it reads in two directions, the BiLSTM can capture context that unfolds across a sentence, learning how the meaning of a word is shaped by what comes before and after it, and it is particularly good at remembering long-range dependencies that simpler models lose track of. Long Short-Term Memory networks were designed specifically to solve the vanishing gradient problem that plagued earlier recurrent networks, allowing them to retain information over many time steps.</p>
<p>The second component is a set of Transformer blocks, the same fundamental technology that powers modern large language models. Transformers rely on a mechanism called self-attention, which lets the model weigh the relevance of every word in a passage against every other word, regardless of distance. Where a recurrent network moves through a sentence one token at a time, a Transformer can attend globally, picking up on structural regularities such as unusually uniform sentence rhythm, repetitive phrasing or the statistically smooth word distributions that language models tend to produce. The irony is deliberate and effective: the very architecture that makes AI text generation possible is here repurposed to detect its output, because the attention layers can highlight the telltale patterns that generative models leave behind.</p>
<p>The third component is a one-dimensional Convolutional Neural Network, or 1D CNN. Convolutional networks slide small filters across the input, and when the input is a sequence of word embeddings, those filters act as local pattern detectors, much like the edge detectors in image-recognition systems. In text, they excel at capturing n-gram-like features, short contiguous sequences of characters or words that recur in machine-generated passages. By stacking convolutional layers with pooling operations, the network builds up from local lexical cues to broader stylistic signatures. The researchers&#8217; insight is that these three views of a document, the sequential memory of the BiLSTM, the global attention of the Transformer and the local pattern sensitivity of the CNN, are complementary, and that a framework which integrates them should outperform any one of them alone.</p>
<p>To train and evaluate the hybrid model, the team used two diverse datasets, DAIGT and HC3, both of which contain thousands of text samples. The corpora pair human-written passages with machine-generated passages produced by a variety of large language models, which is an important design choice. A detector trained only on the output of a single model risks becoming a specialist that fails the moment a different generator appears. By drawing on multiple sources of synthetic text, the datasets force the model to learn general distinguishing features of AI prose rather than the quirks of one particular system. HC3, in particular, has become a widely used benchmark for this task because it pairs ChatGPT-style responses with human answers drawn from question-answering communities, while DAIGT offers a broader mix of generated content for training and testing.</p>
<p>The results were striking. On the DAIGT dataset, the hybrid framework reached a test accuracy of 99.47 percent, meaning it misclassified fewer than six in a thousand documents. On the HC3 dataset, it achieved 97.09 percent, still a level of performance that would leave only a small fraction of texts incorrectly labeled. The authors attribute this performance to the model&#8217;s ability to capture subtle linguistic and stylistic differences between AI-generated and human-written content, differences that are often invisible to human readers but statistically robust. Human writing tends to carry irregularities in rhythm, vocabulary choice and sentence construction, whereas machine-generated text, even when polished, exhibits measurable regularities that the combined networks can learn to recognize.</p>
<p>The implications extend well beyond the laboratory. In education, institutions struggling to uphold academic integrity in the era of freely available chatbots could integrate detectors of this kind into submission workflows, flagging assignments that show a high probability of machine authorship for closer review. In journalism and on social media platforms, where coordinated campaigns of AI-written posts can flood feeds with synthetic opinions, a reliable detector offers a tool for triage at scale. The authors explicitly frame their contribution as supporting applications aimed at maintaining content authenticity and academic integrity, and the near-perfect accuracy figures suggest the approach could withstand the noisy, adversarial conditions of real-world deployment better than single-architecture baselines.</p>
<p>Still, the researchers and independent observers alike caution that this is a moving target. Each new generation of language models produces text that is smoother and harder to distinguish, and detectors must evolve in step. The hybrid design has an advantage here: because it learns from data rather than from hand-crafted rules, it can be retrained as new generators emerge, and its multi-component structure means that even if one architecture&#8217;s advantage fades as models improve, the others may still carry signal. The study was supported by the Key Program of the Natural Science Foundation of China under grant U23A20316 and by the Project of Humanities and Social Sciences of the Ministry of Education under grant 20YJA740033, and the article is published open access, making the full technical details available to any research group that wants to build on it.</p>
<p>What the study ultimately demonstrates is a principle that may define the next phase of the AI era: the same deep learning revolution that created the problem of synthetic text is also supplying the tools to police it. By combining recurrent memory, self-attention and convolutional pattern detection in a single framework, the Zhengzhou-led team has shown that the boundary between human and machine writing, however blurred it appears to the naked eye, remains sharply visible to the right kind of algorithm. As generative models continue to spread through classrooms, newsrooms and social networks, detectors of this hybrid breed are likely to become as routine a part of the digital infrastructure as spam filters are today, quietly sorting authentic human expression from its synthetic imitations.</p>
<p><strong>Subject of Research:</strong> Development of a hybrid deep learning framework combining BiLSTM, Transformer and 1D CNN architectures to detect AI-generated text</p>
<p><strong>Article Title:</strong> Hybrid deep learning framework for AI-generated text detection</p>
<p><strong>Article References:</strong> Sohail, M., Hongying, Z., Guiling, N., Rashid, J., Abdullah, M., Ali, G., Irfan, M., &amp; Gaid, A. S. A. (2026). Hybrid deep learning framework for AI-generated text detection. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02501-2" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02501-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02501-2" rel="noopener noreferrer">10.1007/s40747-026-02501-2</a></p>
<p><strong>Keywords:</strong> AI-generated text detection, deep learning, BiLSTM, Transformer, 1D CNN, large language models, natural language processing, academic integrity, misinformation, DAIGT dataset, HC3 dataset, Complex &amp; Intelligent Systems</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204556</post-id>	</item>
		<item>
		<title>AI Model AVP-Pro Speeds Discovery of Antiviral Peptides</title>
		<link>https://scienmag.com/ai-model-avp-pro-speeds-discovery-of-antiviral-peptides/</link>
		
		<dc:creator><![CDATA[Kristina Jarvis]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:06:20 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI-driven antiviral drug development]]></category>
		<category><![CDATA[Antiviral peptide discovery]]></category>
		<category><![CDATA[antiviral peptides]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[bioinformatics tools for antiviral peptide discovery]]></category>
		<category><![CDATA[BLOSUM62]]></category>
		<category><![CDATA[computational drug design for viral inhibition]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for antiviral activity prediction]]></category>
		<category><![CDATA[ESM-2]]></category>
		<category><![CDATA[functional subtype prediction]]></category>
		<category><![CDATA[high-throughput antiviral peptide screening]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in antiviral peptide identification]]></category>
		<category><![CDATA[OHEM strategy]]></category>
		<category><![CDATA[peptide function prediction]]></category>
		<category><![CDATA[peptide-based antiviral therapeutics]]></category>
		<category><![CDATA[protein language models in peptide research]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[sequence chemistry of antiviral peptides]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[viral membrane interaction peptides]]></category>
		<category><![CDATA[virus-specific antiviral peptide prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202508</guid>

					<description><![CDATA[Researchers have developed AVP-Pro, a two-stage deep learning framework that identifies antiviral peptides and predicts which virus families and specific viruses they target.]]></description>
										<content:encoded><![CDATA[<p>Antiviral peptides have long been viewed as one of the more promising corners of the antiviral toolbox: short chains of amino acids that can interfere with viruses before they gain a foothold in a host cell. Yet identifying which peptides actually possess antiviral activity, and what kind of viral targets they prefer, has remained a slow and expensive experimental problem. A team of researchers in China now reports a computational framework, called AVP-Pro, that aims to compress that search, combining modern protein language models with a two-stage deep learning pipeline that not only flags candidate antiviral peptides but also predicts which virus families and specific viruses they are likely to act against. The work, published in BMC Genomics, was led by Xinru Wen, Weizhong Lin, Zi Liu and Xuan Xiao of the School of Information Engineering at Jingdezhen Ceramic University in Jiangxi Province.</p>
<p>The biological rationale behind the study rests on well-established sequence chemistry. Antiviral peptides tend to share recognizable characteristics: particular amino acid compositions, a net positive charge that helps them interact with negatively charged viral membranes or envelopes, hydrophobicity patterns that govern how they insert into lipid bilayers, and conserved sequence motifs that recur across peptides with similar mechanisms of action. But these features are distributed unevenly and interact in nonlinear ways, which is precisely why simple rule-based screens have struggled. Two peptides can look superficially similar yet differ sharply in antiviral potency, and a peptide that works against one virus family may be inert against another.</p>
<p>Existing computational classifiers, the authors note, have generally treated the problem as a binary one: is a given sequence an antiviral peptide or not? That framing discards valuable information. Because different AVPs exhibit distinct virus-targeting specificities, a binary answer tells a laboratory only half of what it needs to know before committing to synthesis and testing. The team also identified three persistent technical weaknesses in prior methods: difficulty modeling complex, long-range dependencies within peptide sequences; difficulty integrating heterogeneous feature sources, such as learned embeddings and hand-crafted physicochemical descriptors; and difficulty separating highly similar positive and negative samples that crowd the decision boundary between classes.</p>
<p>AVP-Pro addresses the feature-integration problem through what the authors call adaptive multi-representation fusion. The framework draws on two complementary sources of information. The first is a deep sequence representation derived from ESM-2, a large-scale protein language model trained on hundreds of millions of natural protein sequences. Such models learn to encode structural and functional context into dense numerical vectors, capturing subtle patterns that would be difficult to hand-engineer. The second source is a set of ten conventional physicochemical descriptors, the classical encodings of peptide science, covering properties such as charge, hydrophobicity and residue composition. By fusing these representations adaptively, rather than simply concatenating them, the model can weight each information source according to its usefulness for a given input.</p>
<p>Architecture plays a central role in how the framework reads a sequence. Convolutional neural network layers are used to capture local fragment-level features, short contiguous stretches of residues that often carry functional significance. Bidirectional long short-term memory units, or BiLSTM, then sweep across the sequence in both directions, capturing global contextual dependencies that span the full length of the peptide. Self-attention modules sit on top of this, allowing the model to focus dynamically on the positions most informative for the classification decision. An adaptive gating mechanism finally arbitrates among these parallel streams, deciding how much each representation should contribute to the fused output for any particular peptide rather than imposing a fixed weighting scheme across the entire dataset.</p>
<p>Perhaps the most distinctive element of the study is its handling of ambiguous samples. In AVP datasets, positive and negative sequences are often extremely similar, and this similarity produces fuzzy decision boundaries that degrade classifier performance, especially for the borderline cases that matter most in real screening pipelines. The team attacked this from two directions. First, they introduced data augmentation guided by BLOSUM62, the standard amino acid substitution matrix long used in sequence alignment. BLOSUM62 scores encode which residue substitutions are biologically tolerable, so augmenting training data with substitution-informed variants exposes the model to realistic sequence variation without straying into biologically implausible territory. Second, they employed online hard example mining, or OHEM, within a contrastive learning objective. Contrastive learning trains the model to pull similar examples together and push dissimilar ones apart in its internal embedding space; by focusing the loss on the hardest, most easily confused samples, the framework sharpens the boundary where it is thinnest.</p>
<p>The resulting system operates in two stages. In the first stage, AVP-Pro performs general antiviral peptide identification, deciding whether an input sequence is an AVP at all. On the independent test set, the authors report that the model achieved competitive predictive performance, holding its own against existing approaches while offering the richer internal machinery described above. The second stage is where the framework departs more sharply from earlier work. Building on the first-stage model through transfer learning, AVP-Pro predicts functional subtypes: which virus families and which specific viruses a candidate peptide is likely to target. In this stage the framework covered six virus families and eight specific viruses, a level of functional granularity that binary classifiers simply cannot provide.</p>
<p>Across multiple evaluation metrics, the authors report that AVP-Pro showed stable performance in both the general identification task and the functional subtype prediction task on the benchmark datasets they evaluated. Stability across metrics is an important qualifier in this field, since models optimized for a single figure of merit can exhibit lopsided behavior, excelling on accuracy while failing on measures that penalize false negatives or class imbalance. The consistency reported here suggests that the fusion and contrastive components are doing genuine work rather than merely inflating one headline number. The authors position the framework as a tool for sequence-level functional annotation and for prioritizing candidate antiviral peptides before laboratory characterization, a use case in which even modest gains in precision translate into meaningful savings of time and reagent cost.</p>
<p>The broader context makes work of this kind increasingly timely. Peptide-based antivirals occupy an attractive middle ground between small molecules and full-length protein therapeutics: they are typically less immunogenic than antibodies, more specific than broad-spectrum antiviral compounds, and synthetically accessible. But the design space of possible peptide sequences is astronomically large, and experimental screens cover only a vanishing fraction of it. Machine learning filters of the kind embodied in AVP-Pro act as a funnel, narrowing candidate lists so that wet-lab resources are spent on sequences with a computationally justified prior of activity. The addition of subtype prediction pushes the funnel further, allowing researchers to search not just for antiviral activity in general but for activity against a virus family of particular interest.</p>
<p>The study also illustrates a wider trend in computational biology: the convergence of large pretrained protein language models with task-specific deep learning heads and classical domain knowledge. ESM-2 embeddings supply the learned, context-rich representation; BLOSUM62 and physicochemical descriptors supply decades of accumulated biochemical insight; and the attention, recurrent and gating modules supply the flexibility to combine them on a per-instance basis. None of these ingredients is new on its own, but their integration, together with hard-example-focused contrastive training, reflects a maturing methodology for peptide function prediction. The work was supported by grants from the National Natural Science Foundation of China, and the authors declare no competing interests. As with any computational predictor, the model&#8217;s judgments will ultimately need experimental validation, but as a screening instrument it offers researchers a substantially finer-grained map of the antiviral peptide landscape than the binary tools that preceded it.</p>
<p><strong>Subject of Research:</strong> Machine learning-based identification and functional subtype prediction of antiviral peptides.</p>
<p><strong>Article Title:</strong> AVP-Pro: adaptive multi-representation fusion and contrastive learning for antiviral peptide identification and functional subtype prediction</p>
<p><strong>Article References:</strong> Wen, X., Lin, W., Liu, Z., &amp; Xiao, X. (2026). AVP-Pro: adaptive multi-representation fusion and contrastive learning for antiviral peptide identification and functional subtype prediction. <em>BMC Genomics</em>. <a href="https://doi.org/10.1186/s12864-026-13352-z" rel="noopener noreferrer">https://doi.org/10.1186/s12864-026-13352-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12864-026-13352-z" rel="noopener noreferrer">10.1186/s12864-026-13352-z</a></p>
<p><strong>Keywords:</strong> antiviral peptides, machine learning, ESM-2, contrastive learning, transfer learning, BLOSUM62, BiLSTM, self-attention, OHEM strategy, functional subtype prediction, peptide function prediction, deep learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202508</post-id>	</item>
		<item>
		<title>Hybrid AI Model Blends Transformer and BiLSTM to Predict Cancer Drug Synergy</title>
		<link>https://scienmag.com/hybrid-ai-model-blends-transformer-and-bilstm-to-predict-cancer-drug-synergy/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:15:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in Oncology]]></category>
		<category><![CDATA[benchmark dataset classification accuracy]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[BT-Synergy]]></category>
		<category><![CDATA[cancer cell lines]]></category>
		<category><![CDATA[Cancer drug synergy prediction]]></category>
		<category><![CDATA[combination cancer therapy]]></category>
		<category><![CDATA[combination therapy]]></category>
		<category><![CDATA[computational drug discovery]]></category>
		<category><![CDATA[computational pharmacology]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[drug interaction modeling]]></category>
		<category><![CDATA[drug pair screening automation]]></category>
		<category><![CDATA[drug synergy prediction]]></category>
		<category><![CDATA[DrugCombDB]]></category>
		<category><![CDATA[hybrid deep learning models]]></category>
		<category><![CDATA[in vitro synergy assay limitations]]></category>
		<category><![CDATA[molecular and cell-line data encoding]]></category>
		<category><![CDATA[protein embeddings]]></category>
		<category><![CDATA[ProteinBERT]]></category>
		<category><![CDATA[representational learning in pharmacology]]></category>
		<category><![CDATA[SELFIES]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[transformer and BiLSTM integration]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201632</guid>

					<description><![CDATA[Researchers at the University of Qom have developed BT-Synergy, a hybrid BiLSTM-Transformer deep learning model that predicts synergistic cancer drug combinations with 0.8458 accuracy by integrating SELFIES molecular encodings with protein language model cell-line representations.]]></description>
										<content:encoded><![CDATA[<p>One of the most stubborn bottlenecks in modern oncology is not finding new drugs, but finding the right pairs of existing drugs that work better together than either does alone. Combination therapy can amplify treatment efficacy, slow the emergence of resistance, and reduce systemic toxicity, yet the number of possible drug pairings across thousands of compounds and hundreds of cancer cell lines grows so quickly that laboratory screening cannot keep pace. In vitro synergy assays remain slow, expensive, and labor-intensive, leaving most of the chemical space of possible combinations unexplored. A new study published in Discover Artificial Intelligence by Sahar Abbasi Rostami and Amir Lakizadeh of the University of Qom in Iran addresses this gap with a hybrid deep learning architecture called BT-Synergy, which the researchers report achieved an accuracy of 0.8458 on benchmark datasets for classifying synergistic drug combinations.</p>
<p>The central problem that BT-Synergy tackles is representational. Earlier computational approaches to synergy prediction, including AuDNNsynergy, SynPathy, and the widely used DeepSynergy model, relied on engineered molecular descriptors or structured multi-omics inputs. More recent frameworks such as SynergyX, DFFNDDS, SYNPRED, PRODeepSyn, DeepTraSynergy, and CFSSynergy jointly encode drug structures and cell-line characteristics using attention mechanisms, feature fusion modules, or protein-protein interaction networks. Yet many of these methods depend on SMILES string encodings or protein similarity matrices, which can miss higher-order chemical and biological dependencies. The Qom team argues that what is needed is an encoder that simultaneously understands the sequential grammar of a molecule and the long-range contextual relationships distributed across it, joined with a biologically grounded picture of the cell in which the interaction takes place.</p>
<p>To build that encoder, the researchers turned to SELFIES, a self-referencing molecular string representation that guarantees chemically valid outputs, unlike SMILES, which can produce structurally impossible sequences that corrupt downstream learning. Each drug is tokenized into SELFIES subunits, truncated or padded to a fixed length, and then processed by a hybrid module in which a bidirectional long short-term memory network is integrated directly into a Transformer block. In the final configuration, the BiLSTM actually replaces the conventional feed-forward sublayer inside the Transformer encoder. This design choice is deliberate: the Transformer&#8217;s multi-head self-attention excels at capturing long-range, non-local dependencies across a molecular sequence, while the BiLSTM contributes sequential inductive biases, reading the token stream in both forward and backward directions to preserve local structural patterns that attention alone can dilute.</p>
<p>The architecture was not chosen blindly. The team systematically compared variants, including a GRU-Transformer hybrid, a parallel configuration in which BiLSTM and Transformer pathways process the sequence simultaneously before element-wise fusion and layer normalization, and pure BiLSTM or pure Transformer baselines. Model depth also mattered: reducing the Transformer to two layers slightly degraded accuracy, while four or five layers inflated computational cost without commensurate gains. Three layers emerged as the optimal balance and were adopted in the final model. Ablation experiments confirmed that the hybrid design outperformed each single-architecture variant under identical conditions, supporting the premise that global contextual modeling and sequential dependency learning are complementary rather than redundant.</p>
<p>Equally important is how BT-Synergy represents the cellular context. Rather than relying on manually curated similarity networks, the model constructs cell-line embeddings from pre-trained protein language models. For each drug-cell-line instance, the researchers compile the union of proteins that are either annotated drug targets or observed as expressed in the relevant cancer cell line, drawing on drug-protein interaction and cell-line protein expression matrices inherited from the DeepTraSynergy dataset. This union typically spans between 54 and 1,479 proteins per sample, averaging roughly 354. Each protein&#8217;s canonical amino acid sequence is retrieved from UniProt and encoded with ProteinBERT from the TAPE suite, which produces dense vectors capturing local residue motifs and longer-range sequence dependencies. A learnable attention-based pooling layer then weights each protein embedding by its relevance, aggregating them into a single fixed-size cell representation that can be trained end-to-end with the rest of the network.</p>
<p>Fusion of the chemical and biological streams happens through a dual-fusion module designed to capture higher-order cross-modal interactions. The two drug embeddings are concatenated, then combined with the cell-line embedding via element-wise multiplication, addition, and subtraction. Multiplication emphasizes synergistic effects, addition captures complementary relationships, and subtraction highlights contrastive signals between molecular and cellular modalities. This interaction-aware scheme replaces naive concatenation, which a baseline variant confirmed is less effective. Because drug combinations are biologically symmetric, the team also applied order-invariance augmentation, generating mirrored training samples in which the two drugs are swapped. The augmentation paid off: across five cross-validation folds, predictions for original and reversed pairs showed a correlation of 0.9721 with a mean absolute difference of just 0.0511, indicating the model treats drug order symmetrically as biology demands.</p>
<p>Training and evaluation relied on two heterogeneous benchmarks. DrugCombDB contributed 69,436 drug-pair-cell-line observations spanning 764 compounds and 76 cancer cell lines, scored with the Zero Interaction Potency metric, whose values cluster tightly around zero. OncologyScreen, by contrast, contains 4,176 observations across 29 compounds and 21 cell lines, scored with the Loewe additivity model, which spans a far wider numerical range. To harmonize these divergent scales and combat class imbalance, the researchers adopted a quantile-based discretization: pairs in the upper quartile of each dataset&#8217;s score distribution were labeled synergistic, those in the lower quartile non-synergistic, and the ambiguous middle half was excluded. A sensitivity analysis comparing 50/50, 33/67, and 25/75 thresholds showed that including low-confidence pairs introduces substantial label noise. The strictest 25/75 configuration delivered the best trade-off, with accuracy of 0.8458, AUC-ROC of 0.9229, and F1 of 0.8422 on DrugCombDB.</p>
<p>The model also held up under punishing robustness protocols. In leave-one-drug-out evaluation, where all combinations involving held-out drugs are removed from training, BT-Synergy achieved an AUC-ROC of 0.8349; in leave-one-cell-line-out testing it reached 0.8357, suggesting genuine resilience to unseen drugs and biological contexts. When trained exclusively on DrugCombDB and tested on the entirely non-overlapping OncologyScreen dataset, the model retained encouraging predictive performance, providing preliminary evidence of cross-dataset transfer, though the authors caution that differing synergy-scoring systems limit strong generalizability claims. Interpretability analyses reinforced the picture: attention heatmaps revealed both globally distributed attention, integrating distant structural components, and sharply localized focus on chemically salient SELFIES symbols such as branching indicators, double-bond notations, and heteroatom tokens. In a token ablation experiment, masking the highest-attention fragments dropped one predicted synergy probability from 0.476 to 0.175, a striking decrease that suggests the model&#8217;s decisions hinge on specific molecular motifs, although the researchers stress that attention weights are proxy indicators rather than proven mechanisms.</p>
<p>Per-drug subgroup analysis added a biologically coherent note. Among the 29 OncologyScreen compounds, the model performed best on drugs with well-characterized mechanisms of action: 5-fluorouracil, an antimetabolite targeting thymidylate synthase, achieved a per-drug AUC-ROC of 0.9354, methotrexate, which inhibits dihydrofolate reductase, scored 0.9055, and doxorubicin, a DNA-targeting agent, reached 0.8850. Compounds with broad, pleiotropic, or poorly defined pharmacology fared noticeably worse. Across all 21 cancer cell lines, performance remained stable, with AUC-ROC values generally between 0.72 and 0.85, indicating the protein-informed cell representations prevent over-specialization to particular cellular backgrounds. Compared against DeepSynergy, GraphSynergy, NEXGB, DeepTraSynergy, and CFSSynergy, BT-Synergy delivered competitive performance on both benchmarks, an outcome the authors attribute to the combination of chemically valid SELFIES encoding, the BiLSTM-Transformer hybrid, and biologically informed protein embeddings.</p>
<p>The limitations are candidly acknowledged. Quantile-based binarization excludes half of the experimental spectrum, so reported performance reflects clearly defined observations rather than the full continuous distribution of synergy. Differences between ZIP and Loewe scoring constrain interpretations of transfer learning, and data sparsity plus the multi-target nature of complex biology mean performance will vary across contexts. The authors call for future validation using harmonized synergy measurements, continuous-label prediction, and additional independent pharmacological benchmarks, alongside extensions to multi-drug combinations and richer omics modalities. Even with those caveats, BT-Synergy demonstrates that fusing sequence-aware molecular encoders with protein language model embeddings can push drug synergy prediction toward the accuracy and robustness that precision oncology demands, and with the source code released on GitHub and both datasets publicly available, the framework is positioned to be tested, extended, and potentially deployed in the search for the next life-extending drug combination.</p>
<p><strong>Subject of Research:</strong> A hybrid deep learning model combining BiLSTM and Transformer architectures with protein embeddings to predict synergistic cancer drug combinations</p>
<p><strong>Article Title:</strong> A hybrid BiLSTM transformer model for drug synergy prediction</p>
<p><strong>Article References:</strong> Rostami, S. A., &amp; Lakizadeh, A. (2026). A hybrid BiLSTM transformer model for drug synergy prediction. <em>Discover Artificial Intelligence, 6</em>(1), Article 1182. <a href="https://doi.org/10.1007/s44163-026-02262-4" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02262-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02262-4" rel="noopener noreferrer">10.1007/s44163-026-02262-4</a></p>
<p><strong>Keywords:</strong> drug synergy prediction, BT-Synergy, BiLSTM, Transformer, SELFIES, ProteinBERT, combination therapy, cancer cell lines, DrugCombDB, deep learning, computational pharmacology, protein embeddings</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201632</post-id>	</item>
		<item>
		<title>AI Co-Teacher Spots Classroom Distraction in Real Time With 90% Accuracy</title>
		<link>https://scienmag.com/ai-co-teacher-spots-classroom-distraction-in-real-time-with-90-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:44:58 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI classroom monitoring]]></category>
		<category><![CDATA[AI co-teacher systems]]></category>
		<category><![CDATA[AI-assisted teaching tools]]></category>
		<category><![CDATA[AI-powered classroom management tools]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[automated student attention tracking]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[classroom behavior analysis]]></category>
		<category><![CDATA[classroom engagement]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[distraction detection]]></category>
		<category><![CDATA[distraction severity scoring in classrooms]]></category>
		<category><![CDATA[edge computing in education]]></category>
		<category><![CDATA[education technology]]></category>
		<category><![CDATA[educational technology for engagement]]></category>
		<category><![CDATA[fog computing]]></category>
		<category><![CDATA[multimodal fusion]]></category>
		<category><![CDATA[objective classroom observation methods]]></category>
		<category><![CDATA[real-time student distraction detection]]></category>
		<category><![CDATA[scalable student engagement measurement]]></category>
		<category><![CDATA[SDG 4]]></category>
		<category><![CDATA[smart education]]></category>
		<category><![CDATA[speech recognition]]></category>
		<category><![CDATA[student behavior]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196527</guid>

					<description><![CDATA[Researchers have developed AI-TEACH, a multimodal AI framework that detects and scores classroom distraction in real time with 90% accuracy and significantly improved student outcomes.]]></description>
										<content:encoded><![CDATA[<p>Every teacher knows the moment: a lesson that began with real momentum slowly loses its grip as students drift toward their phones, slump over their desks, or slip into whispered side conversations. Distraction is one of the most persistent and least measured problems in education, and the tools educators currently have to detect it—manual observation, student self-reports, and post-lesson surveys—are subjective, delayed, and impractical at scale. A research team now reports a system designed to change that. In a study published in the journal Cognitive Computation, Munish Saini and Harsh Sharma of Guru Nanak Dev University in India, together with Eshan Sengupta of Vilnius Gediminas Technical University in Lithuania, introduce AI-TEACH, an artificial intelligence framework that continuously watches and listens to a classroom, identifies distraction events as they happen, scores their severity, and hands teachers an actionable engagement report at the end of every session.</p>
<p>The core idea behind AI-TEACH is what the researchers call a co-teacher paradigm. Rather than replacing educator judgment, the system acts as a silent, objective partner that augments it. Strategically positioned surveillance cameras with synchronized audio capture stream classroom data to an edge computing layer, where parallel pipelines analyze behavior and speech in near real time. On the video side, frames are preprocessed with OpenCV and passed through YOLO-NAS, a neural architecture search-optimized object detector that isolates each student in the room. A tracking algorithm called ByteTrack then assigns each detected student a persistent identity across frames, preserving the spatial-temporal coherence needed to study individual behavior over time. MediaPipe Pose extracts skeletal landmarks—nose, eyes, shoulders, wrists, hips—while MediaPipe&#8217;s facial model tracks 468 landmark points across the mouth, eyes, and jawline, providing the geometric substrate from which distraction cues are inferred.</p>
<p>From those landmarks, the system derives a structured taxonomy of behavioral distraction. A neck-tilt angle computed between the head axis and torso axis in three-dimensional skeletal space reveals slouching or simulated sleep, especially when combined with prolonged eye closure measured through the eye-aspect ratio and sustained stillness. Wrist-to-hip proximity paired with a downward gaze estimate flags mobile phone use. Rapid arm extension combined with a fast-moving optical flow contour, extracted via the Farnebäck method, identifies thrown objects. Exaggerated facial expressions are caught by measuring the standard deviation of facial landmark displacements across a sliding window of roughly ten frames, and an abrupt displacement of the hip keypoint toward a mapped door location signals a student bolting from class. Each detected incident is logged with a timestamp, an indicator type, and the responsible modality.</p>
<p>The audio pipeline complements the cameras by capturing the verbal disruptions that vision alone cannot see. Incoming sound is first segmented using Silero Voice Activity Detection, a lightweight neural model that separates speech from background noise with low latency. Detected speech segments are then passed through Wav2Vec2, a self-supervised speech model that produces rich acoustic embeddings, which a Support Vector Machine classifies into specific distraction categories: off-task back-talk occurring outside the teacher&#8217;s directional audio profile, loud or chaotic sound events such as shouting or whistling identified through amplitude envelope tracking and spectral flatness, rhythmic tapping detected via autocorrelation and short-time energy, and abusive or offensive language flagged by a toxicity classifier applied to transcribed utterances. The design consciously draws on recent advances in multimodal emotion recognition, including per-sample modality equilibration and adversarial fusion techniques, to keep weaker signals from being drowned out by dominant ones when audio and video streams arrive asynchronously.</p>
<p>Fusing these streams is the job of a Bidirectional Long Short-Term Memory network, a recurrent architecture that reads each multimodal feature sequence in both temporal directions, capturing escalating patterns of disruption that a frame-by-frame classifier would miss. A softmax layer over the pooled hidden states assigns each incident a distraction category, and every category carries an empirically calibrated severity weight reflecting its actual disruptive impact on the class—from a single slouching student to thrown objects or hostile verbal outbursts. Over the course of a session, the fog layer multiplies each indicator&#8217;s occurrence count by its severity score and sums the results into a single Class Distraction Score, alongside timestamped behavioral metadata. Only this structured metadata, not the raw recordings, is transmitted over encrypted channels to the cloud layer, where an interactive dashboard visualizes the score, breaks down each indicator with its frequency and severity, and plots a temporal heatmap of when distraction clustered during the lesson.</p>
<p>The engineering choices are deliberately pragmatic. The fog layer runs on edge hardware comparable to an NVIDIA Jetson Xavier NX, sustaining frame-level video analytics and concurrent audio processing at 25 to 30 frames per second, while the cloud tier requires only a server-class GPU and at least 10 Mbps of upstream bandwidth. Under these conditions the system achieves an average end-to-end latency of roughly 350 milliseconds from frame capture to logged distraction event, consuming about 15 watts at the edge during continuous operation. Teachers need no technical training beyond an estimated one to two hours of onboarding, because the system is designed for post-session reflection rather than in-the-moment alarms: educators review the dashboard, add contextual notes where needed, and adjust pacing, format, or grouping in subsequent lessons.</p>
<p>The validation combined benchmark testing with a live classroom experiment. The team assembled a curated dataset of more than 5,350 labeled classroom images drawn from three open-source repositories, spanning behaviors such as phone use, object throwing, slouching, yawning, gaze aversion, and head-on-desk posture. The video module achieved 90% overall accuracy and a weighted average recall of 92%, with a macro-average F1 score of 92% across nine evaluation rounds, indicating balanced performance on both frequent and rare distraction categories. The audio module performed comparably, with accuracy between 88% and 94% and a macro-average F1 of 92%, maintaining reliable classification despite overlapping speech and ambient noise. In head-to-head comparisons, AI-TEACH&#8217;s overall F1 score of 93.5% substantially outperformed conventional baselines, including a convolutional neural network at 74.14%, a standalone LSTM at 76.04%, HOG plus SVM at 80.52%, OpenPose plus SVM at 83.26%, and even a BiLSTM used alone at 85.95%—evidence that the advantage comes from the fusion of complementary modalities and severity weighting rather than any single component.</p>
<p>The pedagogical test was a controlled experiment with 200 students randomly assigned to experimental and control groups. Control classrooms received traditional instruction, while teachers in the experimental group worked with AI-TEACH&#8217;s real-time feedback. After baseline pretests, the experimental group gained an average of 9 points on post-tests compared with 5 points for the control group, and a two-sample t-test confirmed the difference was statistically significant at the 0.05 level. The calculated effect size, Cohen&#8217;s d of approximately 0.90, qualifies as large by conventional standards—a striking result for an intervention that changes nothing about the curriculum itself and everything about how quickly teachers can perceive and respond to disengagement. The finding aligns with earlier research showing that objectively measured engagement predicts academic performance better than self-reported distraction.</p>
<p>The authors are candid about the ethical terrain. Because several monitored behaviors—fidgeting, repetitive movement, gaze aversion—can be natural self-regulation strategies for neurodivergent students with autism or ADHD, the system is explicitly designed as a class-level tool rather than an individual diagnostic instrument, and future versions will let educators suppress specific indicators for particular students. Algorithmic bias across skin tones, body types, cultural behavioral norms, and languages remains a recognized limitation, as does the single-institution setting of the 200-student trial and the absence of physiological signals such as heart rate variability. Raw audio and video never leave the edge; only encrypted metadata is stored remotely under role-based access control, with informed consent and periodic bias audits recommended as deployment prerequisites. Framed against the United Nations Sustainable Development Goals, particularly SDG 4 on quality education, AI-TEACH represents a broader shift toward classrooms where attention itself becomes measurable, and where teachers—armed with evidence instead of intuition—can reach drifting students before the drifting becomes permanent.</p>
<p><strong>Subject of Research:</strong> Multimodal AI-based real-time assessment of student distraction and engagement in classrooms</p>
<p><strong>Article Title:</strong> Artificial Intelligence Based Framework for Student Engagement Assessment in Classroom Environments</p>
<p><strong>Article References:</strong> Saini, M., Sharma, H., &amp; Sengupta, E. (2026). Artificial Intelligence Based Framework for Student Engagement Assessment in Classroom Environments. <em>Cognitive Computation, 18</em>(1), Article 108. <a href="https://doi.org/10.1007/s12559-026-10629-z" rel="noopener noreferrer">https://doi.org/10.1007/s12559-026-10629-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12559-026-10629-z" rel="noopener noreferrer">10.1007/s12559-026-10629-z</a></p>
<p><strong>Keywords:</strong> artificial intelligence, classroom engagement, distraction detection, computer vision, speech recognition, BiLSTM, fog computing, education technology, multimodal fusion, smart education, SDG 4, student behavior</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196527</post-id>	</item>
	</channel>
</rss>
