<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI-based fault diagnosis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-based-fault-diagnosis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 02:52:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI-based fault diagnosis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Hunt the Cloud&#8217;s Most Dangerous Failures Before They Strike</title>
		<link>https://scienmag.com/ai-learns-to-hunt-the-clouds-most-dangerous-failures-before-they-strike/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 02:52:11 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based fault diagnosis]]></category>
		<category><![CDATA[Alibaba Cluster Trace]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[anomaly detection in cloud computing]]></category>
		<category><![CDATA[autoencoder]]></category>
		<category><![CDATA[cloud computing]]></category>
		<category><![CDATA[Cloud infrastructure failure detection]]></category>
		<category><![CDATA[cloud monitoring limitations]]></category>
		<category><![CDATA[cloud system resilience]]></category>
		<category><![CDATA[distributed system reliability]]></category>
		<category><![CDATA[Google Borg Trace]]></category>
		<category><![CDATA[gray failure]]></category>
		<category><![CDATA[gray failures in data centers]]></category>
		<category><![CDATA[hierarchical reinforcement learning]]></category>
		<category><![CDATA[innovative failure mitigation techniques]]></category>
		<category><![CDATA[machine learning for cloud failure prediction]]></category>
		<category><![CDATA[migration scheduling]]></category>
		<category><![CDATA[proactive workload management]]></category>
		<category><![CDATA[robust PCA]]></category>
		<category><![CDATA[self-healing cloud systems]]></category>
		<category><![CDATA[self-healing systems]]></category>
		<category><![CDATA[SLO violations]]></category>
		<category><![CDATA[subtle hardware and network failures]]></category>
		<category><![CDATA[temporal convolutional network]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=246190</guid>

					<description><![CDATA[A new closed-loop reinforcement learning framework called A3SM detects, localizes, and mitigates gray failures in cloud clusters, cutting recovery times and service violations in trace-driven experiments.]]></description>
										<content:encoded><![CDATA[<p>Some of the most dangerous failures in the world&#8217;s data centers never announce themselves. A server does not crash, a network cable does not snap, and yet somewhere inside a cloud cluster a machine has quietly slowed to a crawl, dropping packets or leaking performance in ways that static alarms simply cannot see. Researchers call these elusive malfunctions gray failures, because the affected components remain technically alive while operating in a degraded twilight state. Now a team of Chinese computer scientists has unveiled a framework that not only detects these weak anomalies but also decides, on its own, when to intervene and where to move the affected workloads. The system, described in the journal Cluster Computing, is called A3SM, and its results suggest that self-healing cloud infrastructure may be closer than many operators think.</p>
<p>The problem A3SM attacks is notoriously subtle. Gray failures were famously described by Microsoft researchers in 2017 as the Achilles&#8217; heel of cloud-scale systems, precisely because they evade the binary logic of conventional monitoring. A node that responds to health checks but serves requests at half its normal speed can poison an entire distributed application, dragging down latency for millions of users while every dashboard shows green. The difficulty, as the authors of the new study emphasize, is not merely detecting these weak anomalies. It is deciding when to intervene, at what granularity to act, and where to place the affected workload once a migration is triggered, all under strict cost constraints. A cloud operator who reacts too aggressively will churn thousands of virtual machines unnecessarily; one who reacts too slowly will watch service-level objectives crumble.</p>
<p>A3SM, short for a closed-loop anomaly-aware migration scheduling framework, closes that detection-to-mitigation gap with a layered architecture built on hierarchical reinforcement learning. At the front end sit four complementary anomaly detectors, each designed to catch a different signature of degradation. A Temporal Convolutional Network performs residual prediction, learning the expected rhythm of system metrics and flagging deviations from forecast behavior. An autoencoder reconstructs incoming metric vectors and measures how badly it fails to compress them, since data that does not fit the learned normal pattern often signals trouble. Robust Principal Component Analysis, a mathematical technique originally developed for separating signal from gross corruption, isolates sparse anomalies hidden in dense metric streams. Finally, graph-consistency analysis examines the relationships among nodes, catching cases where an individual machine looks healthy but behaves inconsistently with its peers.</p>
<p>Detection alone, however, produces only a risk score. To act intelligently, the system must localize the blast radius of an anomaly. A3SM accomplishes this through a three-tier attribution scheme that assigns blame at the task level, the node level, or the group level. A single misbehaving process might warrant moving one task; a sick machine might require evacuating everything running on it; a correlated failure affecting a rack or a service group might demand a broader reshuffle. By pinning down the scope of impact before any action is taken, the framework avoids the costly overreaction that has plagued simpler self-healing schemes, where a minor hiccup could trigger a cascade of unnecessary migrations across an entire cluster.</p>
<p>The heart of A3SM is a two-level reinforcement learning policy that elegantly decouples two decisions that are usually tangled together: whether to intervene at all, and where to send the affected workload. The upper-level policy weighs the evidence of gray-failure risk against the disruption cost of migration and decides whether the moment has come to act. The lower-level policy, activated only when intervention is warranted, selects the target node that best balances load, capacity, and expected benefit. This separation mirrors the way experienced site reliability engineers think, first triaging the situation and only then choosing a remedy, and it allows each policy to learn on its own far more manageable decision space.</p>
<p>What makes the loop truly closed is an online benefit-cost feedback mechanism that continuously updates the system&#8217;s scheduling sensitivity and decision policies. Every intervention produces an outcome, faster recovery or wasted effort, and that outcome feeds back into the learning process, tuning how eagerly the framework responds to future anomalies. In effect, A3SM calibrates its own paranoia. When migrations prove beneficial, it becomes more willing to act; when they prove wasteful, it raises the bar for intervention. This adaptivity is crucial in production clouds, where workload patterns shift hourly and a statically tuned threshold would quickly go stale.</p>
<p>The experimental evidence comes from trace-driven replay of two of the most famous datasets in cluster research: the Alibaba Cluster Trace 2017 and the Google Borg Trace 2020. On the Alibaba data, A3SM achieved an anomaly-detection F1-score of 0.82, a strong result for a problem where weak signals blur into normal noise. More striking are the operational numbers: the framework reduced the violation rate of resource service-level objectives to 3.2 percent and shortened the average recovery time to 185 seconds, all while cutting the number of unnecessary migrations. On the Google Borg trace, the framework transferred with only modest degradation, posting an F1-score of 0.79, a violation rate of 3.8 percent, and a mean time to recovery of 202 seconds. That cross-platform consistency matters, because a mitigation tool that only works on one provider&#8217;s workload profile would be of limited value to the industry.</p>
<p>The authors, led by Juntao Ye and Yuanchen Sun of the University of Shanghai for Science and Technology and Shenyang Normal University, with contributions from Di Wu, Qianqian Duan, and corresponding author Xing Hu, are careful to state the limits of their claims. The experiments rely on trace replay and controlled fault injection rather than live production clusters, and real-world gray failures can be messier than anything captured in a recorded trace. Still, the pattern of results, improved recovery timeliness, fewer service violations, and reduced migration churn, holds across both evaluated environments, suggesting that the underlying principle of closed-loop, anomaly-aware scheduling is robust rather than dataset-specific.</p>
<p>The broader significance of this work lies in its vision of cloud infrastructure that manages its own pathology. Modern hyperscale data centers already automate enormous amounts of routine scheduling, but failure response has remained stubbornly human, dependent on on-call engineers paging through dashboards at three in the morning. A framework that fuses multiple detection paradigms, localizes faults at the right granularity, and learns from the consequences of its own interventions points toward a future in which the cluster itself is the first responder. As cloud-native applications grow more complex and microservice architectures multiply the number of components that can quietly degrade, the gap between detection and mitigation has become one of the most expensive weaknesses in modern computing. A3SM demonstrates that hierarchical reinforcement learning, armed with a diverse ensemble of anomaly detectors and an honest accounting of intervention costs, can bridge that gap, turning the cloud&#8217;s grayest failures from silent killers into managed events.</p>
<p><strong>Subject of Research:</strong> Closed-loop anomaly-aware migration scheduling for gray failure mitigation in cloud clusters using hierarchical reinforcement learning</p>
<p><strong>Article Title:</strong> A3SM: Closed-loop anomaly-aware migration scheduling for gray failure mitigation in cloud clusters</p>
<p><strong>Article References:</strong> Ye, J., Sun, Y., Wu, D., Duan, Q., &amp; Hu, X. (2026). A3SM: Closed-loop anomaly-aware migration scheduling for gray failure mitigation in cloud clusters. <em>Cluster Computing, 29</em>(13), Article 763. <a href="https://doi.org/10.1007/s10586-026-06594-9" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06594-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06594-9" rel="noopener noreferrer">10.1007/s10586-026-06594-9</a></p>
<p><strong>Keywords:</strong> cloud computing, gray failure, anomaly detection, hierarchical reinforcement learning, migration scheduling, autoencoder, temporal convolutional network, robust PCA, Alibaba Cluster Trace, Google Borg Trace, SLO violations, self-healing systems</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">246190</post-id>	</item>
	</channel>
</rss>
