<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>K-ARM &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/k-arm/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 12:55:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>K-ARM &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Hidden Fingerprints: New Method Uncovers Backdoor Targets in Compromised Neural Networks</title>
		<link>https://scienmag.com/hidden-fingerprints-new-method-uncovers-backdoor-targets-in-compromised-neural-networks/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 12:55:00 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[autonomous vehicle neural network safety]]></category>
		<category><![CDATA[backdoor attack vulnerabilities in neural networks]]></category>
		<category><![CDATA[backdoor attacks]]></category>
		<category><![CDATA[backdoor defenses]]></category>
		<category><![CDATA[class transfer pattern in poisoned neural networks]]></category>
		<category><![CDATA[combating backdoor attacks in facial recognition]]></category>
		<category><![CDATA[data poisoning]]></category>
		<category><![CDATA[deep neural networks]]></category>
		<category><![CDATA[detecting adversarial triggers in AI models]]></category>
		<category><![CDATA[hidden fingerprint analysis in deep learning]]></category>
		<category><![CDATA[K-ARM]]></category>
		<category><![CDATA[label noise]]></category>
		<category><![CDATA[methods for uncovering backdoor targets]]></category>
		<category><![CDATA[model security]]></category>
		<category><![CDATA[Neural Cleanse]]></category>
		<category><![CDATA[neural network backdoor detection]]></category>
		<category><![CDATA[neural network security and integrity]]></category>
		<category><![CDATA[statistical fingerprinting for model security]]></category>
		<category><![CDATA[training pipeline security in deep learning]]></category>
		<category><![CDATA[transition matrix]]></category>
		<category><![CDATA[trigger inversion]]></category>
		<category><![CDATA[watermark and trigger detection in AI systems]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=238076</guid>

					<description><![CDATA[Researchers have developed a transition matrix-based framework that identifies the target classes of backdoor attacks in neural networks and sharply improves existing defenses.]]></description>
										<content:encoded><![CDATA[<p>Deep neural networks now power everything from facial recognition systems to autonomous vehicles, but a quiet and dangerous threat has been lurking in their training pipelines. Known as backdoor attacks, these assaults embed invisible triggers into a model during training, so that the network behaves perfectly on ordinary inputs but flips to an attacker-chosen output the moment a specific pattern appears. A stop sign with a tiny sticker, a face with a subtle perturbation, a document with a nearly imperceptible watermark — any of these can hijack the model&#8217;s decision. A new study published in Applied Intelligence by Wenbin Jiang, Shikang Xie, Jiqiang Liu, Nan Jiang and Jian Wang, researchers at Beijing Jiaotong University and Beijing University of Technology, offers a fresh way to fight back by reading the statistical fingerprints that such attacks inevitably leave behind in a model&#8217;s predictions.</p>
<p>The core insight of the work is deceptively simple: attackers have intentions, and those intentions impose structure. When a backdoor is active, poisoned samples are systematically pushed from their true classes toward a single target class chosen by the attacker. This creates what the authors call a class transfer pattern — a measurable, recurring flow of predictions from many source classes into one destination class. While most existing defenses concentrate either on discovering the trigger itself or on unlearning the backdoor behavior, they often struggle to characterize which class the attacker actually targeted. The new framework flips the problem around: instead of hunting for the trigger first, it estimates the pattern of class-to-class prediction transfers and uses that estimate to pinpoint the target class, even when no prior knowledge of the attack exists.</p>
<p>Technically, the method builds on an idea borrowed from the study of label noise. In noisy-label learning, researchers model the probability that a true label is flipped into an incorrect one using a transition matrix — a table whose entries describe how likely each class is to be misread as every other class. The authors adapt this machinery to the backdoor setting. Because poisoned samples are a small minority of the training data, the attack-induced transfers appear as large off-diagonal entries in an otherwise diagonal-dominant matrix: a well-trained model correctly classifies most samples, so the diagonal entries, representing correct predictions, dominate, while systematic misclassifications caused by poisoning stand out as anomalous rows and columns. Estimating this matrix therefore becomes a way of exposing the attack&#8217;s geometry.</p>
<p>The estimation itself proceeds in two steps and is driven by an improved information-theoretic objective. The researchers draw on the determinant of the joint matrix between predictions and labels, a loss function known in the label-noise literature as DMI, which penalizes degenerate transition estimates. In an appendix, the team provides a formal proof bounding the gradient of this loss: using the matrix differential identity for the log-determinant and Weyl&#8217;s inequality on singular values, they show that the per-sample gradient norm is bounded by the inverse of the smallest singular value of the joint matrix. This matters because it formalizes when the estimation is stable — namely, when the diagonal clean-prediction component remains stronger than the off-diagonal attack-induced component, a condition expected to hold whenever the poisoning ratio is moderate and the model retains high clean accuracy.</p>
<p>Once the transition matrix has been estimated, the identified target class becomes a powerful lever for existing defenses. Trigger-inversion methods, such as the widely used Neural Cleanse and the K-ARM optimization approach, work by reverse-engineering the smallest perturbation that can force any input to be classified as a suspect target. The trouble is that these methods must search across every possible target class, reconstructing candidate triggers for each one — a computationally expensive process prone to false positives. By narrowing the search space to the class flagged by the transition matrix, the new framework eliminates unnecessary trigger reconstruction while preserving mitigation effectiveness. The defense knows where to look before it starts looking.</p>
<p>The experimental results are striking. Across multiple datasets and attack types, the enhanced Neural Cleanse variant achieved an average attack success rate of just 2.17 percent, while the enhanced K-ARM variant drove the figure down to 1.46 percent. The strongest baseline, by comparison, averaged an attack success rate of 3.68 percent. Crucially, these gains did not come at the cost of normal performance: the defended models maintained higher clean accuracy than the baselines, meaning legitimate users would notice no degradation in everyday behavior. The evaluation covered standard image benchmarks including CIFAR-10, CIFAR-100 and Tiny ImageNet, all publicly available datasets, and spanned a range of attack designs from classic patch-based triggers to more modern imperceptible variants.</p>
<p>The breadth of attacks considered reflects the evolving threat landscape. The study&#8217;s reference list traces the lineage of backdoor research from BadNets, the 2017 work that first identified vulnerabilities in the machine learning supply chain, through targeted data-poisoning attacks, Trojaning attacks, reflection-based natural backdoors, warping-based WaNet triggers and label-consistent attacks that try to evade detection by aligning poisoned samples with their labels. Each generation of attacks has grown stealthier, and frequency-domain analyses have shown that triggers can hide in spectral regions humans barely perceive. A defense that does not depend on the specific form of the trigger — but instead on the invariant statistical consequence of any targeted attack — is inherently more robust to this arms race.</p>
<p>That invariance is precisely what makes the approach compelling. Whether an attacker uses a visible patch, a transparent overlay, a warping field or a frequency-domain perturbation, the end goal is the same: route inputs from diverse source classes into one chosen target. The transition matrix captures this routing behavior directly. The authors&#8217; theoretical analysis reinforces the point by decomposing the joint matrix into a diagonal component, representing the model&#8217;s legitimate functionality on the majority of samples, and an off-diagonal component, where large entries signal systematic misclassification toward the attacker&#8217;s target. Under the reasonable assumption that no exact linear dependencies exist among classes, the matrix is invertible, and its spectral properties govern the reliability of the estimate.</p>
<p>The practical implications extend well beyond computer vision benchmarks. Backdoor attacks are a supply-chain problem: models downloaded from public repositories, trained by third-party contractors or fine-tuned on crowdsourced data can all carry hidden triggers. Scanning such models before deployment is becoming a standard security practice, and methods that make scanning faster and more accurate have immediate value. By telling trigger-inversion defenses which class to investigate, the transition-matrix framework could reduce the computational cost of auditing large models, an increasingly important consideration as networks grow to billions of parameters. The same reasoning may eventually extend to language models, where backdoor token unlearning has emerged as a parallel research frontier.</p>
<p>Limitations remain, as with any defense. The framework&#8217;s stability guarantee rests on the assumption that correct predictions dominate misclassifications, which could weaken under extreme poisoning ratios or attacks that spread their effect across multiple targets. The authors acknowledge that the approach assumes moderate poisoning and high clean accuracy — conditions that describe most realistic attacks but not necessarily the most aggressive adversarial scenarios. Still, the work represents a meaningful conceptual shift: treating the attack target not as a mystery to be brute-forced but as a statistical signature to be estimated. As deep learning systems take on higher-stakes roles in medicine, infrastructure and transportation, defenses that exploit the inherent structure of the attacker&#8217;s intent — rather than the specifics of any single trigger — may prove to be the durable line of protection the field has been searching for.</p>
<p><strong>Subject of Research:</strong> Backdoor attack defense in deep neural networks using class transfer matrix estimation</p>
<p><strong>Article Title:</strong> Class transfer estimation in neural networks for backdoor target identification and mitigation</p>
<p><strong>Article References:</strong> Class transfer estimation in neural networks for backdoor target identification and mitigation. (n.d.). <a href="https://doi.org/10.1007/s10489-026-07470-0" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07470-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07470-0" rel="noopener noreferrer">10.1007/s10489-026-07470-0</a></p>
<p><strong>Keywords:</strong> deep neural networks, backdoor attacks, backdoor defenses, transition matrix, trigger inversion, Neural Cleanse, K-ARM, label noise, adversarial machine learning, model security, data poisoning, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">238076</post-id>	</item>
	</channel>
</rss>
