<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>self-attention &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/self-attention/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 01:31:42 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>self-attention &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Tiny Transformer Offers Early Warning Against Stealthy Attacks on Industrial IoT</title>
		<link>https://scienmag.com/tiny-transformer-offers-early-warning-against-stealthy-attacks-on-industrial-iot/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 01:31:42 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced persistent threats]]></category>
		<category><![CDATA[AI in industrial network threat monitoring]]></category>
		<category><![CDATA[AI-driven intrusion detection for industrial networks]]></category>
		<category><![CDATA[CICAPT-IIoT dataset]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[Compact AI models for resource-constrained devices]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[Early detection of advanced persistent threats]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[Edge computing security in industrial IoT]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[industrial IoT]]></category>
		<category><![CDATA[Industrial IoT cybersecurity]]></category>
		<category><![CDATA[intrusion detection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Machine learning for IoT attack prevention]]></category>
		<category><![CDATA[Multi-stage cyberattack detection in industrial systems]]></category>
		<category><![CDATA[provenance data]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[Self-attention mechanisms in cyber threat detection]]></category>
		<category><![CDATA[Sequence modeling in industrial cybersecurity]]></category>
		<category><![CDATA[Small transformer models for edge device security]]></category>
		<category><![CDATA[Stealthy cyberattack identification using transformers]]></category>
		<category><![CDATA[Transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224886</guid>

					<description><![CDATA[Researchers in Beijing have built a 0.27-megabyte transformer model that detects multi-stage APT attacks in industrial IoT telemetry with high precision and microsecond latency, positioning lightweight attention-based sequence modeling as a calibrated early-warning layer for edge defenses.]]></description>
										<content:encoded><![CDATA[<p>Industrial systems that once ran in isolation are now stitched into networks of connected sensors, controllers, and gateways, and that connectivity has opened the door to one of the most dangerous categories of cyberattack: the advanced persistent threat, or APT. These intrusions are patient, multi-stage campaigns in which an adversary quietly maps a network, escalates privileges, and moves laterally toward critical assets, all while hiding inside enormous streams of ordinary operational telemetry. A new study published in the International Journal of Machine Learning and Cybernetics by Ramadhani Zuberi Nyangusi and Hongsong Chen of the University of Science and Technology Beijing tackles this problem with an unusually small piece of artificial intelligence: a compact transformer model designed to run on the resource-starved edge devices that guard industrial Internet of Things deployments.</p>
<p>The appeal of transformer architectures in cybersecurity is easy to understand. Since the landmark 2017 paper &#8220;Attention Is All You Need,&#8221; self-attention mechanisms have transformed natural language processing and, more recently, sequence modeling in security applications, because they can weigh the relationships between events in a sequence regardless of how far apart those events occur. For APT detection, that matters enormously. An attacker&#8217;s footprint is rarely a single anomalous packet; it is a chain of individually unremarkable actions whose significance emerges only when they are read together. Larger transformer models, however, carry millions of parameters and demand memory and compute budgets that typical IIoT gateways simply cannot provide, which is why many high-performing research models never leave the laboratory.</p>
<p>Nyangusi and Chen&#8217;s answer is a deliberately stripped-down transformer. Their framework uses just two encoder layers and two attention heads, with a model dimension of 64 and a feed-forward dimension of 128. The result is a network with only 69,057 trainable parameters and an approximate model size of 0.27 megabytes, small enough to plausibly sit on edge hardware rather than requiring a cloud round-trip for every decision. The design philosophy is context-awareness at minimal cost: rather than analyzing entire provenance graphs or long event histories, the system organizes provenance events into short temporal windows, allowing the attention mechanism to capture local temporal behavior while keeping the computational footprint tiny.</p>
<p>Class imbalance is the second central challenge the researchers confront head-on. In real industrial telemetry, malicious events are vanishingly rare compared with benign ones, and models trained naively on such data tend to achieve high accuracy while missing most actual attacks. The framework therefore employs focal loss, an imbalance-aware optimization objective that down-weights easy, well-classified examples and concentrates learning effort on the difficult minority cases that matter most. Just as importantly, the authors are explicit about methodology hygiene: decision thresholds are selected on a validation set before final testing, avoiding the test-set-driven calibration that can silently inflate reported performance in detection research.</p>
<p>The evaluation rests on the CICAPT-IIoT dataset, a publicly available provenance-based APT attack dataset for IIoT environments released by the Canadian Institute for Cybersecurity at the University of New Brunswick. Provenance data records the causal history of system activity, which makes it a natural substrate for spotting multi-stage intrusions. Across five random seeds, the best sequence-level configuration used a four-event temporal window and achieved a malicious precision of 0.8547 plus or minus 0.0205, a recall of 0.5025 plus or minus 0.0353, an F1-score of 0.6321 plus or minus 0.0263, a ROC-AUC of 0.8840 plus or minus 0.0131, and a PR-AUC of 0.5909 plus or minus 0.0193. Reporting across multiple seeds and including variance, rather than a single best run, gives these numbers a credibility that single-shot benchmarks often lack.</p>
<p>Those figures tell an honest and nuanced story. Precision above 0.85 means that when the model raises an alarm, it is right the vast majority of the time, which is exactly what operators of critical infrastructure need, since false alarms in a factory or power grid carry real operational costs. Recall near 0.50, by contrast, means the model catches roughly half of malicious sequences, and the modest PR-AUC reflects the brutal arithmetic of extreme class imbalance. The authors do not paper over this trade-off. Instead, they position the framework explicitly as what it is: a compact, calibrated early-warning component for IIoT APT detection, not a universal replacement for all classical classifiers. In a layered defense, a lightweight sensor that reliably flags high-confidence threats at the edge has clear value even if deeper analysis systems handle the harder cases.</p>
<p>The resource profiling is where the work becomes genuinely striking for anyone thinking about deployment. CPU inference latency measured 0.0609 plus or minus 0.0002 milliseconds per four-event window, a figure so low that the model could, in principle, evaluate thousands of windows per second on modest hardware. Combined with the 0.27-megabyte footprint, this suggests the framework could be embedded directly into gateways, industrial PCs, or even constrained embedded devices, screening provenance streams continuously and escalating only suspicious sequences to heavier backend analysis. That division of labor, tiny models at the edge and heavyweight forensics in the core, is increasingly seen as the realistic architecture for securing sprawling industrial estates.</p>
<p>To test whether the approach generalizes beyond its home dataset, the researchers performed an external validation on Windows-APT 2025, a dataset of APT-inspired attack scenarios on Windows systems. The same temporal-window pipeline showed it could transfer to ATT&amp;CK-mapped Windows host-alert detection, suggesting the design is not merely tuned to the quirks of one provenance dataset. The authors are careful to note that this second setting uses a different telemetry source and a proxy-label structure, so the transfer result is suggestive rather than definitive. Even so, the ability of one lightweight pipeline to operate across both IIoT provenance data and Windows host alerts hints at a portable pattern for early-stage threat detection across heterogeneous environments.</p>
<p>The study also situates itself within a rapidly crowding field. Recent years have produced transformer-based intrusion detectors, hybrid CNN-BiLSTM and Swin-transformer hybrids, diffusion-transformer models for imbalanced IoT learning, provenance-graph frameworks with masked representation learning, and knowledge-distillation approaches aimed at explainable detection. Many of these achieve strong classification metrics, but the Beijing team argues that too few provide evidence of deployment feasibility under edge-oriented resource constraints, and many gloss over the precision-recall trade-off that severe imbalance imposes. By publishing parameter counts, model sizes, latency figures, and seed-level variance alongside detection metrics, this work offers a template for how lightweight security AI should be evaluated: not just how well it detects, but whether it can actually run where the threats arrive.</p>
<p>For the operators of factories, utilities, and critical infrastructure, the takeaway is pragmatic rather than sensational. Advanced persistent threats will not be defeated by a single algorithm, and a detector that catches half of malicious sequences is not a silver bullet. But a 69,000-parameter model that fits in a fraction of a megabyte, responds in microseconds, and delivers high-precision alerts from raw provenance windows represents a meaningful building block for defense in depth. As industrial networks grow and attackers grow more patient, the future of cybersecurity may depend less on ever-larger models in distant data centers and more on swarms of small, fast, honest sentinels watching quietly at the edge, and this research shows exactly what such sentinels can, and cannot, yet do.</p>
<p><strong>Subject of Research:</strong> Lightweight transformer-based detection of advanced persistent threats in industrial Internet of Things environments</p>
<p><strong>Article Title:</strong> A lightweight transformer-based framework for context-aware APT detection in industrial IoT</p>
<p><strong>Article References:</strong> Nyangusi, R. Z., &amp; Chen, H. (2026). A lightweight transformer-based framework for context-aware APT detection in industrial IoT. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 484. <a href="https://doi.org/10.1007/s13042-026-03324-w" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03324-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03324-w" rel="noopener noreferrer">10.1007/s13042-026-03324-w</a></p>
<p><strong>Keywords:</strong> advanced persistent threats, industrial IoT, transformer, intrusion detection, edge computing, provenance data, focal loss, class imbalance, CICAPT-IIoT dataset, self-attention, cybersecurity, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224886</post-id>	</item>
		<item>
		<title>New AI Model Reads the Hidden Rhythms of Evolving Knowledge Graphs</title>
		<link>https://scienmag.com/new-ai-model-reads-the-hidden-rhythms-of-evolving-knowledge-graphs/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 21:04:20 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advances in applied intelligence for temporal data]]></category>
		<category><![CDATA[AI models for dynamic relationship prediction]]></category>
		<category><![CDATA[AI reasoning over changing networks]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[DIRA]]></category>
		<category><![CDATA[dynamic embedding]]></category>
		<category><![CDATA[dynamic embedding models]]></category>
		<category><![CDATA[event forecasting]]></category>
		<category><![CDATA[evolving relational data]]></category>
		<category><![CDATA[extrapolation reasoning]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[handling sparse and incomplete relational data]]></category>
		<category><![CDATA[implicit relation-aware self-attention]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[link prediction]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[modeling entity evolution in knowledge graphs]]></category>
		<category><![CDATA[predicting missing facts in TKGC]]></category>
		<category><![CDATA[rhythm detection in knowledge graphs]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[temporal knowledge graph completion]]></category>
		<category><![CDATA[temporal reasoning]]></category>
		<category><![CDATA[time-stamped fact inference]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223618</guid>

					<description><![CDATA[Researchers have developed DIRA, a model that combines dynamic entity embeddings with implicit relation-aware self-attention to significantly improve temporal knowledge graph completion on sparse, evolving data.]]></description>
										<content:encoded><![CDATA[<p>Knowledge is not static. Political alliances shift, trade relationships form and dissolve, and the actors on the world stage rise and fade with time. For artificial intelligence systems that try to reason over such ever-changing webs of facts, the challenge of predicting what will happen next, or of filling in the gaps in what is already known, has proven stubbornly difficult. A new study published in Applied Intelligence by Shuang Liu, Xiaohui Sun, Peng Chen, and Simon Kolmanič tackles this problem head-on with a model called DIRA, short for Dynamic embedding and Implicit Relation-aware Self-attention, designed specifically for temporal knowledge graph completion, the task of inferring missing facts in time-stamped relational data.</p>
<p>Temporal knowledge graphs, or TKGC as the research community abbreviates the completion task, represent facts as connections between entities that are only valid within particular time windows. A diplomatic visit recorded in 2014 does not imply the same relationship holds in 2019, and a conflict that flared in one decade may be entirely absent from another. Systems built to reason over these graphs must therefore do more than memorize connections; they must understand how entities evolve, how relationships pulse with periodic rhythms, and how sparse or incomplete records can still hint at facts that were never explicitly written down. This is especially critical in extrapolation scenarios, where a model must reason about the future rather than merely reconstruct the past, as in event forecasting and temporal question answering.</p>
<p>According to the authors, existing approaches suffer from two persistent weaknesses. The first is structural sparsity: real-world temporal knowledge graphs capture only a fraction of the facts that actually exist, so the explicit network of recorded connections is riddled with holes that mislead models trained to rely on direct structure alone. The second is inadequate modeling of how entities change across multiple dimensions of time. Most methods treat an entity&#8217;s representation as either frozen or evolving along a single trajectory, failing to capture the interplay between long-term trends, recurring periodic behavior, and stable static attributes. The result, the team argues, is inferior performance in sparse and long-term reasoning scenarios, precisely the settings where temporal reasoning matters most.</p>
<p>DIRA&#8217;s answer begins with a multi-component dynamic embedding scheme that explicitly separates an entity&#8217;s representation into three parts: a trend component that tracks directional change over time, a periodic component that captures cyclical patterns such as annually recurring diplomatic or economic events, and a static component that encodes properties that persist regardless of the timestamp. By modeling these three facets independently and then combining them, the model gains a richer portrait of each entity than a single time-dependent vector can provide. This decomposition reflects an intuition familiar to anyone who has studied time series analysis: the behavior of real-world actors is rarely purely trending or purely seasonal, but usually some blend of both layered over an unchanging core identity.</p>
<p>On top of this dynamic foundation, DIRA strengthens its view of the graph&#8217;s structure through relation-aware graph convolution, a technique descended from graph neural networks that propagates information between connected entities while taking the type of each relationship into account. A military alliance and a trade agreement are not interchangeable edges, and a convolution that respects relational semantics can distinguish the neighborhood of a head of state from that of a corporation. Crucially, the model does not stop at explicit connections. It also constructs what the authors call adaptive implicit semantic completion graphs, which surface latent associations between entities that are not directly linked in the recorded data. These implicit links act as scaffolding across the gaps left by structural sparsity, allowing information to flow between entities whose relationship is suggested by context rather than documented by an edge.</p>
<p>The final stage of the architecture fuses these multiple streams of information. Gating mechanisms, borrowed in spirit from recurrent sequence models, learn to weigh the contribution of each source dynamically rather than fixing their importance in advance. Temporal self-attention, a time-aware descendant of the attention mechanism popularized by the Transformer architecture, then integrates evidence across historical time steps, letting the model decide which moments in the past are most relevant to a query about the present or future. The combination means DIRA can, in effect, decide for itself whether a query is best answered by a recent structural pattern, a periodic signal, or a subtle semantic association that only appears when several time slices are considered together.</p>
<p>To test the approach, the researchers evaluated DIRA on four widely used real-world benchmark datasets drawn from event streams: ICEWS14, ICEWS18, ICEWS05-15, and GDELT. These datasets record political and news events with timestamps and relational annotations, making them standard proving grounds for temporal reasoning systems. Performance was measured using mean reciprocal rank, or MRR, and Hits at one, three, and ten, metrics that capture how highly a model ranks the correct answer among all candidates. Across all four datasets, DIRA consistently outperformed state-of-the-art baseline methods, and the gains were particularly meaningful in extrapolation reasoning, the hardest setting in which models must predict facts beyond the time range they were trained on.</p>
<p>The team did not rely on a single lucky run to support those claims. In a statistical appendix, the authors report that DIRA was executed with five independent random seeds, and its mean MRR was compared against the reported performance of the state-of-the-art baseline HisRES using a one-sample t-test at a significance level of 0.05. On every dataset, the mean difference was positive, the 95 percent confidence intervals excluded zero, and the one-tailed p-values fell below 0.001, indicating that the improvements are statistically significant rather than artifacts of initialization. Narrow confidence intervals further underscored the reliability of the advantage. The experiments were run under a fixed configuration, with a batch size of 1024, a historical window length of three, and an embedding dimension of 200, on a single NVIDIA RTX 3090 GPU, and the authors report that training completed within a reasonable time frame on all datasets, though a comprehensive efficiency comparison with baselines is deferred to future work.</p>
<p>The broader significance of the work lies in its demonstration that explicit and implicit modeling are complementary rather than competing strategies. Explicit dynamic embeddings give a model a principled account of how entities move through time, while implicit relation-aware mechanisms recover the semantic glue that sparse data leaves unstated. The authors argue that combining the two is a promising direction for temporal knowledge graph completion, particularly in sparse and extrapolation scenarios where neither strategy alone suffices. Because the study uses publicly available datasets and follows their usage policies, the results are reproducible by other groups, and the statistical rigor of the evaluation sets a standard that much of the field still lacks.</p>
<p>For applications, the implications reach well beyond the benchmarks. Temporal knowledge graphs underpin event forecasting systems used in political risk analysis, temporal question answering assistants, and recommendation engines that must respect the shelf life of facts. A model that can reason reliably over sparse, evolving data, and that can extrapolate beyond its training window with statistical confidence, brings those systems closer to practical dependability. The research was funded in part by the 2023 Humanities and Social Sciences Research and Planning Fund of the Ministry of Education of China under grant number 23YJA860010, and the authors, based at Dalian Minzu University, Dalian Neusoft University of Information, and the University of Maribor, declare no conflicts of interest. As knowledge graphs continue to grow as the factual backbone of search, dialogue, and decision-support systems, techniques like DIRA suggest that the next leap in machine reasoning may come not from bigger models, but from smarter ways of listening to time itself.</p>
<p><strong>Subject of Research:</strong> Temporal knowledge graph completion using dynamic embeddings and implicit relation-aware self-attention</p>
<p><strong>Article Title:</strong> Dynamic embedding and implicit relation-aware self-attention for TKGC</p>
<p><strong>Article References:</strong> Dynamic embedding and implicit relation-aware self-attention for TKGC. (n.d.). <a href="https://doi.org/10.1007/s10489-026-07481-x" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07481-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07481-x" rel="noopener noreferrer">10.1007/s10489-026-07481-x</a></p>
<p><strong>Keywords:</strong> temporal knowledge graph completion, dynamic embedding, self-attention, graph neural networks, knowledge graphs, extrapolation reasoning, event forecasting, machine learning, link prediction, temporal reasoning, Applied Intelligence, DIRA</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223618</post-id>	</item>
		<item>
		<title>Self-Attention Meets Parallel Memory Networks to Sharpen Text Sentiment Recognition</title>
		<link>https://scienmag.com/self-attention-meets-parallel-memory-networks-to-sharpen-text-sentiment-recognition/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 00:15:15 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced deep neural models for analyzing public opinion]]></category>
		<category><![CDATA[challenges in traditional neural networks for sentiment analysis]]></category>
		<category><![CDATA[combining LSTM with self-attention for improved sentiment detection]]></category>
		<category><![CDATA[data science]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Deep learning for sentiment analysis]]></category>
		<category><![CDATA[emotion detection]]></category>
		<category><![CDATA[emotional signal extraction from social media comments]]></category>
		<category><![CDATA[feature representation]]></category>
		<category><![CDATA[GoEmotions dataset]]></category>
		<category><![CDATA[handling noisy and high-dimensional language data]]></category>
		<category><![CDATA[innovative approaches to]]></category>
		<category><![CDATA[long short-term memory]]></category>
		<category><![CDATA[long short-term memory (LSTM) networks for emotion recognition]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[multimodal neural network architectures for sentiment analysis]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[parallel memory networks in text processing]]></category>
		<category><![CDATA[PLSTM-SA architecture for natural language understanding]]></category>
		<category><![CDATA[recurrent neural networks]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[self-attention mechanisms in NLP]]></category>
		<category><![CDATA[sentiment analysis]]></category>
		<category><![CDATA[text classification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=220194</guid>

					<description><![CDATA[Researchers have developed a parallel long short-term memory network with self-attention that reaches 91.6 percent accuracy in recognizing sentiment in unstructured online text.]]></description>
										<content:encoded><![CDATA[<p>Every minute, millions of comments, reviews, and opinions pour onto social media platforms, e-commerce sites, and discussion forums, forming one of the largest and messiest data streams humanity has ever produced. Hidden inside this torrent of unstructured text is an enormous amount of emotional signal: whether customers are delighted or furious, whether public mood is turning optimistic or anxious, and how people genuinely feel about products, services, and events. A new study published in the International Journal of Data Science and Analytics proposes a deep learning architecture designed to read that emotional signal more reliably than existing approaches, combining parallel long short-term memory networks with a self-attention mechanism into a single sentiment recognition framework the authors call PLSTM-SA.</p>
<p>The research, conducted by Sushadevi Shamrao Adagale of CSMU in Navi Mumbai, Shubhangi Vairagar of the Dr. D.Y. Patil Institute of Technology in Pimpri, and Praveen Gupta, also of CSMU, addresses a stubborn problem in natural language processing. Deep neural networks have transformed sentiment analysis over the past decade, but the authors argue that many current architectures still struggle in two critical ways. First, they handle high-dimensional feature spaces poorly, which becomes a serious liability when models must process the sprawling, noisy vocabulary of real-world online text. Second, and perhaps more fundamentally, many models treat all features as equally important, ignoring the fact that in a sentence like the film was long but never boring, a single word such as boring can flip the entire emotional meaning of the utterance.</p>
<p>The core of the new approach is the parallel long short-term memory network, or PLSTM. Long short-term memory networks, introduced in the late 1990s, are a specialized form of recurrent neural network built to capture dependencies that stretch across long sequences of data. They do this through internal gates, including forget gates, input gates, and output gates, that control what information is stored, discarded, and passed forward at each step of the sequence. This gating machinery allows LSTMs to remember context from many words back, which is essential for sentiment tasks where meaning often depends on phrases like not bad at all or I almost liked it, where negation and qualification reverse surface-level cues.</p>
<p>Where a standard LSTM processes a sequence through a single recurrent pathway, the parallel variant splits the workload across multiple LSTM branches that operate simultaneously on the input. The authors report that this parallel design enhances the model&#8217;s generalization capability and its feature representation, allowing the network to capture both short-term and long-term dependencies in textual features more effectively. Intuitively, running multiple memory pathways in parallel gives the model several complementary views of the same sentence: one branch may specialize in tracking immediate local context, while others preserve longer-range relationships between distant words. The outputs of these branches are then combined, producing a richer representation of the text than any single pathway could deliver alone.</p>
<p>Yet memory alone is not enough, and this is where the second half of the architecture comes in. The researchers pair the PLSTM with a self-attention mechanism, a technique popularized by the landmark 2017 paper Attention Is All You Need, which underlies modern transformer models. Self-attention allows a network to weigh the importance of every token in a sequence relative to every other token, dynamically deciding which words deserve the most focus when building a representation of the whole. In the PLSTM-SA framework, the self-attention layer is used to acquire information about emotional patterns in individual text tokens and to strengthen the correlation between local and global features, linking the fine-grained emotional charge of specific words to the broader sentiment of the entire passage.</p>
<p>This combination is designed to solve the equal-treatment problem directly. Instead of letting every feature contribute equally to the final classification, the attention mechanism assigns learned weights that amplify emotionally decisive tokens and suppress irrelevant ones. A word like excellent or terrible receives high attention, while filler words fade into the background. At the same time, because the attention operates on top of the parallel LSTM representations, it can capture relationships that span the whole sentence, not just neighboring words. The result, according to the authors, is a more context-aware framework that can also handle variable-length sequences, a practical necessity when real-world input ranges from a five-word tweet to a multi-paragraph product review.</p>
<p>To evaluate the model, the researchers trained and tested it on the GoEmotions dataset, an open-access corpus released by Google that contains tens of thousands of Reddit comments annotated with fine-grained emotion categories. The choice of dataset matters: GoEmotions reflects the informal, sarcastic, and grammatically loose language of genuine online conversation, which is far harder to classify than the clean sentences found in many academic benchmarks. The reported results show the PLSTM-SA architecture achieving an overall accuracy of 91.6 percent, with a recall of 0.8766, a precision of 0.8877, and an F1-score of 0.8811. The balance between precision and recall is notable, indicating that the model is neither excessively trigger-happy in labeling text as emotional nor overly conservative in missing genuine sentiment.</p>
<p>The study situates itself within a long lineage of sentiment analysis research, from early lexicon-based systems such as SentiWordNet and SenticNet, which scored words using curated affective dictionaries, through classical machine learning approaches built on support vector machines, to the deep learning era of convolutional and recurrent architectures. More recent work has explored hybrid designs, including convolutional neural networks paired with bidirectional LSTMs, BERT-based transformers combined with classical classifiers, and attention-enhanced deep networks applied to domains ranging from tweets about the Ukraine-Russia conflict to financial commentary and online food delivery reviews. The PLSTM-SA model draws on this accumulated evidence but argues that the parallel structure combined with self-attention offers a distinct advantage in generalization, particularly when the model must cope with unstructured, complex data that was never seen during training.</p>
<p>The practical implications extend well beyond academic benchmarks. Businesses use sentiment analysis to monitor brand reputation and customer satisfaction at scale, governments and public health agencies track collective mood during crises, and recommendation systems increasingly factor emotional response into what they surface. Models that misread sarcasm, negation, or mixed emotions translate directly into flawed business intelligence and misguided automated decisions. An architecture that better captures the interplay between local word-level emotion and global sentence-level meaning could make these downstream applications substantially more trustworthy, especially on the informal, rapidly evolving language of social platforms where traditional lexicons quickly go stale.</p>
<p>The authors acknowledge that developing a reliable sentiment analysis model remains challenging precisely because of the unstructured and complex nature of the data involved, and their contribution should be read as one step in an ongoing effort rather than a final solution. Still, the reported performance figures, combined with the architecture&#8217;s ability to handle variable-length sequences and its explicit mechanism for weighting emotional features, suggest that hybrid designs merging recurrent memory with attention will remain a productive direction for text-based affect recognition. As the volume of opinionated online text continues to grow, the systems that can genuinely understand how people feel, not just what they say, will become an increasingly essential part of the data science toolkit, and this study offers a concrete, measurable demonstration of how parallel memory and self-attention can work together toward that goal.</p>
<p><strong>Subject of Research:</strong> Deep learning architecture combining parallel LSTM networks and self-attention for text sentiment recognition</p>
<p><strong>Article Title:</strong> Text sentiment recognition using a self-attention-based parallel long short term memory</p>
<p><strong>Article References:</strong> Text sentiment recognition using a self-attention-based parallel long short term memory. (n.d.). <a href="https://doi.org/10.1007/s41060-026-01305-4" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01305-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01305-4" rel="noopener noreferrer">10.1007/s41060-026-01305-4</a></p>
<p><strong>Keywords:</strong> sentiment analysis, natural language processing, long short-term memory, self-attention, deep learning, machine learning, GoEmotions dataset, text classification, emotion detection, recurrent neural networks, feature representation, data science</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">220194</post-id>	</item>
		<item>
		<title>Hidden Switches: New Attack Plants Undetectable Backdoors in Vision Transformers</title>
		<link>https://scienmag.com/hidden-switches-new-attack-plants-undetectable-backdoors-in-vision-transformers/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 23:11:11 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial attacks on computer vision models]]></category>
		<category><![CDATA[adversarial machine learning]]></category>
		<category><![CDATA[AI safety]]></category>
		<category><![CDATA[backdoor attack]]></category>
		<category><![CDATA[backdoor attack in machine learning]]></category>
		<category><![CDATA[backdoor detection in vision transformers]]></category>
		<category><![CDATA[covert backdoor triggers in deep learning]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[cybersecurity threats in AI]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning model robustness]]></category>
		<category><![CDATA[hidden switches in neural networks]]></category>
		<category><![CDATA[high-stakes AI application vulnerabilities]]></category>
		<category><![CDATA[image classification]]></category>
		<category><![CDATA[machine learning security]]></category>
		<category><![CDATA[medical imaging model security]]></category>
		<category><![CDATA[model integrity in autonomous systems]]></category>
		<category><![CDATA[model poisoning]]></category>
		<category><![CDATA[prompt tuning]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[self-attention mechanism exploitation]]></category>
		<category><![CDATA[supply chain security]]></category>
		<category><![CDATA[vision transformer security vulnerabilities]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=219986</guid>

					<description><![CDATA[Researchers have demonstrated a switchable backdoor attack that injects dual tokens through the layers of Vision Transformers, achieving up to 99 percent attack success while leaving clean accuracy intact.]]></description>
										<content:encoded><![CDATA[<p>Vision Transformers have quietly become the backbone of modern computer vision, powering everything from image classifiers and object detectors to medical imaging pipelines and autonomous driving prototypes. Their self-attention mechanism, which lets the model weigh relationships between every patch of an image, has displaced convolutional networks in many high-stakes applications. But a new study from researchers at the University of Information Technology, Ho Chi Minh City, and Vietnam National University Ho Chi Minh City suggests that the very architecture making these models so powerful may also make them dangerously easy to compromise. In a paper published in the International Journal of Machine Learning and Cybernetics, Dung Minh Do and Khang Nguyen describe a backdoor attack that achieves attack success rates of up to 99 percent while leaving the model&#8217;s performance on clean, unmodified images essentially untouched.</p>
<p>Backdoor attacks are among the most insidious threats in machine learning security. Unlike adversarial examples, which exploit a model at inference time by perturbing individual inputs, a backdoor is baked into the model itself during training or fine-tuning. A model with a backdoor behaves perfectly normally on ordinary data, matching the accuracy and reliability its developers expect. But when a specific trigger, a subtle pattern chosen by the attacker, appears in an input, the model&#8217;s behavior flips, producing whatever output the attacker has designated. For a deployed system, this means an adversary could, for example, cause a traffic sign classifier to misread a stop sign or a security screening system to wave through a prohibited item, all without any visible degradation in everyday performance that might raise alarms.</p>
<p>What makes the new attack notable is where it operates. Rather than modifying the weights of the transformer or poisoning the training labels in conventional ways, the method injects two specially designed tokens into the model&#8217;s token stream, progressively working from the shallowest layers to the deepest ones. Vision Transformers process images by splitting them into patches, embedding each patch as a token, and then passing those tokens through a stack of transformer blocks in which self-attention layers let tokens exchange information. By inserting crafted tokens at multiple depths, the attack ensures that the malicious signal is reinforced and refined as it travels through the network, rather than being diluted or overwritten by the model&#8217;s normal processing.</p>
<p>The progressive, layer-by-layer nature of the injection is central to the attack&#8217;s stealthiness. Defenses that inspect a single layer, or that look for anomalous attention patterns at one point in the network, can miss a signal that is distributed across the entire depth of the model. Because the injected tokens interact with the image tokens through the standard attention mechanism, they can steer the model&#8217;s internal representations toward the attacker&#8217;s chosen target class only when the trigger is present. On clean inputs, the tokens remain effectively inert, which is why the authors report that the method maintains competitive clean accuracy even as it delivers near-perfect attack success rates across multiple visual datasets.</p>
<p>The word switchable in the study&#8217;s title points to another dimension of the threat. The attack draws on a growing line of research into switchable backdoors against pre-trained vision transformers, in which a single compromised model can carry multiple behaviors that an attacker can toggle. This means a defender who discovers one trigger and filters it out cannot assume the model is safe; another trigger may remain dormant, waiting to be activated. The authors position their work as an analysis and exploitation of ViT vulnerabilities, and the breadth of prior work they survey, from BadNets and weight-poisoning attacks to attention hijacking and prompt-based backdoors, underscores how rapidly this attack surface has expanded as transformers have spread through computer vision.</p>
<p>The connection to prompt tuning is particularly significant for the current state of the field. Prompt-based methods, in which small learnable tokens are prepended to a model&#8217;s input or inserted into its layers, have become a popular way to adapt large pre-trained models to new tasks without expensive full fine-tuning. Visual prompt tuning and related techniques are widely used because they are efficient and effective. But the same mechanism that makes prompts useful, namely the ability to inject learned tokens that influence the model&#8217;s attention and representations, is exactly what this attack weaponizes. A malicious actor with access to a fine-tuning pipeline, or able to distribute a poisoned adapter or checkpoint, could embed a backdoor that looks indistinguishable from a legitimate prompt-based adaptation.</p>
<p>The supply chain implications are sobering. Modern machine learning practice relies heavily on pre-trained models downloaded from public hubs, fine-tuned adapters shared between teams, and third-party datasets. Each of these channels is a potential delivery mechanism for a backdoor. Earlier research has shown that weight poisoning attacks on pre-trained models can survive downstream fine-tuning, and that backdoors can be hidden in ways that evade standard inspection. The new study adds to this picture by showing that the token-based machinery of vision transformers offers attackers a particularly clean injection point, one that does not require the crude modifications of earlier attacks that made them easier to detect.</p>
<p>The authors evaluated their method on multiple visual datasets, measuring not only attack success rate and clean accuracy but also stealthiness and robustness. The headline figure, an attack success rate of up to 99 percent, is alarming enough, but the more troubling result is the combination of that success with preserved clean performance, since it means conventional accuracy-based validation would reveal nothing amiss. The paper also situates the work against existing defenses, which include fine-pruning approaches that remove rarely activated neurons, neural attention distillation that tries to erase trigger-related attention patterns, and input-level detection methods that look for inconsistencies in a model&#8217;s predictions under image transformations. Because the new attack distributes its signal across shallow and deep layers through dual tokens, many of these defenses, which were designed with convolutional networks or single-point injections in mind, face a harder problem.</p>
<p>The researchers are explicit about the warning their findings carry for real-world deployment. Vision Transformers are increasingly used in settings where a silent failure mode could have serious consequences, including medical diagnosis support, surveillance, industrial inspection, and safety-critical perception systems. A backdoored model in any of these contexts could be remotely triggered by an input crafted to contain the attacker&#8217;s pattern, and the compromise would be invisible in routine testing. The authors make their code publicly available, which serves the defensive side of the field as well: reproducible attack implementations are essential for developing and benchmarking countermeasures, and the history of adversarial machine learning shows that security research advances fastest when attacks are fully documented.</p>
<p>For the broader community, the study is a reminder that architectural progress and security progress have been badly out of step. The references in the paper trace a decade of deep learning breakthroughs, from early convolutional networks through EfficientNet and the original Vision Transformer, alongside a parallel literature of backdoor learning surveys, prompt injection analyses, and defense proposals. Yet the authors note that security threats to ViTs, particularly backdoor attacks, have not received research attention commensurate with the architecture&#8217;s adoption. Closing that gap will likely require defenses designed specifically for token-based architectures: methods that audit injected tokens, verify the provenance of fine-tuned checkpoints, test models against a family of triggers rather than a single known pattern, and treat the entire depth of the network, not just its input layer, as a potential attack surface. Until such defenses mature, the near-perfect stealth and effectiveness demonstrated by progressive dual-token injection stands as a stark warning that the models powering tomorrow&#8217;s vision systems may harbor switches that only their attackers know how to flip.</p>
<p><strong>Subject of Research:</strong> Backdoor attack vulnerabilities in Vision Transformers via progressive dual-token injection</p>
<p><strong>Article Title:</strong> Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers</p>
<p><strong>Article References:</strong> Do, D. M., &amp; Nguyen, K. (2026). Switchable backdoor attack in vision transformers via progressive dual-token injection from shallow to deep layers. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 489. <a href="https://doi.org/10.1007/s13042-026-03326-8" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03326-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03326-8" rel="noopener noreferrer">10.1007/s13042-026-03326-8</a></p>
<p><strong>Keywords:</strong> Vision Transformers, backdoor attack, machine learning security, self-attention, prompt tuning, adversarial machine learning, model poisoning, supply chain security, deep learning, cybersecurity, image classification, AI safety</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">219986</post-id>	</item>
		<item>
		<title>Attention Explained: A sweeping survey maps the engine behind modern AI</title>
		<link>https://scienmag.com/attention-explained-a-sweeping-survey-maps-the-engine-behind-modern-ai/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 23:40:55 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[applications of attention in medical imaging]]></category>
		<category><![CDATA[attention in natural language processing]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[Attention mechanism in artificial intelligence]]></category>
		<category><![CDATA[BERT]]></category>
		<category><![CDATA[comprehensive survey of attention mechanisms]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[evolution of attention-based models]]></category>
		<category><![CDATA[Flash Attention]]></category>
		<category><![CDATA[GPT]]></category>
		<category><![CDATA[importance of dynamic weighting in neural networks]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[mapping of attention families]]></category>
		<category><![CDATA[mathematical foundations of attention]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[neural network importance scoring]]></category>
		<category><![CDATA[role of attention in AI systems]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[significance of attention in modern AI]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[Transformer architecture variants]]></category>
		<category><![CDATA[vision transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211286</guid>

					<description><![CDATA[A new survey in Machine Learning provides an extensive taxonomy of attention mechanisms and more than thirty Transformer variants that underpin modern artificial intelligence.]]></description>
										<content:encoded><![CDATA[<p>A single mathematical idea has quietly become the beating heart of nearly every transformative artificial intelligence system of the past decade: the attention mechanism. From the chatbots that draft your emails to the vision models that flag tumors in medical scans, attention decides, moment by moment, which pieces of information matter most. Now, a comprehensive survey published in the journal Machine Learning by Farhad Mortezapour Shiri of Universiti Putra Malaysia, together with Fateme Memar of the University of Kansas and Maryam Parhizgar of Islamic Azad University, offers one of the most complete maps yet of this computational landscape, cataloging fourteen distinct families of attention and more than thirty Transformer variants that together define the state of the art.</p>
<p>The core insight behind attention is deceptively simple. Instead of forcing a neural network to process every element of an input with equal weight, attention lets the model assign importance scores dynamically, amplifying the parts of a sentence, image, or time series that are most relevant to the current task and dampening the rest. In its standard formulation, each input is projected into three vectors called queries, keys, and values. The similarity between a query and each key produces a set of weights, typically normalized through a softmax function, and the model outputs a weighted sum of the values. This means a language model deciding what word to predict next can look back across an entire sentence, or an entire document, and pull together exactly the context it needs.</p>
<p>The survey traces the mechanism&#8217;s origins to neural machine translation, where researchers led by Dzmitry Bahdanau showed in 2014 that letting a decoder peek back at the most relevant source words, rather than squeezing an entire sentence into a single fixed-length vector, dramatically improved translation quality. That early alignment-based attention has since blossomed into a sprawling taxonomy. The authors organize the field into hierarchical attention, which operates at multiple levels of granularity; bidirectional attention, which fuses information flowing in both directions; and multi-head attention, the workhorse of the original Transformer, which runs several attention operations in parallel so the model can simultaneously track different kinds of relationships, such as syntax in one head and long-range semantic links in another.</p>
<p>Beyond those foundations, the survey catalogs a generation of efficiency-driven refinements. Multi-query and grouped-query attention reduce the number of key and value heads shared across queries, slashing the memory bandwidth required during the fast autoregressive decoding that powers large language models. Graph attention extends the mechanism to irregular, network-structured data. Channel attention, exemplified by the influential squeeze-and-excitation paradigm, learns which feature channels of a convolutional network to emphasize, while spatial attention highlights which regions of an image deserve focus; combining the two yields the channel-spatial hybrids now common in computer vision, remote sensing, and medical imaging. Temporal and spatial-temporal variants bring the same selectivity to video, sensor streams, and time-series forecasting, and cross attention lets one modality interrogate another, forming the connective tissue of image captioning and text-to-image generation. Axial attention factors a full two-dimensional attention map into separate row and column passes, taming the quadratic cost of images, while Flash Attention, introduced by Tri Dao and colleagues, recomputes attention in memory-efficient tiles that stay close to the processor&#8217;s fast on-chip memory, delivering exact results at a fraction of the usual input-output cost.</p>
<p>All of these innovations converge in the Transformer, the architecture introduced in 2017 under the slogan that attention is all you need. By dispensing with recurrence entirely and stacking layers of multi-head self-attention and feedforward networks, the Transformer made it possible to train on entire sequences in parallel, unlocking the scale that defines modern AI. The survey devotes extensive attention to the architecture&#8217;s evolution. Encoder-only models such as BERT learn deep bidirectional language representations by masking words and predicting them from context, and they have seeded domain-specific descendants for biomedical text, finance, climate science, and electronic health records. Decoder-only families, including the GPT series and the openly released LLaMA models, generate text autoregressively and now anchor the large language model boom. Encoder-decoder systems such as BART and the text-to-text framework T5 unify translation, summarization, and comprehension within a single sequence-to-sequence mold.</p>
<p>Handling long contexts remains one of the field&#8217;s defining challenges, because naive self-attention scales quadratically with sequence length. The survey details a rich arsenal of responses. Transformer-XL introduces a recurrence mechanism that caches hidden states across text segments, extending context far beyond a fixed window, while XLNet rethinks autoregressive pretraining with permutation-based objectives. Longformer and BIGBIRD employ sparse, local-plus-global attention patterns to process documents thousands of tokens long. Reformer buckets similar tokens together using locality-sensitive hashing, Linforcer-style approximations appear in the Linformer&#8217;s low-rank projections of the attention matrix, and the Performer replaces the softmax kernel entirely with random feature maps that make the computation linear in sequence length. Positional information itself has been reengineered, with the rotary embeddings of RoFormer and the linear biases of ALiBi letting models extrapolate to lengths never seen during training. At the extreme end, the Switch Transformer pairs attention with sparse mixture-of-experts routing, activating only a fraction of a trillion-parameter model for each token.</p>
<p>Perhaps the most striking chapter of the story is attention&#8217;s conquest of computer vision. The Vision Transformer, or ViT, chops an image into fixed-size patches, embeds them like words, and feeds them through a standard Transformer encoder, matching or beating state-of-the-art convolutional networks when trained at scale. A rapidly expanding family tree followed: DeiT showed that distillation through attention makes such models trainable on modest data; DeepViT and T2T ViT refined how tokens are constructed and deepened the stack; CrossViT fused information across multiple patch scales with cross attention; the Pyramid Vision Transformer adapted the architecture for dense prediction tasks such as segmentation and detection; and the Swin Transformer introduced shifted local windows that give vision models a hierarchical, convolution-like multi-scale structure. DETR reframed object detection as a set prediction problem solved with an encoder-decoder and object queries, while MViT and ViViT extended the recipe to video, and the Deformable Attention Transformer taught attention to sample only the most informative spatial locations.</p>
<p>The survey also points toward the frontier of neuromorphic computing, where Spiking Transformers merge attention with spiking neural networks that communicate through discrete, event-driven pulses. By replacing costly floating-point multiply-accumulate operations with sparse spike accumulation, architectures such as Spikformer, the spike-driven transformer family, and related spiking vision models promise dramatically lower energy consumption, an increasingly critical consideration as the computational appetite of large models collides with hardware and environmental limits.</p>
<p>What emerges from the authors&#8217; comparative analysis is a field in vigorous, creative flux, unified by one principle. Whether the domain is natural language processing, computer vision, recommender systems, speech recognition, weather forecasting, or sensor data analysis, the ability to selectively weigh what matters has proven to be the common denominator of success. The survey&#8217;s contribution is not a single breakthrough but a synthesis: by laying out the general framework of attention, the trade-offs among its many variants, and the strengths and limitations of each Transformer descendant, it gives researchers and practitioners a coherent map of how a decade of scattered innovations fits together, and where the next ones are likely to come from. As the authors emphasize, attention-based models, and the Transformer architecture above all, have not merely contributed to modern deep learning; they have reshaped it, and this comprehensive account makes clear that the reshaping is far from over.</p>
<p><strong>Subject of Research:</strong> Attention mechanisms and Transformer architectures in deep learning</p>
<p><strong>Article Title:</strong> What is Attention Mechanism? A Comprehensive Survey of Attention Methods and Transformer Models</p>
<p><strong>Article References:</strong> What is Attention Mechanism? A Comprehensive Survey of Attention Methods and Transformer Models. (n.d.). <a href="https://doi.org/10.1007/s10994-026-07131-w" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07131-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07131-w" rel="noopener noreferrer">10.1007/s10994-026-07131-w</a></p>
<p><strong>Keywords:</strong> attention mechanism, Transformer, deep learning, machine learning, BERT, GPT, vision transformer, self-attention, Flash Attention, large language models, computer vision, natural language processing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211286</post-id>	</item>
		<item>
		<title>New Self-Attention Method Strips Redundant Filters From CNNs to Boost Plant Disease Detection</title>
		<link>https://scienmag.com/new-self-attention-method-strips-redundant-filters-from-cnns-to-boost-plant-disease-detection/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 23:54:26 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced techniques for CNN filter pruning]]></category>
		<category><![CDATA[agricultural AI]]></category>
		<category><![CDATA[CNN-based plant leaf disease analysis]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[convolutional neural networks for agriculture]]></category>
		<category><![CDATA[cosine similarity]]></category>
		<category><![CDATA[cosine similarity in feature detection]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning model optimization in agriculture]]></category>
		<category><![CDATA[feature diversity]]></category>
		<category><![CDATA[filter redundancy]]></category>
		<category><![CDATA[filter redundancy in deep learning]]></category>
		<category><![CDATA[improving CNN efficiency for plant disease classification]]></category>
		<category><![CDATA[model efficiency]]></category>
		<category><![CDATA[neural network pruning]]></category>
		<category><![CDATA[plant disease classification]]></category>
		<category><![CDATA[plant disease detection]]></category>
		<category><![CDATA[Plant Pathology 2020]]></category>
		<category><![CDATA[real-time plant disease diagnosis]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[self-attention in CNNs]]></category>
		<category><![CDATA[self-attention mechanisms in convolutional neural networks]]></category>
		<category><![CDATA[suppression of duplicate feature detectors]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=208903</guid>

					<description><![CDATA[Researchers have developed CNN-SA-RFR, a cosine similarity-based self-attention framework that suppresses redundant filters during CNN training, boosting plant disease classification accuracy on two benchmark datasets.]]></description>
										<content:encoded><![CDATA[<p>Deep learning has transformed the way computers interpret images, and nowhere is that transformation more consequential than in agriculture, where a fast and accurate diagnosis of diseased leaves can mean the difference between a saved harvest and a lost season. Convolutional neural networks, the workhorses of modern computer vision, have repeatedly demonstrated outstanding performance in plant disease classification. Yet beneath their impressive accuracy lies a persistent and costly inefficiency: as these networks train, many of their learned filters end up detecting essentially the same features, a phenomenon known as filter redundancy. A new study published in Multimedia Tools and Applications tackles this problem head-on with a framework called CNN-SA-RFR, which uses cosine similarity-based self-attention to identify and suppress duplicate feature detectors while a model is still learning.</p>
<p>The research, conducted by Saloua Lagnaoui and Khalid Haddouch of the Laboratory of Applied Sciences and Emerging Technologies at ENSA, Sidi Mohamed Ben Abdellah University in Fez, Morocco, together with Zakariae En-naimani of the Laboratory Computer Science, Artificial Intelligence and Cyber Security at ENSET, Hassan II University of Casablanca, introduces a fundamentally different way of thinking about redundancy. Rather than waiting until training is complete and then surgically removing filters, or randomly resetting them in the hope that diversity will emerge, the proposed method models the relationships between filters dynamically during training itself. It continuously measures how similar the responses of different filters are and applies similarity-aware weighting that dampens redundant feature responses before they can dominate the network&#8217;s internal representations.</p>
<p>To understand why this matters, it helps to consider what filters actually do inside a convolutional neural network. Each filter is a small pattern detector that slides across an image, responding strongly to particular visual features such as edges, textures, color gradients, or the characteristic lesions and discolorations that signal disease on a leaf. In an ideal network, every filter would specialize in a distinct feature, collectively forming a rich and diverse vocabulary of visual descriptors. In practice, however, gradient-based training frequently drives multiple filters toward nearly identical behavior. When that happens, the network wastes computational capacity re-learning the same feature, its effective representational power shrinks, and its convergence during training can become unstable as redundant filters compete for the same gradient signals.</p>
<p>Previous approaches to this problem have generally fallen into two camps. Pruning techniques, such as those that rank filters by importance and remove the least useful ones after or during training, treat redundancy as something to be excised. Stochastic resetting methods, which randomly reinitialize filters mid-training, gamble that randomness will restore diversity. Both strategies have drawbacks: pruning can be brittle and often requires careful tuning and retraining, while stochastic approaches introduce nondeterminism that makes results harder to reproduce and optimize. The CNN-SA-RFR framework, by contrast, is deterministic and adaptive. By embedding a cosine similarity measure into a self-attention mechanism, it quantifies the angular alignment between filter responses and uses that information to reweight the network&#8217;s feature maps continuously, suppressing responses that duplicate what other filters already capture.</p>
<p>The choice of cosine similarity is technically significant. Unlike Euclidean distance, which is sensitive to the magnitude of activation vectors, cosine similarity measures the orientation between vectors, making it a natural gauge of whether two filters are responding to the same underlying pattern regardless of response strength. Combined with self-attention, a mechanism popularized by the transformer architecture that allows a model to weigh the relationships among different elements of its input, this creates a layer that is acutely aware of its own internal redundancy. The attention weights effectively act as a soft, learned form of redundancy control: filters whose responses are highly similar to others receive diminished influence, while distinctive, informative features are amplified and allowed to propagate through the network.</p>
<p>The authors evaluated their framework on two publicly available benchmark datasets: Plant Pathology 2020, released for the FGVC7 Kaggle competition, and the Plant Disease Recognition dataset, also hosted on Kaggle. These datasets present genuinely difficult classification challenges, requiring models to distinguish healthy leaves from those affected by multiple diseases under varying lighting conditions, viewing angles, and stages of disease progression. The results were striking. The proposed method achieved up to 78 percent accuracy on the Plant Pathology 2020 dataset and 94 percent on the Plant Disease Recognition dataset, outperforming baseline convolutional neural network models on both benchmarks.</p>
<p>Beyond raw accuracy, the study reports improvements in three additional dimensions that matter greatly for practical deployment. Feature diversity increased, indicating that the network&#8217;s filters genuinely specialized rather than duplicating one another, which suggests the similarity-aware weighting is doing exactly what it was designed to do. Convergence stability also improved, meaning training runs were less prone to the oscillations and plateaus that plague redundant networks, a benefit that translates directly into reduced engineering time and compute cost. Together, these gains position the method as a practical optimization strategy rather than a laboratory curiosity, particularly for agricultural applications where models may need to run on edge devices with limited processing power.</p>
<p>The work builds on a growing body of research at the intersection of attention mechanisms and efficient convolutional networks. Self-attention has already proven its value in plant disease recognition, with prior studies showing that attention-enhanced convolutional models can better localize disease symptoms and ignore background clutter. At the same time, research into kernel and filter redundancy reduction has explored low-rank expansions, knowledge distillation, and differentiable pruning masks. What distinguishes CNN-SA-RFR is the fusion of these two threads: it uses attention not to improve the network&#8217;s view of the input image, but to improve the network&#8217;s view of itself, continuously auditing its own filters for redundancy and correcting it in real time.</p>
<p>The implications extend well beyond plant pathology. Filter redundancy is a general problem in convolutional architecture design, affecting applications from medical imaging to autonomous driving, and any deterministic, training-time method that mitigates it without sacrificing accuracy is of broad interest to the machine learning community. For agriculture specifically, the stakes are high. Plant diseases cause substantial global crop losses each year, and smartphone-based diagnostic tools powered by efficient CNNs are increasingly seen as a frontline defense for farmers who lack access to expert agronomists. Leaner, more diverse, and more stable networks make such tools cheaper to train, faster to run, and more reliable in the field.</p>
<p>The Moroccan team&#8217;s findings also carry a methodological lesson for the wider deep learning community: architecture and training procedure need not be designed independently. By making redundancy awareness an intrinsic property of the network layer rather than an external post-processing step, CNN-SA-RFR suggests a future in which models are not merely trained and then optimized, but are self-optimizing from the first forward pass. The datasets used in the study are publicly available, and the authors note that the code can be obtained from the corresponding author, opening the door for other researchers to replicate, refine, and extend the approach across new domains and larger architectures.</p>
<p><strong>Subject of Research:</strong> A cosine similarity-based self-attention framework for reducing filter redundancy in convolutional neural networks applied to plant disease classification</p>
<p><strong>Article Title:</strong> CNN-SA-RFR: A cosine similarity-based self-attention approach for reducing filter redundancy in CNNs for plant disease classification</p>
<p><strong>Article References:</strong> Lagnaoui, S., En-naimani, Z., &amp; Haddouch, K. (2026). CNN-SA-RFR: A cosine similarity-based self-attention approach for reducing filter redundancy in CNNs for plant disease classification. <em>Multimedia Tools and Applications, 85</em>(10), Article 764. <a href="https://doi.org/10.1007/s11042-026-21921-3" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21921-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21921-3" rel="noopener noreferrer">10.1007/s11042-026-21921-3</a></p>
<p><strong>Keywords:</strong> convolutional neural networks, self-attention, filter redundancy, cosine similarity, plant disease classification, deep learning, computer vision, agricultural AI, model efficiency, feature diversity, Plant Pathology 2020, neural network pruning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">208903</post-id>	</item>
		<item>
		<title>New Multi-Branch AI Model Predicts Winter Wheat Yields Weeks Before Harvest, Even Under Extreme Weather</title>
		<link>https://scienmag.com/new-multi-branch-ai-model-predicts-winter-wheat-yields-weeks-before-harvest-even-under-extreme-weather/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:50:26 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[agricultural forecasting under climate change]]></category>
		<category><![CDATA[AI in food security]]></category>
		<category><![CDATA[climate-resilient crop modeling]]></category>
		<category><![CDATA[convolutional neural network]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning crop forecasting]]></category>
		<category><![CDATA[drought and frost stress prediction]]></category>
		<category><![CDATA[early harvest yield prediction]]></category>
		<category><![CDATA[extreme climate events]]></category>
		<category><![CDATA[extreme climate indices]]></category>
		<category><![CDATA[extreme weather impact on wheat]]></category>
		<category><![CDATA[Food security]]></category>
		<category><![CDATA[genetic algorithm]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[MBF-HybridNet model]]></category>
		<category><![CDATA[nonlinear crop response modeling]]></category>
		<category><![CDATA[remote sensing]]></category>
		<category><![CDATA[satellite data for agriculture]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[soil and weather data integration]]></category>
		<category><![CDATA[Thiessen polygons]]></category>
		<category><![CDATA[winter wheat]]></category>
		<category><![CDATA[winter wheat yield prediction]]></category>
		<category><![CDATA[yield prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204132</guid>

					<description><![CDATA[A new multi-branch deep learning model fusing satellite, weather, soil and extreme climate data predicts winter wheat yields in China with high accuracy about 30 days before harvest, even in extreme climate years.]]></description>
										<content:encoded><![CDATA[<p>As heat waves, frosts and droughts increasingly batter the world&#8217;s wheat fields, a team of researchers in China has unveiled a deep learning model that can predict winter wheat yields with remarkable accuracy—and do so roughly a month before harvest, even in years dominated by extreme climate events. The model, called MBF-HybridNet, was developed and tested across six counties in Qingdao City, a major agricultural region on China&#8217;s eastern coast, and its results suggest that fusing satellite data, weather records, soil properties and explicit extreme-climate indices can push crop forecasting to a new level of precision.</p>
<p>The stakes could hardly be higher. Wheat is a cornerstone of China&#8217;s national food security, yet escalating global warming has made extreme climate events considerably more frequent and severe, threatening the stability of production. Traditional tools for forecasting yields have struggled in exactly those years when forecasts matter most. Mechanistic crop models, which simulate plant growth using genotype parameters, weather and soil data, often perform poorly under extreme weather because of heavy data demands and structural rigidity. Statistical models, meanwhile, typically assume linear relationships between vegetation indices and yield, blinding them to the nonlinear dynamics that govern crop responses to stress.</p>
<p>Machine learning approaches such as Random Forest, Support Vector Regression and neural networks have improved on these limits, but they frequently fail to extract spatiotemporal dynamics from complex datasets, and most existing yield studies focus on late growth stages, delaying predictions until it is nearly too late to act. Deep learning architectures have begun to close this gap. Convolutional neural networks excel at extracting spatial features, while Long Short-Term Memory networks, with their memory cells and gating mechanisms, are adept at modeling the time-series character of crop growth. Earlier hybrid CNN-LSTM models outperformed either architecture alone, but they relied mainly on one-dimensional convolutions, ignored static variables such as soil, and—critically—did not account for extreme climate events at all, introducing systematic bias into their predictions.</p>
<p>MBF-HybridNet addresses all three shortcomings at once. The model adopts a multi-branch parallel architecture with three distinct modules: a Dynamic Variables Module that processes daily remote sensing and meteorological data, a Dynamic ECIs Module that handles yearly-scale extreme climate index data, and a Static Variables Module that ingests soil properties. The dynamic module stacks three two-dimensional convolutional layers with asymmetric 1 by 2 kernels, each followed by batch normalization and ReLU activation, and inserts a Self-Attention mechanism after the first two layers to focus on the most informative features. The extracted features are then reshaped into sequences and fed into a two-layer LSTM network with 256 units per layer and dropout to curb overfitting. The static module processes soil organic carbon, cation exchange capacity, pH, sand and clay content through its own convolutional stream. A staged fusion strategy then concatenates these streams through successive fully connected layers to produce the final yield estimate.</p>
<p>A key innovation lies in how the data are prepared. Rather than relying on county-level averages, the team used the Thiessen Polygon method to partition the winter wheat planting area into 288 uniform analysis cells centered on evenly distributed sampling points, preserving fine-scale environmental information while avoiding contamination from non-planting surfaces such as mountains and water bodies. The data were then organized as pseudo-2D image tensors, with rows representing time steps, columns representing feature variables and a third dimension representing the aggregated subregions. The growing season was divided into three progressively cumulative windows: the vegetative growth phase from sowing to pre-jointing, the vegetative-reproductive phase spanning jointing to heading, and the reproductive phase from heading to maturity.</p>
<p>To capture the fingerprint of extreme weather, the researchers computed nine extreme climate indices covering heat, frost and precipitation extremes—hot days, heat stress intensity, consecutive hot days, frost days, cold stress intensity, consecutive cold days, heavy precipitation days, consecutive wet days and consecutive dry days—calculated for each growth stage. Because feeding all 21 growth-stage indices into the model risked dimensionality and overfitting, a genetic algorithm was deployed to select the most informative subset for each stage: three indices for the vegetative phase, five for the vegetative-reproductive window, and eleven for the full season.</p>
<p>The performance gains were substantial. Validated with leave-one-year-out cross-validation across 2004 to 2019, MBF-HybridNet achieved R-squared values of 0.756 to 0.765 across the three cumulative growth stages, with mean absolute percentage errors around 4.2 percent, compared to the baseline LSTM&#8217;s R-squared values of 0.652 to 0.671 and errors approaching 4.9 percent. Relative to the baseline, the new model cut root mean square error by up to 55.67 kilograms per hectare and mean absolute error by up to 48.12 kilograms per hectare. An ablation study confirmed that the gains arise from the complementary contributions of spatial representation, self-attention and temporal modeling rather than any single component: CNN alone reached an R-squared of 0.646, adding self-attention lifted it to 0.674, and the full CNN-SA-LSTM stack reached 0.763.</p>
<p>The extreme climate indices proved especially valuable in anomalous years. When the study years were split into normal and extreme groups—with 2006, 2013, 2014 and 2019 flagged as extreme—the model without the indices showed visibly degraded accuracy in extreme years, with R-squared values dropping to around 0.72 to 0.74. Adding all indices raised extreme-year R-squared values to as high as 0.799, and the genetically optimized subsets performed even better, reaching 0.803 in the vegetative-reproductive window while reducing computational cost. Across the full record, the GA-based models improved R-squared by 1.6 to 2.1 percentage points over the index-free model. SHAP interpretability analysis revealed a clear phenological pattern: pre-flowering low-temperature events such as frost days and cold stress intensity dominated early stages, while post-flowering heat and water stress—consecutive hot days, heavy precipitation days and consecutive dry days—took over as the key drivers during grain filling. Wind speed, temperature, precipitation and vegetation indices, particularly solar-induced chlorophyll fluorescence, rounded out the most influential predictors.</p>
<p>Perhaps the most striking result is the model&#8217;s early-warning capability. Prediction accuracy improved as seasonal information accumulated but plateaued at the vegetative-reproductive stage, meaning that winter wheat yields can be reasonably estimated approximately 30 days before harvest. The model also corrected a persistent weakness of simpler networks: the tendency to overestimate low yields and underestimate high ones, a bias rooted in the imbalanced distribution of yield samples concentrated in the 5000 to 7000 kilograms per hectare range. Residual analysis showed most county-level errors stayed within plus or minus 400 kilograms per hectare, with the strongest performance in medium- and high-yield areas.</p>
<p>The researchers caution that the framework, built and tested in Qingdao&#8217;s temperate monsoon climate, would need regional recalibration elsewhere: extreme-climate thresholds should be adjusted for arid or subtropical zones, growth-stage windows redefined by local phenology, and topographic factors such as elevation and slope added for hillier terrain. Management variables—irrigation, fertilization and cultivar choice—were not explicitly modeled and may explain some residual uncertainty. Still, because every input is drawn from public datasets, the model&#8217;s architecture offers a transferable blueprint. As climate extremes intensify, tools like MBF-HybridNet could give farmers and policymakers the lead time they need to protect harvests before the damage is done.</p>
<p><strong>Subject of Research:</strong> A multi-branch fusion deep learning model that integrates remote sensing, meteorological, soil and extreme climate index data to estimate winter wheat yields under extreme climate events in Qingdao, China.</p>
<p><strong>Article Title:</strong> Develop a multi-branch fusion deep learning model to estimate the winter wheat yields under extreme climate events</p>
<p><strong>Article References:</strong> Jiang, X., Kong, D., Zhang, J., Zhang, S., Ma, Z., Yu, L., Yang, S., Bai, Y., Ali, S., &amp; Ullah, H. (2026). Develop a multi-branch fusion deep learning model to estimate the winter wheat yields under extreme climate events. <em>Artificial Intelligence in Agriculture</em>. <a href="https://doi.org/10.1016/j.aiia.2026.09.001" rel="noopener noreferrer">https://doi.org/10.1016/j.aiia.2026.09.001</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.aiia.2026.09.001" rel="noopener noreferrer">10.1016/j.aiia.2026.09.001</a></p>
<p><strong>Keywords:</strong> winter wheat, yield prediction, deep learning, extreme climate events, remote sensing, LSTM, convolutional neural network, self-attention, extreme climate indices, Thiessen polygons, food security, genetic algorithm</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204132</post-id>	</item>
		<item>
		<title>AI Model Predicts How We Remember Emotional Experiences Using Brain Signals</title>
		<link>https://scienmag.com/ai-model-predicts-how-we-remember-emotional-experiences-using-brain-signals/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:13:44 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[affective computing]]></category>
		<category><![CDATA[AI prediction of emotional experiences]]></category>
		<category><![CDATA[brain activity and memory evaluation]]></category>
		<category><![CDATA[brain signal analysis for emotional memory]]></category>
		<category><![CDATA[Brain-Computer Interface]]></category>
		<category><![CDATA[clinical applications of emotion prediction]]></category>
		<category><![CDATA[Cognitive Computation]]></category>
		<category><![CDATA[cognitive psychology of emotional memories]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for emotion recognition]]></category>
		<category><![CDATA[EEG]]></category>
		<category><![CDATA[EEG-based affective computing]]></category>
		<category><![CDATA[emotion recognition]]></category>
		<category><![CDATA[emotional memory prediction using brain signals]]></category>
		<category><![CDATA[event-related potentials]]></category>
		<category><![CDATA[impact of emotional peaks and endings]]></category>
		<category><![CDATA[LPP]]></category>
		<category><![CDATA[neural correlates of emotional recall]]></category>
		<category><![CDATA[peak-end effect]]></category>
		<category><![CDATA[peak-end rule in emotional experience]]></category>
		<category><![CDATA[retrospective emotional assessment]]></category>
		<category><![CDATA[retrospective evaluation]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[temporal convolutional network]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202672</guid>

					<description><![CDATA[A new deep learning system embeds the peak-end rule directly into its architecture, predicting how people will retrospectively judge emotional experiences from their EEG brain signals with unprecedented accuracy.]]></description>
										<content:encoded><![CDATA[<p>When people look back on an emotional experience, they rarely judge it by its full duration. Instead, decades of research in cognitive psychology have shown that our overall memory of an event is dominated by two moments: the emotional peak and the way it ends. This phenomenon, known as the peak-end rule, was famously articulated by Daniel Kahneman and colleagues through pain perception experiments, and it has since been confirmed across consumer experiences, everyday well-being, and clinical symptom reporting. A recent study published in the journal Cognitive Computation has now taken this behavioral insight and translated it directly into the architecture of a deep learning system, building an artificial intelligence model that can predict how a person will retrospectively evaluate an emotional experience from their brain activity alone.</p>
<p>The research, conducted by Zhongtang Guo of Duke Kunshan University, addresses a long-standing gap in EEG-based affective computing. Most existing studies of brain-driven emotion recognition have focused on classifying a person&#8217;s immediate emotional state at a given moment. Far less attention has been paid to predicting the summary judgment a person forms after an entire emotional sequence has unfolded. Yet it is precisely this retrospective evaluation that shapes clinical symptom recall, customer satisfaction, and everyday judgments of happiness or distress. The new work proposes that the neural signatures of the peak-end effect can be captured in event-related potentials, or ERPs, the millisecond-scale electrical ripples the brain produces in response to each stimulus.</p>
<p>To gather the necessary data, thirty right-handed participants aged 18 to 28 completed a sequential emotion induction paradigm built from the International Affective Picture System. Each trial presented six images in a row, arranged along a preset arousal curve so that each sequence contained a clear emotional peak and a defined endpoint. The researchers systematically manipulated where the peak occurred within the sequence and how intense the final image was, using a Latin square design to cross these variables fully. After each sequence, participants rated their overall experience on a nine-point scale, and these ratings became the training targets for the predictive model. Brain activity was recorded from a 64-channel EEG cap at a sampling rate of 1000 Hz, with rigorous preprocessing including independent component analysis to remove ocular, muscular, and cardiac artifacts.</p>
<p>Three ERP components formed the neural backbone of the system. The Early Posterior Negativity, appearing roughly 150 to 350 milliseconds after stimulus onset, indexes early attentional capture by emotionally salient material. The P300, measured between 300 and 500 milliseconds over central-parietal electrodes, reflects cognitive evaluation and stimulus categorization. The Late Positive Potential, or LPP, spanning roughly 400 to 800 milliseconds over midline parietal sites, is widely regarded as the core electrophysiological marker of emotional arousal and motivational salience. The study confirmed that all three components were significantly modulated by arousal intensity, with high-arousal negative images producing the largest LPP amplitudes, and the peak of each sequence was operationally defined as the image evoking the maximum absolute LPP amplitude.</p>
<p>On top of these neural features, the researchers built a hybrid deep learning architecture called TCN-Attention-PeakEnd Gate, abbreviated TAPE. The model begins with a temporal convolutional network, or TCN, whose dilated causal convolutions expand the receptive field exponentially across layers, allowing the network to capture long-range temporal dependencies without the vanishing-gradient problems of recurrent networks. The TCN output then passes through an eight-head multi-head self-attention module, which models global relationships among all time steps in the sequence. The distinctive element, however, is the peak-end gating module. This module takes the ERP feature vectors of the peak and endpoint moments and computes a learnable sigmoid gating signal that adaptively reweights the contribution of each time step. In effect, the peak-end rule is embedded into the network as a differentiable operation, constraining the model to mirror the cognitive bias that human memory actually exhibits.</p>
<p>The performance results were striking. Under leave-one-subject-out cross-validation, the most demanding evaluation scheme in EEG research, in which each participant serves as the test subject exactly once, TAPE achieved a mean absolute error of 1.038 rating points, a Pearson correlation of 0.654 between predicted and actual retrospective ratings, and a three-level classification accuracy of 70.4 percent. These figures significantly outperformed eight baselines spanning shallow regression, classical EEG feature engineering, and deep architectures including LSTM, CNN-LSTM, and Transformer models, with all differences confirmed by Holm-Bonferroni-corrected paired-sample t-tests. Across fifty independent fivefold cross-validation evaluations, TAPE simultaneously attained the highest median correlation of 0.682 and the smallest interquartile range of 0.054, indicating that its advantage was stable rather than an artifact of favorable data splits.</p>
<p>External validation on the publicly available SEED dataset replicated the model&#8217;s superiority, where TAPE reached 84.7 percent three-class accuracy compared with 81.9 percent for the strongest deep-learning baseline, a Transformer. An ablation study then dissected the contribution of each module. Removing the peak-end gating module increased error by roughly 11 percent, removing self-attention reduced the correlation by nearly 10 percent, and replacing the TCN encoder with an LSTM cost 6.3 percentage points of accuracy, confirming that each component makes an independent and synergistic contribution to performance.</p>
<p>Perhaps the most scientifically compelling findings came from the interpretability analysis. When the researchers visualized the attention weights learned by the trained model, they discovered that under high-arousal conditions the model automatically assigned its highest attention to the sequence positions where peak stimuli most frequently occurred, precisely as peak-end theory predicts. Under neutral conditions, where no salient peak existed, attention became nearly uniform across positions, consistent with the theoretical corollary that retrospective evaluation reverts to an averaging strategy when arousal variability is low. The model also revealed a previously underexplored valence asymmetry: endpoint attention weights were higher for positive sequences than negative ones, hinting at a neural basis for the everyday wisdom of ending experiences on a high note. Correlation analysis further showed that the gating signals tracked LPP amplitude more strongly than the other ERP components, with a correlation of 0.483, providing computational-level evidence that sustained emotional elaboration carried by the LPP is the primary neural signal underlying peak-end integration.</p>
<p>The implications reach well beyond the laboratory. In clinical psychology, retrospective symptom assessments such as the PHQ-9 depression questionnaire are known to be vulnerable to peak-end recall bias, and a theory-constrained decoding system of this kind could eventually support wearable-EEG monitoring that quantifies and corrects such distortions in depression, anxiety, and post-traumatic stress. In user experience research, combining EEG with peak-end gating offers a more objective way to identify the key moments that dominate a customer&#8217;s lasting impression than traditional questionnaires ever could. The authors also outline a roadmap for future work, including multi-center validation on larger and more diverse cohorts, real-time implementation on edge devices through model compression, multimodal integration with heart rate, skin conductance, and eye-tracking, and longitudinal clinical translation. The study&#8217;s broader methodological message may prove its most durable contribution: embedding mature cognitive theories directly into neural network architectures as differentiable constraints can simultaneously improve predictive accuracy, stability, and interpretability, offering a template for a new generation of theory-guided affective computing systems.</p>
<p><strong>Subject of Research:</strong> A deep learning framework that predicts retrospective evaluations of emotional experiences from ERP neural markers of the peak-end effect</p>
<p><strong>Article Title:</strong> A Deep Learning-Based Retrospective Evaluation Prediction System for Emotional Experiences: Temporal Dynamic Feature Extraction and ERP Neural Mechanisms of the Peak-End Effect</p>
<p><strong>Article References:</strong> Guo, Z. (2026). A Deep Learning-Based Retrospective Evaluation Prediction System for Emotional Experiences: Temporal Dynamic Feature Extraction and ERP Neural Mechanisms of the Peak-End Effect. <em>Cognitive Computation, 18</em>(1), Article 111. <a href="https://doi.org/10.1007/s12559-026-10657-9" rel="noopener noreferrer">https://doi.org/10.1007/s12559-026-10657-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12559-026-10657-9" rel="noopener noreferrer">10.1007/s12559-026-10657-9</a></p>
<p><strong>Keywords:</strong> peak-end effect, deep learning, EEG, event-related potentials, emotion recognition, temporal convolutional network, self-attention, LPP, retrospective evaluation, affective computing, Cognitive Computation, brain-computer interface</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202672</post-id>	</item>
		<item>
		<title>AI Model AVP-Pro Speeds Discovery of Antiviral Peptides</title>
		<link>https://scienmag.com/ai-model-avp-pro-speeds-discovery-of-antiviral-peptides/</link>
		
		<dc:creator><![CDATA[Kristina Jarvis]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:06:20 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI-driven antiviral drug development]]></category>
		<category><![CDATA[Antiviral peptide discovery]]></category>
		<category><![CDATA[antiviral peptides]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[bioinformatics tools for antiviral peptide discovery]]></category>
		<category><![CDATA[BLOSUM62]]></category>
		<category><![CDATA[computational drug design for viral inhibition]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for antiviral activity prediction]]></category>
		<category><![CDATA[ESM-2]]></category>
		<category><![CDATA[functional subtype prediction]]></category>
		<category><![CDATA[high-throughput antiviral peptide screening]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in antiviral peptide identification]]></category>
		<category><![CDATA[OHEM strategy]]></category>
		<category><![CDATA[peptide function prediction]]></category>
		<category><![CDATA[peptide-based antiviral therapeutics]]></category>
		<category><![CDATA[protein language models in peptide research]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[sequence chemistry of antiviral peptides]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[viral membrane interaction peptides]]></category>
		<category><![CDATA[virus-specific antiviral peptide prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202508</guid>

					<description><![CDATA[Researchers have developed AVP-Pro, a two-stage deep learning framework that identifies antiviral peptides and predicts which virus families and specific viruses they target.]]></description>
										<content:encoded><![CDATA[<p>Antiviral peptides have long been viewed as one of the more promising corners of the antiviral toolbox: short chains of amino acids that can interfere with viruses before they gain a foothold in a host cell. Yet identifying which peptides actually possess antiviral activity, and what kind of viral targets they prefer, has remained a slow and expensive experimental problem. A team of researchers in China now reports a computational framework, called AVP-Pro, that aims to compress that search, combining modern protein language models with a two-stage deep learning pipeline that not only flags candidate antiviral peptides but also predicts which virus families and specific viruses they are likely to act against. The work, published in BMC Genomics, was led by Xinru Wen, Weizhong Lin, Zi Liu and Xuan Xiao of the School of Information Engineering at Jingdezhen Ceramic University in Jiangxi Province.</p>
<p>The biological rationale behind the study rests on well-established sequence chemistry. Antiviral peptides tend to share recognizable characteristics: particular amino acid compositions, a net positive charge that helps them interact with negatively charged viral membranes or envelopes, hydrophobicity patterns that govern how they insert into lipid bilayers, and conserved sequence motifs that recur across peptides with similar mechanisms of action. But these features are distributed unevenly and interact in nonlinear ways, which is precisely why simple rule-based screens have struggled. Two peptides can look superficially similar yet differ sharply in antiviral potency, and a peptide that works against one virus family may be inert against another.</p>
<p>Existing computational classifiers, the authors note, have generally treated the problem as a binary one: is a given sequence an antiviral peptide or not? That framing discards valuable information. Because different AVPs exhibit distinct virus-targeting specificities, a binary answer tells a laboratory only half of what it needs to know before committing to synthesis and testing. The team also identified three persistent technical weaknesses in prior methods: difficulty modeling complex, long-range dependencies within peptide sequences; difficulty integrating heterogeneous feature sources, such as learned embeddings and hand-crafted physicochemical descriptors; and difficulty separating highly similar positive and negative samples that crowd the decision boundary between classes.</p>
<p>AVP-Pro addresses the feature-integration problem through what the authors call adaptive multi-representation fusion. The framework draws on two complementary sources of information. The first is a deep sequence representation derived from ESM-2, a large-scale protein language model trained on hundreds of millions of natural protein sequences. Such models learn to encode structural and functional context into dense numerical vectors, capturing subtle patterns that would be difficult to hand-engineer. The second source is a set of ten conventional physicochemical descriptors, the classical encodings of peptide science, covering properties such as charge, hydrophobicity and residue composition. By fusing these representations adaptively, rather than simply concatenating them, the model can weight each information source according to its usefulness for a given input.</p>
<p>Architecture plays a central role in how the framework reads a sequence. Convolutional neural network layers are used to capture local fragment-level features, short contiguous stretches of residues that often carry functional significance. Bidirectional long short-term memory units, or BiLSTM, then sweep across the sequence in both directions, capturing global contextual dependencies that span the full length of the peptide. Self-attention modules sit on top of this, allowing the model to focus dynamically on the positions most informative for the classification decision. An adaptive gating mechanism finally arbitrates among these parallel streams, deciding how much each representation should contribute to the fused output for any particular peptide rather than imposing a fixed weighting scheme across the entire dataset.</p>
<p>Perhaps the most distinctive element of the study is its handling of ambiguous samples. In AVP datasets, positive and negative sequences are often extremely similar, and this similarity produces fuzzy decision boundaries that degrade classifier performance, especially for the borderline cases that matter most in real screening pipelines. The team attacked this from two directions. First, they introduced data augmentation guided by BLOSUM62, the standard amino acid substitution matrix long used in sequence alignment. BLOSUM62 scores encode which residue substitutions are biologically tolerable, so augmenting training data with substitution-informed variants exposes the model to realistic sequence variation without straying into biologically implausible territory. Second, they employed online hard example mining, or OHEM, within a contrastive learning objective. Contrastive learning trains the model to pull similar examples together and push dissimilar ones apart in its internal embedding space; by focusing the loss on the hardest, most easily confused samples, the framework sharpens the boundary where it is thinnest.</p>
<p>The resulting system operates in two stages. In the first stage, AVP-Pro performs general antiviral peptide identification, deciding whether an input sequence is an AVP at all. On the independent test set, the authors report that the model achieved competitive predictive performance, holding its own against existing approaches while offering the richer internal machinery described above. The second stage is where the framework departs more sharply from earlier work. Building on the first-stage model through transfer learning, AVP-Pro predicts functional subtypes: which virus families and which specific viruses a candidate peptide is likely to target. In this stage the framework covered six virus families and eight specific viruses, a level of functional granularity that binary classifiers simply cannot provide.</p>
<p>Across multiple evaluation metrics, the authors report that AVP-Pro showed stable performance in both the general identification task and the functional subtype prediction task on the benchmark datasets they evaluated. Stability across metrics is an important qualifier in this field, since models optimized for a single figure of merit can exhibit lopsided behavior, excelling on accuracy while failing on measures that penalize false negatives or class imbalance. The consistency reported here suggests that the fusion and contrastive components are doing genuine work rather than merely inflating one headline number. The authors position the framework as a tool for sequence-level functional annotation and for prioritizing candidate antiviral peptides before laboratory characterization, a use case in which even modest gains in precision translate into meaningful savings of time and reagent cost.</p>
<p>The broader context makes work of this kind increasingly timely. Peptide-based antivirals occupy an attractive middle ground between small molecules and full-length protein therapeutics: they are typically less immunogenic than antibodies, more specific than broad-spectrum antiviral compounds, and synthetically accessible. But the design space of possible peptide sequences is astronomically large, and experimental screens cover only a vanishing fraction of it. Machine learning filters of the kind embodied in AVP-Pro act as a funnel, narrowing candidate lists so that wet-lab resources are spent on sequences with a computationally justified prior of activity. The addition of subtype prediction pushes the funnel further, allowing researchers to search not just for antiviral activity in general but for activity against a virus family of particular interest.</p>
<p>The study also illustrates a wider trend in computational biology: the convergence of large pretrained protein language models with task-specific deep learning heads and classical domain knowledge. ESM-2 embeddings supply the learned, context-rich representation; BLOSUM62 and physicochemical descriptors supply decades of accumulated biochemical insight; and the attention, recurrent and gating modules supply the flexibility to combine them on a per-instance basis. None of these ingredients is new on its own, but their integration, together with hard-example-focused contrastive training, reflects a maturing methodology for peptide function prediction. The work was supported by grants from the National Natural Science Foundation of China, and the authors declare no competing interests. As with any computational predictor, the model&#8217;s judgments will ultimately need experimental validation, but as a screening instrument it offers researchers a substantially finer-grained map of the antiviral peptide landscape than the binary tools that preceded it.</p>
<p><strong>Subject of Research:</strong> Machine learning-based identification and functional subtype prediction of antiviral peptides.</p>
<p><strong>Article Title:</strong> AVP-Pro: adaptive multi-representation fusion and contrastive learning for antiviral peptide identification and functional subtype prediction</p>
<p><strong>Article References:</strong> Wen, X., Lin, W., Liu, Z., &amp; Xiao, X. (2026). AVP-Pro: adaptive multi-representation fusion and contrastive learning for antiviral peptide identification and functional subtype prediction. <em>BMC Genomics</em>. <a href="https://doi.org/10.1186/s12864-026-13352-z" rel="noopener noreferrer">https://doi.org/10.1186/s12864-026-13352-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12864-026-13352-z" rel="noopener noreferrer">10.1186/s12864-026-13352-z</a></p>
<p><strong>Keywords:</strong> antiviral peptides, machine learning, ESM-2, contrastive learning, transfer learning, BLOSUM62, BiLSTM, self-attention, OHEM strategy, functional subtype prediction, peptide function prediction, deep learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202508</post-id>	</item>
		<item>
		<title>New AI Tracker Fuses Visible and Thermal Vision to Stay on Target in Any Weather</title>
		<link>https://scienmag.com/new-ai-tracker-fuses-visible-and-thermal-vision-to-stay-on-target-in-any-weather/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 17:47:43 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI object tracking]]></category>
		<category><![CDATA[AI-powered tracking in adverse weather]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[benchmark datasets for object tracking]]></category>
		<category><![CDATA[BFA-HARF tracking framework]]></category>
		<category><![CDATA[bidirectional feature adapter]]></category>
		<category><![CDATA[combined thermal and visual sensor systems]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[cross-modal feature enhancement]]></category>
		<category><![CDATA[hybrid attention]]></category>
		<category><![CDATA[multi-sensor visual tracking]]></category>
		<category><![CDATA[multimodal fusion]]></category>
		<category><![CDATA[object tracking]]></category>
		<category><![CDATA[real-time object tracking in darkness]]></category>
		<category><![CDATA[receptive fields]]></category>
		<category><![CDATA[RGB-T tracking]]></category>
		<category><![CDATA[RGB-T tracking technology]]></category>
		<category><![CDATA[robust target tracking in low-light conditions]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[thermal and visible spectrum fusion]]></category>
		<category><![CDATA[thermal infrared]]></category>
		<category><![CDATA[thermal infrared imaging applications]]></category>
		<category><![CDATA[thermal-visible camera integration]]></category>
		<category><![CDATA[tracking drift]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197103</guid>

					<description><![CDATA[Researchers have developed BFA-HARF, a tracking framework that aligns visible and thermal features bidirectionally before fusing them with hybrid attention, reducing drift and ambiguity in challenging conditions.]]></description>
										<content:encoded><![CDATA[<p>Tracking a single object through a video stream sounds like a solved problem until the lights go out. When a camera plunges into darkness, smoke fills the frame, or a pedestrian steps behind a car whose hot engine glows in the infrared, even the best visual trackers lose their grip. A research team in China now reports a new approach designed to keep computers locked onto targets under exactly these punishing conditions, by teaching an artificial intelligence to blend what ordinary cameras see with what thermal sensors feel. The work, published in the journal Applied Intelligence, introduces a tracking framework called BFA-HARF that its authors say achieves performance comparable to state-of-the-art methods across four public benchmark datasets.</p>
<p>The challenge the researchers set out to address is known as RGB-T object tracking, where RGB refers to the standard red-green-blue color imagery captured by visible-light cameras and T stands for thermal infrared. Thermal cameras detect heat rather than light, which makes them nearly immune to darkness, glare, and many kinds of visual clutter. Visible cameras, meanwhile, deliver rich texture and color detail that thermal sensors lack. In principle, combining the two should produce a tracker that works around the clock and in almost any weather. In practice, the fusion is far from trivial, because the two modalities carry fundamentally different kinds of information, and naively merging them can inject as much noise as signal.</p>
<p>According to the authors, Can Xu of East China Normal University, Weidai Xia of Central South University, Lingmin Fan of Shenergy Group, and Yue Zhang, also of East China Normal University, existing methods often suffer from insufficient feature representation and redundant cross-modal invalid information. In plain terms, the networks behind many current trackers do not extract rich enough descriptions of the target, and when they combine visible and thermal streams they frequently drag along information from one modality that is useless or misleading in the other. The result is a familiar failure mode in the tracking literature: the bounding box that is supposed to hug the target begins to drift, sometimes sliding onto a nearby distractor or ballooning into an ambiguous region that no longer corresponds to anything in the scene.</p>
<p>The team&#8217;s answer is a two-stage design philosophy they describe as align-then-fuse. Rather than throwing the two modalities together at a single point and hoping the network sorts things out, BFA-HARF first makes sure the visible and thermal features are progressively aligned and mutually enhanced, and only then applies a dedicated fusion mechanism. This sequencing, the authors argue, is what allows the tracker to build a comprehensive feature representation instead of a muddled one, and it is the conceptual core of the paper.</p>
<p>The first of the framework&#8217;s two synergistic modules is the Bidirectional Feature Adapter, or BFA. Adapters are lightweight neural components inserted into a larger network, a technique that has become popular because it lets researchers adapt powerful pretrained backbones to new tasks without retraining everything from scratch. What distinguishes the BFA is its direction of information flow. Instead of letting the two modalities exchange information only once, at a single fusion layer, the BFA facilitates a continuous bidirectional information flow between the RGB and thermal branches throughout the backbone network. At every stage of feature extraction, each modality receives a steady stream of guidance from its counterpart, so that the visible features gradually absorb thermal cues about where heat signatures lie, and the thermal features gradually absorb visible cues about texture and boundary structure. By the time the features reach the fusion stage, they are no longer two parallel, loosely related descriptions of the scene; they are two mutually refined representations that already share a common frame of reference.</p>
<p>The second module, Hybrid Attention with Receptive Fields, or HARF, takes over once the alignment is done. Its job is to process the aligned features by collaboratively capturing two complementary kinds of structure. On one side, self-attention, the mechanism that powers modern transformer architectures, lets every position in the feature map weigh the relevance of every other position, capturing global dependencies that span the entire search region. This is invaluable when a target is small, distant, or surrounded by context that matters for disambiguation. On the other side, convolutional operations excel at fine-grained local patterns, detecting edges, corners, and textures within a small neighborhood. Convolution is also constrained by its receptive field, the limited window of the input it can see at any given layer, which is precisely the weakness that attention compensates for. By hybridizing the two, the HARF module captures both the forest and the trees: the sweeping global relationships that attention provides and the sharp local detail that convolution preserves.</p>
<p>The practical payoff of this architecture, the authors report, is a measurable reduction in two of the most stubborn failure modes in multimodal tracking. The first is bounding box ambiguity, in which the predicted box becomes uncertain about exactly what it should contain, often because the fused features have blended target and background information indiscriminately. The second is tracking drift, the slow accumulation of error in which a tracker that is slightly off in one frame becomes further off in the next, eventually losing the target entirely. Because the BFA ensures that each modality continuously corrects and enriches the other, and the HARF module fuses only features that have already been aligned, the framework is better equipped to keep the target&#8217;s identity stable across frames even when one sensor&#8217;s view degrades.</p>
<p>The evidence comes from extensive experiments on four public RGB-T benchmark datasets, the standard proving grounds for this subfield, which include sequences annotated with challenging attributes such as low light, thermal crossover, occlusion, and distractors. On these benchmarks the proposed algorithm achieved performance comparable to state-of-the-art methods, according to the paper, while specifically alleviating the ambiguity and drift problems the design targets. The work builds on a deep lineage of RGB-T research, from early sparse-representation approaches that treated grayscale-thermal fusion as a collaborative coding problem, through Siamese network trackers that learned shared embeddings for both modalities, to recent transformer-based and adapter-based designs such as bi-directional adapters for multimodal tracking and prompt-driven trackers. BFA-HARF&#8217;s contribution within that lineage is the insistence that alignment and fusion are distinct problems deserving distinct, staged solutions.</p>
<p>The broader significance of the research lies in its application space. RGB-T tracking underpins technologies where failure is costly: autonomous driving systems that must keep sight of pedestrians at night, surveillance platforms operating through smoke or fog, search-and-rescue drones scanning for body heat in rubble, and traffic monitoring systems that must function in rain and glare. Thermal-infrared object detection for autonomous driving has been an active topic in the same journal, and the new tracker&#8217;s emphasis on robustness in complex environments speaks directly to those safety-critical uses. A tracker that resists drift when the visible channel fails could mean the difference between a system that reliably follows a person through a dark parking structure and one that silently loses them.</p>
<p>The authors acknowledge support from the Shanghai Special Program for Promoting High-Quality Industrial Development, and they report no conflicts of interest. The data underlying the study will be made available upon request. For the field, the paper adds a clear architectural lesson: in multimodal perception, the order of operations matters. Aligning features before fusing them, and letting that alignment happen continuously rather than at a single bottleneck, appears to squeeze more value out of each sensor and less noise into the final representation. As cameras and thermal imagers become cheaper and more common on everything from cars to consumer drones, frameworks like BFA-HARF point toward vision systems that do not blink when the lights do.</p>
<p><strong>Subject of Research:</strong> Robust RGB-T object tracking via bidirectional feature alignment and hybrid attention fusion of visible and thermal infrared imagery</p>
<p><strong>Article Title:</strong> BFA-HARF: Robust RGB-T tracking via bidirectional feature adapter and hybrid attention with receptive fields</p>
<p><strong>Article References:</strong> Xu, C., Xia, W., Fan, L., &amp; Zhang, Y. (2026). BFA-HARF: Robust RGB-T tracking via bidirectional feature adapter and hybrid attention with receptive fields. <em>Applied Intelligence, 56</em>(14), Article 416. <a href="https://doi.org/10.1007/s10489-026-07466-w" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07466-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07466-w" rel="noopener noreferrer">10.1007/s10489-026-07466-w</a></p>
<p><strong>Keywords:</strong> RGB-T tracking, thermal infrared, multimodal fusion, cross-modal feature enhancement, bidirectional feature adapter, hybrid attention, receptive fields, self-attention, tracking drift, computer vision, object tracking, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197103</post-id>	</item>
	</channel>
</rss>
