<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>cognitive-inspired language processing &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/cognitive-inspired-language-processing/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 17:08:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>cognitive-inspired language processing &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Noise-Proof AI Translation Brings Low-Resource Languages Into the Digital Age</title>
		<link>https://scienmag.com/noise-proof-ai-translation-brings-low-resource-languages-into-the-digital-age/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 17:08:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adversarial training]]></category>
		<category><![CDATA[AI translation resilience to messy input]]></category>
		<category><![CDATA[back-translation]]></category>
		<category><![CDATA[BLEU score]]></category>
		<category><![CDATA[cognitive information processing]]></category>
		<category><![CDATA[cognitive-inspired language processing]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[enhancing multilingual communication through robust AI]]></category>
		<category><![CDATA[functional equivalence]]></category>
		<category><![CDATA[handling typos and dialectal variation in translation]]></category>
		<category><![CDATA[improving translation accuracy with limited data]]></category>
		<category><![CDATA[Low-resource language machine translation]]></category>
		<category><![CDATA[low-resource languages]]></category>
		<category><![CDATA[machine translation]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[neural computing for low-resource languages]]></category>
		<category><![CDATA[neural network robustness in translation models]]></category>
		<category><![CDATA[noise injection]]></category>
		<category><![CDATA[noise-proof neural translation frameworks]]></category>
		<category><![CDATA[noise-resistant AI translation systems]]></category>
		<category><![CDATA[scalable translation models for endangered languages]]></category>
		<category><![CDATA[scarce data language translation solutions]]></category>
		<category><![CDATA[Transformer model]]></category>
		<category><![CDATA[translation studies]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=228711</guid>

					<description><![CDATA[Researchers have built an anti-interference machine translation model that combines translation-training corpora, adversarial noise injection, and functional equivalence constraints to boost robustness for low-resource languages.]]></description>
										<content:encoded><![CDATA[<p>Machine translation has transformed how hundreds of millions of people read, work, and communicate across language barriers, but the revolution has quietly left most of the world&#8217;s languages behind. The neural systems that power fluent translation between English and Chinese, or Spanish and French, depend on enormous collections of parallel texts, sentence-by-sentence translations curated over decades. For the thousands of languages spoken by smaller communities, such resources simply do not exist, and the models trained on the scraps that are available tend to be fragile, collapsing when confronted with the typos, dialectal variation, and messy input that characterize real-world text. A new study published in Neural Computing and Applications by Yanjun Zhou of Guilin University of Electronic Technology and Shuling Zhou of No. 6 High School in Lengshuijiang, China, proposes a way to harden these models against interference, even when training data is desperately scarce.</p>
<p>The researchers frame the problem within the framework of cognitive-based information processing and applications, an approach that treats translation not as a purely statistical exercise but as a process that should mimic how human translators cope with ambiguity and noise. Their central insight is that robustness cannot be bolted onto a model after the fact; it must be built into training from the ground up. To do this, the team systematically integrated three ingredients: translation training data resources, a dynamic noise injection mechanism, and theoretical constraints drawn from functional equivalence, a concept borrowed from translation studies that judges a translation by whether it produces the same effect on its reader as the original text does on its own audience.</p>
<p>The first ingredient is unusually creative. Rather than scraping the internet for whatever parallel text exists, the authors harvested corpora generated during translation training itself, the pedagogical process by which student translators learn their craft. These corpora include complete text pairs consisting of student translations alongside professional reference translations. The pairing is valuable in a subtle way: student translations naturally contain errors, awkward phrasings, and deviations from the professional standard, so each pair captures not only what a correct translation looks like but also what plausible mistakes look like. That structure gives a machine learning system a rich signal about the boundaries between acceptable variation and genuine distortion of meaning, precisely the boundary a robust model must learn to respect.</p>
<p>On top of this foundation, the researchers applied back-translation, a well-established data augmentation technique in which a model translates target-language text back into the source language to synthesize new parallel pairs. Back-translation has long been a lifeline for low-resource machine translation because it multiplies limited data without requiring new human annotation. But synthetic data alone does not prepare a model for hostile conditions. The team therefore expanded the augmented training set by deliberately injecting noise, corrupting sentences in controlled ways so the model would encounter degraded input during training rather than only in deployment. The strategy echoes a broader trend in natural language processing, where data augmentation for limited-data learning has become one of the most reliable tools for squeezing performance out of small datasets.</p>
<p>The most technically distinctive component is the combination of adversarial training with a dynamically adjusted noise injection ratio governed by cognitively inspired attention weighting. In adversarial training, the model is trained against deliberately challenging examples, an arms race in which the system learns to resist perturbations that would otherwise derail it. Here, the adversary is a noise generator whose intensity is not fixed in advance but modulated as training proceeds, with attention mechanisms determining where interference should be concentrated to simulate realistic interference scenarios. The cognitive framing matters: human readers do not experience noise uniformly across a sentence, and their attention gravitates toward content-bearing words that carry the core meaning. By weighting noise injection in a way that reflects this cognitive reality, the training regime produces a model that fails, when it fails, in more human-like and less catastrophic ways.</p>
<p>The final piece of the architecture constrains the decoder, the component of a Transformer model that generates the output translation, using functional equivalence theory. Instead of allowing the decoder to optimize purely for surface-level likelihood, the constraint pushes it toward outputs that preserve the semantic function of the source text. This is a notable example of importing a humanistic concept from translation studies into the mathematical machinery of deep learning. Functional equivalence, developed in the tradition of dynamic equivalence in Bible translation and elaborated by generations of translation theorists, asks whether a translated text accomplishes the same communicative purpose as the original. Encoding that requirement as an optimization constraint gives the model a semantic anchor that survives even when the input is heavily corrupted.</p>
<p>The experimental results suggest the approach delivers substantial gains where they are needed most. At a noise level of 20 percent, a level of corruption that would seriously degrade a conventional model, the performance improvement for languages with high morphological complexity was approximately 26.0 percent. Morphologically complex languages, in which a single word can encode grammatical information that English spreads across several words, are particularly vulnerable to noise because a single corrupted character can destroy grammatical features the model depends on. The fact that the largest gains appear in exactly this group indicates the method is addressing a structural weakness of standard architectures rather than delivering a marginal, across-the-board bump.</p>
<p>Even more striking are the results under extreme data scarcity. With only 5,000 sentence pairs available for training, a volume that would be considered laughably small for conventional neural machine translation, the proposed model achieved a BLEU score of 16.3 in the high morphological complexity language group. BLEU, the Bilingual Evaluation Understudy, is the standard automated metric for translation quality, and every point represents a meaningful improvement in how closely machine output matches human reference translations. The 16.3 score was 3.7 points higher than a standard Transformer model trained on the same tiny dataset. For context, the Transformer architecture introduced in 2017 remains the backbone of virtually all modern machine translation, so beating it by nearly four BLEU points in a low-resource regime is a significant result for the field.</p>
<p>The study arrives amid growing concern about what researchers have called digital sidelining, the phenomenon by which speakers of low-resource languages are effectively excluded from the benefits of modern language technology. Recent surveys of neural machine translation for low-resource languages, including comprehensive reviews in the ACM Computing Surveys and analyses from Chinese-centric and multilingual perspectives, have documented how the gap between well-resourced and under-resourced languages widens as models scale. Language modeling bias has even been argued to produce forms of epistemic injustice, since communities whose languages are poorly represented by algorithms face systematic disadvantages in accessing information. Techniques like the one developed by the Zhou team offer a practical countermeasure, extracting maximum value from data that already exists rather than waiting for expensive annotation campaigns that may never come.</p>
<p>The work also carries implications beyond translation itself. The combination of adversarial noise injection, cognitively motivated attention weighting, and theory-driven output constraints could inform other natural language processing tasks in low-resource settings, from speech recognition for dialects, where noise robustness has been a persistent challenge, to sentiment analysis, cross-lingual summarization, and text classification across languages. The research was conducted as part of a project on the translation of intangible cultural heritage in Guangxi, a region of China home to numerous minority languages and traditions, underscoring the practical stakes: when a language lacks robust machine translation, its literature, oral history, and cultural documentation remain locked away from the wider digital world. By demonstrating that a carefully engineered training regime can make small-data translation models dramatically more resilient, the study offers a template for bringing the next few thousand languages, and the people who speak them, into the conversation.</p>
<p><strong>Subject of Research:</strong> Anti-interference neural machine translation for low-resource languages using translation training data, noise injection, and functional equivalence constraints</p>
<p><strong>Article Title:</strong> Construction of anti-interference translation model for low-resource languages based on translation training</p>
<p><strong>Article References:</strong> Zhou, Y., &amp; Zhou, S. (2026). Construction of anti-interference translation model for low-resource languages based on translation training. <em>Neural Computing and Applications, 38</em>(17), Article 738. <a href="https://doi.org/10.1007/s00521-026-12361-z" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12361-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12361-z" rel="noopener noreferrer">10.1007/s00521-026-12361-z</a></p>
<p><strong>Keywords:</strong> machine translation, low-resource languages, Transformer model, adversarial training, back-translation, data augmentation, functional equivalence, noise injection, BLEU score, natural language processing, cognitive information processing, translation studies</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">228711</post-id>	</item>
	</channel>
</rss>
