<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>semantic pattern integration &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/semantic-pattern-integration/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 15:37:01 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>semantic pattern integration &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition</title>
		<link>https://scienmag.com/new-ai-model-reads-chinese-text-at-multiple-scales-to-sharpen-entity-recognition/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 15:37:01 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in Chinese NLP technology]]></category>
		<category><![CDATA[AI models for Chinese language]]></category>
		<category><![CDATA[benchmark dataset performance]]></category>
		<category><![CDATA[benchmark datasets]]></category>
		<category><![CDATA[character and word level analysis]]></category>
		<category><![CDATA[Chinese named entity recognition]]></category>
		<category><![CDATA[Chinese text analysis]]></category>
		<category><![CDATA[Complex & Intelligent Systems]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[Cross Transformer]]></category>
		<category><![CDATA[cross-granularity contrastive learning]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for Chinese NLP]]></category>
		<category><![CDATA[entity boundary detection]]></category>
		<category><![CDATA[information extraction]]></category>
		<category><![CDATA[lattice framework]]></category>
		<category><![CDATA[lexical knowledge]]></category>
		<category><![CDATA[multi-scale entity recognition]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[representation learning]]></category>
		<category><![CDATA[semantic pattern integration]]></category>
		<category><![CDATA[semantic representations]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=228415</guid>

					<description><![CDATA[Researchers have developed ALTAI, a neural network that aligns character, word, and semantic representations through cross-granularity contrastive learning to achieve state-of-the-art performance on Chinese named entity recognition benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Chinese named entity recognition, the task of automatically finding and classifying names of people, organizations, places, and other key terms in Chinese text, has long been one of the trickiest problems in natural language processing. Unlike English, where spaces and capitalization offer obvious clues about where a name begins and ends, Chinese is written as an unbroken stream of characters, forcing machines to infer both the boundaries and the meaning of entities from context alone. A new study published in Complex &amp; Intelligent Systems by Shun Mao of Guangzhou Maritime University, Zefeng Feng and Yuncheng Jiang of South China Normal University, and their colleagues proposes a fresh attack on this problem. Their model, called ALTAI, short for A noveL cross-granulariTy contrAstive learnIng network, blends information from characters, words, and broader semantic patterns in a way that previous systems have not managed, and it consistently beats strong baseline models across four benchmark datasets.</p>
<p>To understand why ALTAI matters, it helps to grasp the peculiar challenge that Chinese poses to machines. In English, a named entity such as a company or a person is usually delimited by spaces and signaled by an initial capital letter, so a model can often spot boundaries with relative ease. Chinese text offers no such signposts. The same character sequence can be segmented into words in several plausible ways, and the meaning of a character often depends on which words it belongs to. Consider that a single character might stand alone as a word, join with a neighbor to form a two-character word, or sit inside a longer multi-character phrase. A model that only looks at characters misses the lexical knowledge embedded in dictionaries, while a model that only looks at words risks losing sight of the fine-grained compositional structure of the language. The best-performing systems therefore try to use both levels of information simultaneously.</p>
<p>For years, the dominant approach has been the so-called word-character lattice framework. The idea is to build a lattice structure over the sentence: characters form the backbone, and every dictionary word that matches a span of characters is attached as an additional node. This allows the model to consult word-level information without committing to a single segmentation. A well-known implementation of this idea is the Flat-Lattice Transformer, which folds lattice nodes into a transformer architecture so that characters and words attend to one another. Lattice-based models have delivered solid gains, and they remain widely used in Chinese NER research. Yet, according to the authors of the new study, these methods share a common blind spot.</p>
<p>That blind spot concerns how lexical information is actually integrated. In existing lattice systems, word-level knowledge is typically injected through a dedicated encoder architecture, and the interactions between character representations and word representations are handled implicitly, as a byproduct of attention mechanisms. The authors argue that this design overlooks the complex interactions across different granularities between character-level and word-level representations. In other words, the model may receive word information, but it does not explicitly learn how that information should reshape its understanding of the individual characters that compose the words. Semantic dependencies, too, are underexploited, particularly in their interactions with character representations. The result is that lexical knowledge, one of the most valuable resources available for Chinese NER, is not fully exploited.</p>
<p>ALTAI addresses this weakness with two tightly coupled components. The first is a Cross Transformer, a module designed to explicitly model interactions among three kinds of representations: the character representations that form the basic units of the text, the word-level lexical representations drawn from matching dictionary entries, and semantic representations that capture higher-level meaning. Rather than letting these views of the sentence mix only incidentally, the Cross Transformer forces a direct exchange of information across the different granularities. Each character can consult the words it participates in, and the semantic layer can modulate how both characters and words are interpreted. This explicit cross-granularity modeling is what distinguishes ALTAI from earlier lattice encoders, where the dialogue between levels of representation was left largely to chance.</p>
<p>The second component is Cross-Granularity Contrastive Learning, a training strategy borrowed from the recent wave of self-supervised learning research. Contrastive learning works by teaching a model to pull related representations together and push unrelated ones apart in the embedding space. In ALTAI, the technique is applied across granularities: character representations are aligned with their corresponding lexical and semantic views, so that the model is encouraged to produce consistent representations across the character, word, and semantic spaces. A character that participates in a meaningful word should end up with an embedding that reflects that word&#8217;s presence, and the semantic view of the sentence should agree with what the character-level view suggests. This alignment acts as an additional training signal, shaping the representation space so that entities become more discriminative, meaning that the vectors for genuine entities stand apart more clearly from those of ordinary text.</p>
<p>The combination of these two ideas produces what the authors describe as discriminative entity representations by modeling cross-granularity interactions. The intuition is straightforward: when a character-level representation has been explicitly refined by lexical and semantic information, and when the training process has actively enforced consistency among those views, the final representation of an entity span carries far richer evidence than a representation built from characters alone. Downstream, the model&#8217;s sequence labeling decisions, deciding where an entity starts, where it ends, and what type it is, rest on this enriched foundation. The contrastive objective also serves as a form of regularization, discouraging the model from latching onto spurious patterns that appear in one granularity but not in others.</p>
<p>The empirical case for ALTAI rests on experiments across four benchmark datasets for Chinese named entity recognition. The authors report that ALTAI consistently outperforms strong baseline models, including the lattice-based architectures that have defined the state of the art in this area. Consistency across multiple datasets is a meaningful claim in NER research, where a method that excels on one corpus, say one dominated by news text, may falter on another drawn from social media or specialized domains. By holding its advantage across all four benchmarks, ALTAI suggests that its benefits stem from the general principle of explicit cross-granularity interaction rather than from quirks of any particular dataset. The work was supported by the National Natural Science Foundation of China and several Guangdong provincial research programs, reflecting sustained institutional investment in Chinese-language AI research.</p>
<p>The significance of this study extends beyond Chinese NER itself. Named entity recognition is a foundational task that feeds countless downstream applications: information extraction, question answering, knowledge graph construction, search, and recommendation systems all depend on reliably identifying entities in raw text. Improvements at this level propagate through entire pipelines. Moreover, the core insight of ALTAI, that representations at different granularities should be explicitly interacted and aligned rather than merely concatenated or implicitly mixed, is a design principle that could transfer to other languages and tasks where multiple levels of linguistic structure coexist. Languages without clear word boundaries, and tasks such as relation extraction or event detection that similarly straddle character, word, and sentence levels, are natural candidates.</p>
<p>The study also highlights a broader trend in machine learning: the migration of contrastive learning from computer vision into structured language tasks. Originally popularized for learning image representations without labels, contrastive objectives are now being adapted to align representations across views, modalities, and, as here, linguistic granularities. For Chinese NER specifically, the arrival of ALTAI signals a shift in how researchers think about lexical knowledge, not as a static feature to be injected into an encoder, but as a set of interacting perspectives that a model must learn to reconcile. As the field moves forward, the open-access publication of this work, published on 1 September 2026 with a permanent DOI, means that researchers worldwide can examine, reproduce, and build upon the approach, potentially accelerating progress on a problem that sits at the heart of making machines truly literate in one of the world&#8217;s most widely spoken languages.</p>
<p><strong>Subject of Research:</strong> Cross-granularity contrastive learning for Chinese named entity recognition</p>
<p><strong>Article Title:</strong> Semantic-aware cross-granularity contrastive learning for Chinese named entity recognition</p>
<p><strong>Article References:</strong> Mao, S., Feng, Z., Li, J., Sun, C., Xiang, D., &amp; Jiang, Y. (2026). Semantic-aware cross-granularity contrastive learning for Chinese named entity recognition. <em>Complex &amp;amp; Intelligent Systems</em>. <a href="https://doi.org/10.1007/s40747-026-02483-1" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02483-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02483-1" rel="noopener noreferrer">10.1007/s40747-026-02483-1</a></p>
<p><strong>Keywords:</strong> Chinese named entity recognition, natural language processing, contrastive learning, Cross Transformer, lattice framework, lexical knowledge, semantic representations, representation learning, benchmark datasets, Complex &amp; Intelligent Systems, deep learning, information extraction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">228415</post-id>	</item>
	</channel>
</rss>
