Friday, October 2, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition

October 2, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition

New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition

New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Chinese named entity recognition, the task of automatically finding and classifying names of people, organizations, places, and other key terms in Chinese text, has long been one of the trickiest problems in natural language processing. Unlike English, where spaces and capitalization offer obvious clues about where a name begins and ends, Chinese is written as an unbroken stream of characters, forcing machines to infer both the boundaries and the meaning of entities from context alone. A new study published in Complex & Intelligent Systems by Shun Mao of Guangzhou Maritime University, Zefeng Feng and Yuncheng Jiang of South China Normal University, and their colleagues proposes a fresh attack on this problem. Their model, called ALTAI, short for A noveL cross-granulariTy contrAstive learnIng network, blends information from characters, words, and broader semantic patterns in a way that previous systems have not managed, and it consistently beats strong baseline models across four benchmark datasets.

To understand why ALTAI matters, it helps to grasp the peculiar challenge that Chinese poses to machines. In English, a named entity such as a company or a person is usually delimited by spaces and signaled by an initial capital letter, so a model can often spot boundaries with relative ease. Chinese text offers no such signposts. The same character sequence can be segmented into words in several plausible ways, and the meaning of a character often depends on which words it belongs to. Consider that a single character might stand alone as a word, join with a neighbor to form a two-character word, or sit inside a longer multi-character phrase. A model that only looks at characters misses the lexical knowledge embedded in dictionaries, while a model that only looks at words risks losing sight of the fine-grained compositional structure of the language. The best-performing systems therefore try to use both levels of information simultaneously.

For years, the dominant approach has been the so-called word-character lattice framework. The idea is to build a lattice structure over the sentence: characters form the backbone, and every dictionary word that matches a span of characters is attached as an additional node. This allows the model to consult word-level information without committing to a single segmentation. A well-known implementation of this idea is the Flat-Lattice Transformer, which folds lattice nodes into a transformer architecture so that characters and words attend to one another. Lattice-based models have delivered solid gains, and they remain widely used in Chinese NER research. Yet, according to the authors of the new study, these methods share a common blind spot.

That blind spot concerns how lexical information is actually integrated. In existing lattice systems, word-level knowledge is typically injected through a dedicated encoder architecture, and the interactions between character representations and word representations are handled implicitly, as a byproduct of attention mechanisms. The authors argue that this design overlooks the complex interactions across different granularities between character-level and word-level representations. In other words, the model may receive word information, but it does not explicitly learn how that information should reshape its understanding of the individual characters that compose the words. Semantic dependencies, too, are underexploited, particularly in their interactions with character representations. The result is that lexical knowledge, one of the most valuable resources available for Chinese NER, is not fully exploited.

ALTAI addresses this weakness with two tightly coupled components. The first is a Cross Transformer, a module designed to explicitly model interactions among three kinds of representations: the character representations that form the basic units of the text, the word-level lexical representations drawn from matching dictionary entries, and semantic representations that capture higher-level meaning. Rather than letting these views of the sentence mix only incidentally, the Cross Transformer forces a direct exchange of information across the different granularities. Each character can consult the words it participates in, and the semantic layer can modulate how both characters and words are interpreted. This explicit cross-granularity modeling is what distinguishes ALTAI from earlier lattice encoders, where the dialogue between levels of representation was left largely to chance.

The second component is Cross-Granularity Contrastive Learning, a training strategy borrowed from the recent wave of self-supervised learning research. Contrastive learning works by teaching a model to pull related representations together and push unrelated ones apart in the embedding space. In ALTAI, the technique is applied across granularities: character representations are aligned with their corresponding lexical and semantic views, so that the model is encouraged to produce consistent representations across the character, word, and semantic spaces. A character that participates in a meaningful word should end up with an embedding that reflects that word’s presence, and the semantic view of the sentence should agree with what the character-level view suggests. This alignment acts as an additional training signal, shaping the representation space so that entities become more discriminative, meaning that the vectors for genuine entities stand apart more clearly from those of ordinary text.

The combination of these two ideas produces what the authors describe as discriminative entity representations by modeling cross-granularity interactions. The intuition is straightforward: when a character-level representation has been explicitly refined by lexical and semantic information, and when the training process has actively enforced consistency among those views, the final representation of an entity span carries far richer evidence than a representation built from characters alone. Downstream, the model’s sequence labeling decisions, deciding where an entity starts, where it ends, and what type it is, rest on this enriched foundation. The contrastive objective also serves as a form of regularization, discouraging the model from latching onto spurious patterns that appear in one granularity but not in others.

The empirical case for ALTAI rests on experiments across four benchmark datasets for Chinese named entity recognition. The authors report that ALTAI consistently outperforms strong baseline models, including the lattice-based architectures that have defined the state of the art in this area. Consistency across multiple datasets is a meaningful claim in NER research, where a method that excels on one corpus, say one dominated by news text, may falter on another drawn from social media or specialized domains. By holding its advantage across all four benchmarks, ALTAI suggests that its benefits stem from the general principle of explicit cross-granularity interaction rather than from quirks of any particular dataset. The work was supported by the National Natural Science Foundation of China and several Guangdong provincial research programs, reflecting sustained institutional investment in Chinese-language AI research.

The significance of this study extends beyond Chinese NER itself. Named entity recognition is a foundational task that feeds countless downstream applications: information extraction, question answering, knowledge graph construction, search, and recommendation systems all depend on reliably identifying entities in raw text. Improvements at this level propagate through entire pipelines. Moreover, the core insight of ALTAI, that representations at different granularities should be explicitly interacted and aligned rather than merely concatenated or implicitly mixed, is a design principle that could transfer to other languages and tasks where multiple levels of linguistic structure coexist. Languages without clear word boundaries, and tasks such as relation extraction or event detection that similarly straddle character, word, and sentence levels, are natural candidates.

The study also highlights a broader trend in machine learning: the migration of contrastive learning from computer vision into structured language tasks. Originally popularized for learning image representations without labels, contrastive objectives are now being adapted to align representations across views, modalities, and, as here, linguistic granularities. For Chinese NER specifically, the arrival of ALTAI signals a shift in how researchers think about lexical knowledge, not as a static feature to be injected into an encoder, but as a set of interacting perspectives that a model must learn to reconcile. As the field moves forward, the open-access publication of this work, published on 1 September 2026 with a permanent DOI, means that researchers worldwide can examine, reproduce, and build upon the approach, potentially accelerating progress on a problem that sits at the heart of making machines truly literate in one of the world’s most widely spoken languages.

Subject of Research: Cross-granularity contrastive learning for Chinese named entity recognition

Article Title: Semantic-aware cross-granularity contrastive learning for Chinese named entity recognition

Article References: Mao, S., Feng, Z., Li, J., Sun, C., Xiang, D., & Jiang, Y. (2026). Semantic-aware cross-granularity contrastive learning for Chinese named entity recognition. Complex & Intelligent Systems. https://doi.org/10.1007/s40747-026-02483-1

Image Credits: AI Generated

DOI: 10.1007/s40747-026-02483-1

Keywords: Chinese named entity recognition, natural language processing, contrastive learning, Cross Transformer, lattice framework, lexical knowledge, semantic representations, representation learning, benchmark datasets, Complex & Intelligent Systems, deep learning, information extraction

Cite Scienmag News

Denise Maddox. (October 2, 2026). New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition. Scienmag. https://scienmag.com/new-ai-model-reads-chinese-text-at-multiple-scales-to-sharpen-entity-recognition/

Denise Maddox. "New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition." Scienmag, 2 October 2026, https://scienmag.com/new-ai-model-reads-chinese-text-at-multiple-scales-to-sharpen-entity-recognition/. Accessed 2 October 2026.

Denise Maddox. "New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition." Scienmag. October 2, 2026. https://scienmag.com/new-ai-model-reads-chinese-text-at-multiple-scales-to-sharpen-entity-recognition/

Tags: advancements in Chinese NLP technologyAI models for Chinese languagebenchmark dataset performancebenchmark datasetscharacter and word level analysisChinese named entity recognitionChinese text analysisComplex & Intelligent Systemscontrastive learningCross Transformercross-granularity contrastive learningdeep learningdeep learning for Chinese NLPentity boundary detectioninformation extractionlattice frameworklexical knowledgemulti-scale entity recognitionnatural language processingrepresentation learningsemantic pattern integrationsemantic representations
Share26Tweet16
Previous Post

Appetite Loss During Chemotherapy Drains Diet and Quality of Life in Gynecologic Cancer Patients

Next Post

Invasive House Crows Are Reshaping Lizard Life in Tanzania’s Capital

Related Posts

Smart Contracts Emerge as the Key Driver of Blockchain Adoption in Food Supply Chains
Technology and Engineering

Smart Contracts Emerge as the Key Driver of Blockchain Adoption in Food Supply Chains

October 2, 2026
Raman Imaging Team Defends Label-Free Spectral Maps Against Criticism
Technology and Engineering

Raman Imaging Team Defends Label-Free Spectral Maps Against Criticism

October 2, 2026
Recycling’s Hidden Workhorse: How Black Mass Separation Could Make Battery Reuse Pay
Technology and Engineering

Recycling’s Hidden Workhorse: How Black Mass Separation Could Make Battery Reuse Pay

October 2, 2026
Hidden Math Cap in AI Blood Cell Models Revealed by New Study
Technology and Engineering

Hidden Math Cap in AI Blood Cell Models Revealed by New Study

October 2, 2026
Machine learning reveals investor demand as the main driver of Malaysian IPO underpricing
Technology and Engineering

Machine learning reveals investor demand as the main driver of Malaysian IPO underpricing

October 2, 2026
Pines planted far from home grow faster and shrug off drought better
Medicine

Pines planted far from home grow faster and shrug off drought better

October 2, 2026
Next Post
Invasive House Crows Are Reshaping Lizard Life in Tanzania’s Capital

Invasive House Crows Are Reshaping Lizard Life in Tanzania's Capital

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Invasive House Crows Are Reshaping Lizard Life in Tanzania’s Capital
  • New AI Model Reads Chinese Text at Multiple Scales to Sharpen Entity Recognition
  • Appetite Loss During Chemotherapy Drains Diet and Quality of Life in Gynecologic Cancer Patients
  • Liver Drug Shows Promise for Restoring Sperm Quality in Metabolic Syndrome

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading