Deep in the archives of Chinese and international libraries, thousands of pages of Tangut manuscripts sit waiting to be read. The script they contain was created nearly a thousand years ago for the Tangut Empire, known in Chinese historical records as the Western Xia or Xixia, which flourished across parts of northwest China between the eleventh and thirteenth centuries. The writing system is one of the most visually intricate ever devised: characters are built from geometric components stacked and interlocked in ways that can make two entirely different words look almost identical at a glance. For decades, scholars have slowly deciphered this script, but the sheer volume of surviving manuscripts, many of them damaged, faded, or handwritten in inconsistent styles, has made comprehensive digitization a daunting task. Now, a team of researchers in China has developed an artificial intelligence framework designed specifically to tackle this problem, and their approach hinges on teaching a neural network to see Tangut characters the way a philologist does: not as whole pictures, but as structured assemblies of parts.
The research, published in npj Heritage Science, introduces a model called SPMTNet, a structural-prior-guided multi-task framework for recognizing cropped Tangut character images. The work was led by Xueyan Wang of the School of Information Engineering at Ningxia University, together with Jun Zhao of Ningxia University’s School of Economics and Management, Ruixiang Yan of the College of Intelligence and Computing at Tianjin University, and Xiangqian Peng of the Academy of Xixia Studies at Ningxia University. The collaboration itself is telling: it pairs computer scientists with specialists from an academy devoted to the study of the Xixia, reflecting a growing trend in which heritage science projects bring domain experts and machine learning researchers to the same table. The paper was supported by the Major Program for Rare and Unique Studies Teams of the National Social Science Fund of China, under a project devoted to surveying and building a database of Xia translations of Chinese classical texts.
Why is Tangut so hard for machines to read? The authors identify several interlocking obstacles. First, the characters are structurally complex, often composed of multiple components arranged in tight spatial configurations. Second, many glyphs are highly similar: two characters may share nearly the same overall contour while differing only in the identity or placement of one component, much as the Latin letters a and o differ by a small internal detail, but with far higher stakes for meaning. Third, the data available for training follows a long-tailed distribution, meaning that a handful of characters appear constantly across manuscripts while thousands of others are rare, so a recognition system trained naively will become excellent at common characters and poor at the rare ones that often carry the most historical interest. Finally, annotated data is limited; creating ground-truth labels for ancient scripts requires scarce expert labor, which caps how large training datasets can grow.
Existing automatic recognition methods, the researchers note, have mostly relied on holistic appearance, treating each character image as an undivided visual pattern and learning to match overall shapes. That strategy works reasonably well when characters are visually distinct, but it struggles precisely where the hardest cases live: characters with similar contours but different internal component arrangements. A model that only memorizes what a character looks like as a whole has no principled way to notice that the left-hand component changed while the right-hand one stayed the same. The insight behind SPMTNet is that this weakness can be fixed by supplying the network with a structural prior, an explicit expectation about how Tangut characters are composed, so that the model learns to reason about parts and their spatial relationships rather than relying on gross visual similarity alone.
To make such structural learning possible, the team first had to build the right kind of data. They constructed a new dataset called XiXiaDB, a multi-source collection that contains not only whole-character images but also component annotations and explicit character-component correspondences. That last element is crucial. It means that for each character in the dataset, the system knows which components make it up and how they are arranged, giving the neural network a supervised signal about internal structure rather than just a final label. Multi-source collection also matters for real-world robustness: manuscripts scanned from different collections, written by different hands, or reproduced under different imaging conditions vary in ways that can silently break a model trained on a single homogeneous source. By drawing from multiple sources, XiXiaDB allows the researchers to test whether their method generalizes across the kinds of variation that actual digitization projects encounter.
Architecturally, SPMTNet pursues two complementary strategies. The first is multi-task learning: alongside the main task of identifying the whole character, the framework is trained on component composition and spatial relation learning, subsidiary tasks that force the network’s internal representations to encode which components are present and where they sit relative to one another. In effect, the model is asked to explain the anatomy of each character, and that anatomical understanding feeds back into better whole-character discrimination. The second strategy is a dedicated module the authors call Residual Spatial-Channel Attention, or RSCA. Attention mechanisms of this kind let a network dynamically emphasize the most informative parts of an image: the spatial dimension of RSCA highlights where in the glyph the discriminative detail lies, while the channel dimension highlights which feature types matter. Because the module is residual, it refines existing features rather than replacing them, strengthening the local cues, such as a single distinguishing stroke or component, that separate near-twin characters.
The experimental results bear out the design. On both XiXiaDB and a second benchmark called CharSample, SPMTNet outperformed representative baseline methods, with the largest gains appearing exactly where the problem is hardest: among highly similar characters and in long-tailed categories, where rare characters would otherwise be misclassified or ignored. The authors also conducted cross-source evaluations, training and testing across different data sources within Tangut character recognition, and found that the framework remained robust to source variations. That robustness is more than a technical nicety. Digitization campaigns for historical manuscripts inevitably mix material from multiple archives and imaging campaigns, and a recognition system that only works on one source’s images would require costly retraining for every new collection. A model that holds its accuracy across sources can be deployed more broadly and more cheaply.
The broader significance of the work lies in what automatic recognition unlocks for scholarship. Once characters in manuscript images can be identified reliably at scale, the resulting transcriptions become searchable text, enabling philologists to trace word usage across corpora, compare multiple copies of translated texts, and spot previously unnoticed variants. This matters especially for the Xixia corpus, which includes a remarkable body of translations of Chinese classical and Buddhist texts; systematic comparison of the Tangut versions against their Chinese sources is a major avenue of research, and the funding project behind this study is explicitly aimed at building a comprehensive database of such translations. Machine assistance does not replace expert decipherment, but it can multiply the amount of material a scholar can examine, converting years of manual lookup into hours of verification.
The approach may also travel beyond Tangut. Many historical writing systems, from Egyptian hieroglyphs to Mayan glyphs to various scripts of Central and East Asia, share the same structural logic: characters composed of meaningful parts arranged in space, with visual similarity between distinct signs and skewed frequency distributions in surviving documents. A framework that explicitly models component composition and spatial relations, supported by datasets that record character-component correspondences, offers a template that other endangered-script projects could adapt, provided comparable structural annotations can be produced. The paper, published open access under a Creative Commons license, was received in May 2026, accepted in September 2026, and published on October 1, 2026, carrying the DOI 10.1038/s40494-026-03012-6. As the authors and their funders continue building out the Xixia textual database, the quiet work of teaching machines to parse the anatomy of a medieval script may prove to be one of the more consequential collaborations between artificial intelligence and the humanities, turning a dead empire’s dense, beautiful characters into data that living scholars can search, compare, and finally read at scale.
Subject of Research: Structural-prior-guided deep learning for automatic recognition of Tangut characters in historical manuscript digitization
Article Title: Structural-prior-guided Tangut character recognition for digital preservation of historical manuscripts
Article References: Wang, X., Zhao, J., Yan, R., & Peng, X. (2026). Structural-prior-guided Tangut character recognition for digital preservation of historical manuscripts. npj Heritage Science. https://doi.org/10.1038/s40494-026-03012-6
Image Credits: AI Generated
DOI: 10.1038/s40494-026-03012-6
Keywords: Tangut script, Xixia, character recognition, deep learning, multi-task learning, attention mechanism, digital preservation, heritage science, manuscript digitization, XiXiaDB, long-tailed distribution, structural priors
Cite Scienmag News
Blake Davidson. (October 11, 2026). AI Learns to Read Tangut Script by Decoding How Characters Are Built. Scienmag. https://scienmag.com/ai-learns-to-read-tangut-script-by-decoding-how-characters-are-built/
Blake Davidson. "AI Learns to Read Tangut Script by Decoding How Characters Are Built." Scienmag, 11 October 2026, https://scienmag.com/ai-learns-to-read-tangut-script-by-decoding-how-characters-are-built/. Accessed 11 October 2026.
Blake Davidson. "AI Learns to Read Tangut Script by Decoding How Characters Are Built." Scienmag. October 11, 2026. https://scienmag.com/ai-learns-to-read-tangut-script-by-decoding-how-characters-are-built/

