<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>challenges in visual variation of ancient scripts &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/challenges-in-visual-variation-of-ancient-scripts/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 12:47:40 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>challenges in visual variation of ancient scripts &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Recognize Ancient Chinese Characters It Has Never Seen Before</title>
		<link>https://scienmag.com/ai-learns-to-recognize-ancient-chinese-characters-it-has-never-seen-before/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 12:47:40 +0000</pubDate>
				<category><![CDATA[Anthropology]]></category>
		<category><![CDATA[AI understanding of Chinese character mutation over time]]></category>
		<category><![CDATA[AI-based recognition of oracle-bone and bronze inscriptions]]></category>
		<category><![CDATA[Ancient Chinese character recognition using AI]]></category>
		<category><![CDATA[ancient Chinese characters]]></category>
		<category><![CDATA[bronze inscriptions]]></category>
		<category><![CDATA[challenges in visual variation of ancient scripts]]></category>
		<category><![CDATA[character retrieval]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[cross-temporal Chinese character identification]]></category>
		<category><![CDATA[cross-validation]]></category>
		<category><![CDATA[deep learning for Chinese historical script analysis]]></category>
		<category><![CDATA[digital heritage]]></category>
		<category><![CDATA[digital paleography and historical script analysis]]></category>
		<category><![CDATA[digital tools for Chinese cultural heritage preservation]]></category>
		<category><![CDATA[DINOv2]]></category>
		<category><![CDATA[linking ancient Chinese characters through graphic structure]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for character retrieval across Chinese dynasties]]></category>
		<category><![CDATA[modeling evolution of Chinese script forms]]></category>
		<category><![CDATA[npj Heritage Science]]></category>
		<category><![CDATA[oracle-bone script]]></category>
		<category><![CDATA[paleography]]></category>
		<category><![CDATA[retrieval vs classification in AI paleography]]></category>
		<category><![CDATA[structural descriptors]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=247746</guid>

					<description><![CDATA[Researchers at the Beijing Institute of Technology show that adding expert-vetted relational graphic structure to AI models enables retrieval of later historical forms of ancient Chinese characters wholly unseen during training.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers at the Beijing Institute of Technology has shown that artificial intelligence can retrieve the later historical forms of ancient Chinese characters it has never encountered during training, provided the models are given explicit information about how the graphic structures of those characters relate to one another. The study, published in npj Heritage Science, tackles one of the most stubborn problems in digital paleography: a single character identity can persist across thousands of years while its written form mutates dramatically from the incised oracle-bone script of the Shang dynasty through bronze inscriptions to the scripts of the Spring and Autumn and Warring States periods. A machine learning system trained only on images tends to treat each visual variant as a separate object, and it collapses when asked to link a form it has never seen to the identity it truly belongs to.</p>
<p>The researchers framed the challenge as a retrieval problem rather than a classification problem. Instead of asking a model to assign a label from a fixed list, they asked it to rank candidate images so that images belonging to the same character identity, even from different historical stages, rise to the top of the list. Performance was measured with mean reciprocal rank, a metric that rewards systems for placing the correct match near the top of their rankings. Crucially, the evaluation used a whole-identity protocol: in each of five cross-validation folds, every image of a test identity was excluded not only from training and validation but also from descriptor standardization and model selection. This means the models were genuinely tested on identities wholly absent from every stage of the learning pipeline, a far stricter standard than the image-level splits that dominate much of computer vision research.</p>
<p>Building a trustworthy dataset was itself a major undertaking. The team derived eighty candidate relations describing how graphic forms connect across historical stages from metadata, and then submitted those candidates to independent review by five specialists. Only relations that survived expert scrutiny were kept, yielding a curated corpus of sixty-four character identities represented by 9,540 quality-controlled images spanning oracle-bone, bronze, Spring and Autumn, and Warring States scripts. This expert-in-the-loop filtering matters because noisy or speculative relations would teach a model the wrong structural lessons. By grounding the relational information in specialist consensus, the researchers ensured that the signal their models learned reflected genuine paleographic continuity rather than artifacts of automated metadata extraction.</p>
<p>The experimental design compared three increasingly informed approaches. The baseline used DINOv2, a self-supervised vision transformer pretrained on natural images, in frozen form, meaning its weights were not adjusted to the ancient scripts. Such general-purpose models extract powerful generic visual features, and DINOv2 alone achieved an identity-macro mean reciprocal rank of 0.469, a surprisingly strong result for imagery as distant from natural photographs as incised bone and cast bronze. Supervised contrastive learning, which trains the network to pull images of the same identity together and push images of different identities apart, improved the score to 0.554. The decisive gain, however, came from augmenting the learned features with explicit structural knowledge.</p>
<p>That structural knowledge took the form of twenty-four image-derived descriptors capturing relational graphic structure: measurements of topology, skeleton organization, component arrangement, and the geometric relationships among strokes and parts. When these descriptors were combined with the contrastively learned features, the mean reciprocal rank rose to 0.591, an improvement of 0.0369 over the contrastive baseline, with a 95 percent confidence interval of 0.0199 to 0.0540. The confidence interval excludes zero, indicating that the gain is statistically reliable rather than a fluke of a particular data split. In practical terms, the descriptors gave the models a vocabulary for talking about how a character is built, and that vocabulary transferred across the visual upheavals of script evolution.</p>
<p>An analysis of which descriptors carried the most transferable signal produced one of the study&#8217;s most interesting findings. Topology and skeleton organization, the coarse properties describing how many connected components a character has and how its structural backbone is arranged, contributed more to cross-stage retrieval than fine-grained surface details. This aligns with how paleographers themselves reason: when a character&#8217;s strokes are reshaped, elongated, or stylized over centuries, the underlying skeletal plan and topological organization often persist even as the surface appearance transforms. The machine results thus echo human expert intuition, suggesting that the relational structure encoded in these descriptors captures something real about the diachronic stability of character identity.</p>
<p>The implications extend beyond Chinese paleography. Many cultural heritage domains face analogous problems, from tracking iconographic motifs across centuries of manuscript illumination to linking ceramic fragments from different excavation layers. The study demonstrates a general recipe: pair modern self-supervised visual representations with expertly vetted relational metadata and interpretable structural descriptors, then evaluate under whole-identity exclusion so that reported performance reflects genuine generalization to unseen categories. The finding that a frozen, off-the-shelf model like DINOv2 already performs respectably on such alien imagery also suggests that heritage researchers need not always train large models from scratch; the leverage lies in the structured knowledge layered on top.</p>
<p>The authors are careful to draw a firm boundary around what their system does not do. The results, they emphasize, do not constitute automated decipherment. Retrieving the later forms of a known character identity is not the same as reading an undeciphered script, and the researchers explicitly position the work as evidence-surfacing support for expert palaeographic judgment rather than a replacement for it. In this framing, the model functions as a scholarly assistant: given an unfamiliar form, it surfaces ranked candidates and the structural relationships that connect them, giving specialists a starting point for analysis. The final interpretive act, deciding what a character means and how its history should be read, remains a human one. This division of labor reflects a mature understanding of where machine learning can reliably contribute to the humanities.</p>
<p>The study also stands out for its methodological hygiene. The five-fold whole-identity evaluation, the independent five-specialist review of candidate relations, the quality control applied to all 9,540 images, and the reporting of confidence intervals around the key performance difference collectively set a high bar for reproducibility in computational heritage research. The work received no external funding, and the authors report no competing interests. Released as open access with supplementary data and software available, the study invites other teams to test whether relational graphic structure delivers similar gains for other writing traditions, from Egyptian hieroglyphs to Linear B, where forms also evolve while identities persist. For a field where every newly legible character can reshape historical understanding, tools that reliably surface structural evidence across millennia are a meaningful step forward.</p>
<p><strong>Subject of Research:</strong> Machine learning retrieval of unseen ancient Chinese character identities using relational graphic structure</p>
<p><strong>Article Title:</strong> Relational graphic structure supports diachronic retrieval of wholly unseen ancient Chinese character identities</p>
<p><strong>Article References:</strong> Yang, Y., Qi, Y., Yang, S., Lin, Z., Zhang, L., &amp; Zhou, Y. (2026). Relational graphic structure supports diachronic retrieval of wholly unseen ancient Chinese character identities. <em>npj Heritage Science</em>. <a href="https://doi.org/10.1038/s40494-026-03044-y" rel="noopener noreferrer">https://doi.org/10.1038/s40494-026-03044-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s40494-026-03044-y" rel="noopener noreferrer">10.1038/s40494-026-03044-y</a></p>
<p><strong>Keywords:</strong> ancient Chinese characters, paleography, machine learning, DINOv2, contrastive learning, oracle-bone script, bronze inscriptions, character retrieval, digital heritage, structural descriptors, cross-validation, npj Heritage Science</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">247746</post-id>	</item>
	</channel>
</rss>
