<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>painting identification using knowledge graphs &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/painting-identification-using-knowledge-graphs/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 01:25:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>painting identification using knowledge graphs &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns Art History: Knowledge Graphs Help Machines Link Paintings to the World</title>
		<link>https://scienmag.com/ai-learns-art-history-knowledge-graphs-help-machines-link-paintings-to-the-world/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 01:25:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI cultural heritage understanding]]></category>
		<category><![CDATA[AI systems for art attribution]]></category>
		<category><![CDATA[art analysis]]></category>
		<category><![CDATA[art history and AI collaboration]]></category>
		<category><![CDATA[art history knowledge graphs]]></category>
		<category><![CDATA[cultural heritage]]></category>
		<category><![CDATA[cultural heritage AI applications]]></category>
		<category><![CDATA[entity disambiguation]]></category>
		<category><![CDATA[heterogeneous graph neural networks]]></category>
		<category><![CDATA[integrating cultural context in machine learning]]></category>
		<category><![CDATA[knowledge discovery]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in art recognition]]></category>
		<category><![CDATA[multimodal entity linking]]></category>
		<category><![CDATA[multimodal entity linking in art]]></category>
		<category><![CDATA[painting identification using knowledge graphs]]></category>
		<category><![CDATA[resolving art ambiguities with AI]]></category>
		<category><![CDATA[semantic web]]></category>
		<category><![CDATA[structured knowledge in AI]]></category>
		<category><![CDATA[vision-language models]]></category>
		<category><![CDATA[visual language models for art]]></category>
		<category><![CDATA[Wikidata]]></category>
		<category><![CDATA[WikiMuSA dataset]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=260694</guid>

					<description><![CDATA[Researchers have developed KARAMEL, a knowledge-aware framework that fuses vision-language encoding with heterogeneous graph neural networks to link artworks to Wikidata entities, outperforming strong baselines on ambiguous cultural heritage mentions.]]></description>
										<content:encoded><![CDATA[<p>Art is full of riddles. A painting of a woman in a blue headscarf might be Vermeer&#8217;s Girl with a Pearl Earring, or it might be one of a hundred similar portraits hanging in a provincial museum. A caption reading &#8220;depiction of Venus&#8221; could refer to the Roman goddess, the Botticelli masterpiece, or an obscure canvas by a minor Flemish painter. Humans resolve these ambiguities effortlessly by drawing on years of accumulated cultural knowledge. Machines, even the most powerful vision-language models available today, often cannot. A new study published in the journal Machine Learning presents a system designed to close that gap, and its results suggest that structured knowledge, not just raw pattern recognition, may be the missing ingredient in how artificial intelligence understands cultural heritage.</p>
<p>The system, called KARAMEL, was developed by Raffaele Scaringi, Gennaro Vessio, and Giovanna Castellano of the University of Bari Aldo Moro in Italy, together with Alejandro Sierra-Múnera of the Hasso Plattner Institute at the University of Potsdam and Ralf Krestel of the University of Kiel and the ZBW Leibniz Information Centre for Economics. It tackles a task known as multimodal entity linking, or MEL: given a mention of an entity that appears both as an image and as text, the system must decide which specific entity in a knowledge base, in this case Wikidata, the mention refers to. Entity linking is a foundational technology for search engines, digital libraries, and question-answering systems, but the multimodal version of the problem, where visual and textual evidence must be combined, is far harder, and the art domain may be its most unforgiving test case.</p>
<p>The reason art is so difficult is that the relevant knowledge is often implicit and scattered across sources. A vision-language model can match the visual appearance of a painting to an image of the correct artwork in its database, and it can read a caption and compare it with a textual description. But when the mention is abstract or underspecified, for example a reference to an iconographic theme, a mythological subject, or an artistic movement, surface-level similarity between pixels and words is simply not enough. The researchers argue that what is needed is sociohistorical and contextual knowledge: facts about who painted what, when, where, and in what style, and how artworks, artists, locations, and movements relate to one another in a web of relationships.</p>
<p>KARAMEL&#8217;s architecture reflects that argument. It fuses two complementary components. The first is a vision-language encoder that processes the image and the textual mention, producing representations that capture what the artwork looks like and what the text says. The second is a heterogeneous graph neural network, a type of neural network designed to operate on graphs whose nodes and edges come in different types. In this case, the graph encodes structured contextual knowledge drawn from Wikidata, connecting artworks to their creators, depictions, periods, materials, and countless other relational facts. By propagating information across this heterogeneous graph, the model can perform contextual reasoning over complex relational structures, allowing it to weigh evidence that no single image or sentence could provide on its own. The graph component builds on ideas from inductive representation learning on large graphs, adapting them to the specific demands of cultural heritage data, where entity types are diverse and relationships carry rich semantics.</p>
<p>A crucial part of the contribution is a new benchmark dataset. The researchers observed that existing resources for multimodal entity linking focus on domains such as social media posts or news, where mentions are relatively concrete. To benchmark multimodal discovery in the arts, they built WikiMuSA, short for Wikidata-based Multimodal Semantic data for Art, which links artworks to Wikidata entities through images, textual descriptions, and structured contextual knowledge. Constructing the dataset required considerable care. The team extracted Wikipedia articles from a November 2023 snapshot, matching titles and languages against Wikidata, and used Wikipedia redirect queries to track down articles whose titles had changed over time. For artworks lacking an English Wikipedia article, they translated an article in another language, filtering for languages supported by the NLLB machine translation system, selecting the longest available article, translating it paragraph by paragraph, and concatenating the results.</p>
<p>Length itself posed a problem. Original Wikipedia articles about artworks ranged from about twenty words to well over a thousand, an imbalance that would bias any model trained on them. To even things out, the researchers used Llama 3, a large language model, to summarize each article with a fixed instruction prompt. The resulting summaries cluster between thirty and one hundred words, and the team measured the quality of each compression using ROUGE-1 precision, a standard metric for comparing generated text against a reference. The result is a dataset in which textual mentions are balanced in length and detail, giving competing models a fair and consistent playing field. Both the dataset and the KARAMEL code have been released publicly on Zenodo and GitHub, allowing other groups to reproduce the experiments and extend the work.</p>
<p>The experiments compared KARAMEL against strong baselines, including recent systems that apply large language models to multimodal entity linking, such as UniMEL and GEMEL, on both WikiMuSA and MELArt, an earlier multimodal entity linking dataset for art. KARAMEL surpassed these baselines, with its advantage most pronounced on implicit or abstract mentions, precisely the cases where surface-level multimodal alignment fails and contextual reasoning matters most. The comparisons required some practical accommodations: the baseline systems were trained with the authors&#8217; original configurations but with a batch size of one to fit within memory limits, a detail that underscores how computationally demanding large-language-model-based approaches to this task can be. KARAMEL&#8217;s ability to compete with and outperform such heavyweight competitors suggests that architectural knowledge integration can be more efficient than brute-force scaling.</p>
<p>Perhaps the most scientifically interesting findings come from the ablation studies, in which components of the system are removed one at a time to measure their contribution. These experiments highlight the critical role of domain-specific, structured context in learning expressive representations. When the structured knowledge from the graph is stripped away, performance drops, particularly on the ambiguous and underspecified mentions that make the art domain distinctive. In other words, the graph is not a decorative addition; it is doing real inferential work. This aligns with a growing body of evidence in the entity linking literature that knowledge graph context improves disambiguation, and it extends that evidence into the multimodal setting, where visual and textual signals must be reconciled with symbolic, relational knowledge.</p>
<p>The implications reach well beyond museums. Entity linking is the connective tissue of the semantic web, the technology that lets a mention of a person, place, or thing be resolved to a canonical identifier so that data from different sources can be joined together. Discovery science, the broader enterprise of extracting knowledge from heterogeneous data, faces the same problem everywhere: relevant knowledge is implicit and distributed across modalities. The authors suggest that KARAMEL has potential for other scientific applications where knowledge is critical for better entity linking, and the general design, a vision-language encoder fused with a heterogeneous graph neural network, could transfer to domains such as biomedicine, food science, or biodiversity, where structured ontologies already exist alongside images and text. The work also speaks to a larger debate in artificial intelligence about whether scale alone can substitute for structure. KARAMEL&#8217;s results are a data point in favor of hybrid approaches: models that combine the perceptual strength of neural encoders with the relational precision of knowledge graphs appear to reason in ways that purely pattern-based systems do not. For the cultural heritage sector, which is digitizing millions of artworks and needs automated tools to catalog and connect them, that is more than an academic result. It is a step toward machines that can genuinely read the art historical record, resolving a painting not just by what it looks like, but by what it means.</p>
<p><strong>Subject of Research:</strong> Multimodal entity linking of artworks using knowledge-aware heterogeneous graph neural networks</p>
<p><strong>Article Title:</strong> KARAMEL: Knowledge-Aware Ranking for Multimodal Entity Linking in the Arts via Heterogeneous Graphs</p>
<p><strong>Article References:</strong> Scaringi, R., Sierra-Múnera, A., Vessio, G., Castellano, G., &amp; Krestel, R. (2026). KARAMEL: Knowledge-Aware Ranking for Multimodal Entity Linking in the Arts via Heterogeneous Graphs. <em>Machine Learning, 115</em>(10), Article 243. <a href="https://doi.org/10.1007/s10994-026-07172-1" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07172-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07172-1" rel="noopener noreferrer">10.1007/s10994-026-07172-1</a></p>
<p><strong>Keywords:</strong> multimodal entity linking, knowledge graphs, heterogeneous graph neural networks, cultural heritage, Wikidata, art analysis, vision-language models, machine learning, entity disambiguation, WikiMuSA dataset, knowledge discovery, semantic web</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">260694</post-id>	</item>
	</channel>
</rss>
