Ancient murals are among the most fragile witnesses to human history, and the damage they accumulate—flaking pigment, hairline cracks, salt efflorescence, insidious erosion—is often so small, fragmented, and visually entangled with the painted surface itself that even trained conservators struggle to map it precisely. Now a team of researchers at Yunnan University in Kunming, China, has developed an artificial intelligence framework designed to do exactly that: read the faint, boundary-dominated signatures of deterioration that ordinary computer vision systems routinely miss. The system, called SFMD-Net, is described in a peer-reviewed study published in npj Heritage Science, and its results suggest that a new generation of AI models could become a practical tool for the condition assessment and digital restoration of cultural heritage.
The core problem the researchers set out to solve is deceptively simple to state and notoriously hard to solve. When a neural network looks at a mural, it sees texture everywhere: brushstrokes, pigment gradients, mineral grain, and the natural mottling of aged plaster. Damage regions are frequently tiny relative to the whole image, broken into scattered fragments rather than clean contiguous patches, and defined more by their irregular boundaries than by their interiors. Worst of all, damaged areas often look locally almost identical to intact painted texture. A conventional segmentation model—one that assigns a class label to every pixel—tends to either drown in the texture noise, hallucinating damage where none exists, or smooth over genuine deterioration because its weak visual cues are overwhelmed by stronger, more salient features elsewhere in the image.
Before building their solution, the team did something unusually rigorous: they characterised the problem itself. Rather than assuming mural damage segmentation was hard, they quantified how hard, and in what ways, by analysing the statistical structure of mural and cultural-heritage segmentation datasets alongside general natural-image and medical-image benchmarks. This cross-domain comparison pinpointed one dataset, known as YMDA, as carrying a particularly severe burden of small-region damage, fragmentation, boundary dominance, and local texture overlap with the surrounding paint. In other words, YMDA concentrates precisely the failure modes that defeat standard segmentation architectures. That diagnosis shaped every design decision that followed, turning the project from a generic application of deep learning into a targeted engineering response to a measured challenge.
The backbone of SFMD-Net is DINOv3, a state-of-the-art self-supervised vision foundation model that has been trained on enormous quantities of imagery and produces rich, general-purpose visual representations. Foundation models of this kind encode semantic information—what an object is—remarkably well, but the researchers recognised that semantic understanding alone is not enough for damage detection. Damage is not an object in the ordinary sense; it is a deviation from expected structure, often signalled as much by subtle frequency-domain patterns, such as fine-grained spectral changes in the image signal, as by what the scene depicts. SFMD-Net therefore wraps the DINOv3 backbone in purpose-built modules that inject and exploit exactly the information the backbone tends to underweight.
The first of these modules is the semantic–frequency interaction adaptor. Its job is to let the model converse between two complementary views of the same image: the semantic space of objects and regions that DINOv3 understands natively, and the frequency space where fine, repetitive, or anomalous patterns reveal themselves. By interacting these representations, the adaptor preserves weak damage cues that would otherwise be washed out, while actively suppressing the spurious responses triggered by ordinary painted texture. This is the mechanism that addresses the local texture overlap problem head-on: instead of asking the network to ignore texture, the architecture gives it a representation in which damage-related signals and texture-related signals can be separated and weighed against each other.
The second component, a high-fidelity projection bridge, tackles a different failure mode. Foundation models typically downsample images aggressively as they process them, which is efficient but destroys the fine spatial detail on which small, boundary-dominated damage regions depend. The projection bridge serves as a conduit that carries high-resolution spatial information from earlier stages of the network forward, projecting the deep semantic features back into a high-fidelity space so that the fine structure of cracks and flaking edges is not lost on the way to the final prediction. In effect, it ensures that the model’s abstract understanding of where damage is remains anchored to the precise pixels where that damage actually lives.
The third element, the multi-scale damage-aware decoder, is where the final pixel-level map is produced. Because damage appears at many scales—a crack may span a few pixels while a region of pigment loss may cover hundreds—the decoder aggregates features across multiple resolutions, attending specifically to the characteristics of damage regions: their small size, their fragmented distribution, and their boundary-dominated geometry. This design choice reflects the study’s central insight that damage segmentation is not merely a harder version of ordinary segmentation but a structurally distinct task that rewards architectures explicitly tuned to its statistics.
The experimental results bear this out. On YMDA, the dataset identified as the most punishing benchmark for this task, SFMD-Net achieved the highest Dice coefficient, intersection over union, precision, and recall among all the compared models—four standard metrics that together measure how accurately the predicted damage maps overlap the ground truth and how reliably the model avoids both false alarms and missed detections. The team also evaluated the framework on ARTeFACT, a heterogeneous benchmark of analogue-media damage that spans material types beyond murals. Under full target-domain training, SFMD-Net again obtained the highest mean Dice and intersection over union scores, demonstrating that its advantages are not confined to a single dataset or imaging style.
Perhaps the most practically significant finding concerns transfer. Training a high-performing segmentation model from scratch for every new heritage site or material is expensive, and labelled damage data is scarce. The researchers showed that fine-tuning SFMD-Net after initialising it on YMDA improved performance beyond what target-only training could achieve, indicating that the representations the model learns on mural data transfer meaningfully to heterogeneous heritage imagery. This supervised adaptation pathway points toward a workflow in which a single strong model, pre-trained on well-annotated mural damage data, can be efficiently adapted to new collections, sites, and media—precisely the kind of scalability that heritage conservation, with its backlogs of undocumented deterioration, urgently needs.
The implications extend beyond murals. Condition assessment is the foundation on which conservation planning, prioritisation, and digital restoration all rest, and the pixel-level maps that SFMD-Net produces could feed directly into assisted digital restoration pipelines, guiding where intervention is needed and documenting change over time. The work, supported by the National Natural Science Foundation of China and Yunnan provincial funding programmes, also illustrates a broader trend in applied AI: rather than simply deploying ever-larger general-purpose models, the most effective systems increasingly combine foundation-model representations with task-specific architectural innovations grounded in a quantitative understanding of the problem. For the painted walls of ancient buildings—slowly, invisibly deteriorating—that combination may mean the difference between damage noticed too late and damage caught, mapped, and repaired in time.
Subject of Research: Deep learning-based pixel-level damage segmentation in ancient murals and cultural heritage imagery
Article Title: SFMD-Net: mural damage prediction with semantic-frequency feature fusion and multi-scale structural decoding
Article References: Yuan, W., Wu, H., Yuan, G., Hu, N., & Zhang, W. (2026). SFMD-Net: mural damage prediction with semantic-frequency feature fusion and multi-scale structural decoding. npj Heritage Science. https://doi.org/10.1038/s40494-026-03015-3
Image Credits: AI Generated
DOI: 10.1038/s40494-026-03015-3
Keywords: SFMD-Net, ancient murals, cultural heritage, damage segmentation, DINOv3, deep learning, frequency-domain features, semantic segmentation, condition assessment, digital restoration, npj Heritage Science, YMDA dataset
Cite Scienmag News
Blake Davidson. (October 10, 2026). AI Learns to Spot Hidden Damage in Ancient Murals by Fusing Vision and Frequency. Scienmag. https://scienmag.com/ai-learns-to-spot-hidden-damage-in-ancient-murals-by-fusing-vision-and-frequency/
Blake Davidson. "AI Learns to Spot Hidden Damage in Ancient Murals by Fusing Vision and Frequency." Scienmag, 10 October 2026, https://scienmag.com/ai-learns-to-spot-hidden-damage-in-ancient-murals-by-fusing-vision-and-frequency/. Accessed 10 October 2026.
Blake Davidson. "AI Learns to Spot Hidden Damage in Ancient Murals by Fusing Vision and Frequency." Scienmag. October 10, 2026. https://scienmag.com/ai-learns-to-spot-hidden-damage-in-ancient-murals-by-fusing-vision-and-frequency/

