<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>noisy data preprocessing &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/noisy-data-preprocessing/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 07:48:07 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>noisy data preprocessing &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Framework Cleans Up Noisy Social Media Posts to Extract Relationships More Accurately</title>
		<link>https://scienmag.com/new-ai-framework-cleans-up-noisy-social-media-posts-to-extract-relationships-more-accurately/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 07:48:07 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI frameworks for messy data]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[information extraction]]></category>
		<category><![CDATA[knowledge graph construction]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for social media]]></category>
		<category><![CDATA[MNRE dataset]]></category>
		<category><![CDATA[multimedia data integration]]></category>
		<category><![CDATA[multimodal data cleaning]]></category>
		<category><![CDATA[multimodal relation extraction]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[noise mitigation]]></category>
		<category><![CDATA[noise reduction in social media posts]]></category>
		<category><![CDATA[noisy data preprocessing]]></category>
		<category><![CDATA[relationship extraction from images and text]]></category>
		<category><![CDATA[social media]]></category>
		<category><![CDATA[social media content analysis]]></category>
		<category><![CDATA[social media data analysis]]></category>
		<category><![CDATA[social media information retrieval]]></category>
		<category><![CDATA[text-image alignment]]></category>
		<category><![CDATA[visual augmentation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221170</guid>

					<description><![CDATA[Researchers in Shanghai have developed DNMMA, a framework that cleans noise from both text and images in social media posts to improve multimodal relation extraction for knowledge graph construction.]]></description>
										<content:encoded><![CDATA[<p>Social media has become one of the richest sources of information about how people, organizations, and events connect to one another in the real world. Every day, millions of short posts pair informal text with photographs, screenshots, and memes, creating a vast stream of paired data that could, in principle, feed search engines, recommendation systems, and knowledge graphs. Turning that raw stream into structured knowledge, however, requires a machine to answer a deceptively simple question: given a post and its image, what is the relationship between the two entities mentioned in the text? This task, known as multimodal relation extraction, sits at the heart of information retrieval and knowledge graph construction, and it has long been hampered by a fundamental problem: social media data is messy. A new framework called DNMMA, published in the journal Knowledge and Information Systems, tackles that messiness head-on by cleaning up noise in both the text and the images at the same time.</p>
<p>The research, conducted by Yachuan Zhang and Yi Guo of East China University of Science and Technology in Shanghai, identifies two problems that previous work has largely overlooked. The first concerns the text itself. Social media posts are short, informal, and riddled with abbreviations, slang, emojis, and grammatical irregularities. When a language model tries to extract semantic representations from such input, the intrinsic noise of the writing style makes the resulting meaning representations unreliable. Most earlier studies concentrated on filtering noise from the visual side of the pairing, treating the text as a comparatively stable anchor. Zhang and Guo argue that this assumption is flawed: if the textual foundation is shaky, everything built on top of it inherits the instability. Their first contribution is therefore a Textual Token Refinement network, designed to distill concise, dependable semantic representations from noisy inputs before any cross-modal reasoning takes place.</p>
<p>The second problem concerns how images are used. In many existing systems, visual information enters the model as coarse-grained, global features, essentially a single vector summarizing the whole picture. The authors point out that such coarse features often introduce irrelevant interference rather than useful relational evidence. A photograph accompanying a post may contain dozens of objects, backgrounds, and incidental details that have nothing to do with the two entities whose relationship the system is trying to determine. Feeding the entire image into the fusion process risks drowning the signal in visual clutter. DNMMA addresses this with a multi-granularity visual augmentation strategy that works at several levels of detail simultaneously, combining hierarchical fusion of global features with carefully selected local regions, so that the model can draw on both the overall scene and the small patches that actually carry relational meaning.</p>
<p>One of the more inventive components of the framework is the Adaptive Semantic Mixup module. Data augmentation, the practice of creating synthetic training examples to improve model robustness, is well established in computer vision, but blindly mixing images can produce combinations that contradict the accompanying text. The Adaptive Semantic Mixup balances visual diversity against text-image consistency by adaptively fusing synthetic and original images. In practice, the system decides how much of a generated or mixed image to blend with the original photograph, ensuring that the augmented visuals remain semantically compatible with the post they accompany. This matters because the framework builds on modern generative tools: the underlying literature the authors draw on includes latent diffusion models for high-resolution image synthesis and bootstrapped language-image pre-training approaches, techniques that can produce plausible synthetic imagery but require careful control when the goal is faithful relation extraction rather than creative image generation.</p>
<p>Another key element is the Role-Aware Attention mechanism, which brings textual knowledge to bear on the visual stream. In a social media post, the entities of interest play specific roles: one may be the subject of an action and the other its object. DNMMA explicitly incorporates these textual entity roles when processing the image, allowing the model to suppress visual features that are noisy or irrelevant to the entities in question. If a post mentions two politicians and the attached photo shows them at a podium surrounded by reporters and flags, the attention mechanism can learn to focus on the regions depicting the two individuals and their interaction, downplaying the background elements. Complementing this, a Salient Visual Augmentation component strengthens the discriminative power of local key regions, sharpening the model&#8217;s ability to distinguish the visual evidence that actually supports one relationship type over another. Object detection and visual grounding research, including open-set detection systems referenced in the paper&#8217;s bibliography, provides the technical backdrop for locating these salient regions.</p>
<p>Once both modalities have been refined, DNMMA performs joint entity relation optimization, integrating the cleaned textual and visual representations to make the final relation prediction. The architecture thus follows a clear pipeline: first stabilize the text, then enrich and filter the image at multiple granularities, then fuse the two streams under the guidance of entity roles, and finally optimize the relation classification jointly across the modalities. This staged design reflects a broader lesson from the multimodal machine learning literature, where researchers have repeatedly found that simply concatenating text and image features performs poorly compared with approaches that respect the distinct noise profiles and informational roles of each modality. The authors&#8217; framework can be read as a systematic application of that lesson to the specific challenges of social media content.</p>
<p>The experimental evaluation was carried out on two benchmark datasets, MNRE and MRE-MI, both of which are publicly available. MNRE, introduced in 2021, is a challenging multimodal dataset for neural relation extraction with visual evidence in social media posts, and it has become the standard testing ground for this line of research. MRE-MI, published more recently, extends the task to posts containing multiple images, a setting that better reflects real-world social media but also amplifies the noise and alignment problems the new framework is designed to solve. According to the paper, extensive experiments on both datasets demonstrate the effectiveness of DNMMA in social media multimodal relation extraction, with the dual-modality noise mitigation and multi-granularity augmentation components each contributing to the overall performance gains.</p>
<p>The significance of this work extends beyond a single benchmark. Relation extraction is a foundational technology for knowledge graph construction, the process of building machine-readable networks of entities and their relationships that power search engines, question-answering systems, and recommendation platforms. As the field&#8217;s own survey literature notes, relation extraction has entered a new era in which large language models and multimodal inputs are reshaping what is possible. Yet the web&#8217;s most dynamic relational data, the constant chatter of social media, remains difficult to harvest because of its low quality. A framework that explicitly models and mitigates noise in both text and images, while augmenting the visual modality at multiple granularities, offers a template for making that harvest practical. The authors&#8217; emphasis on fine-grained alignment, in particular, addresses a recognized gap: earlier systems that relied on coarse visual features were often injecting as much confusion as evidence.</p>
<p>The publication also situates itself within a rapidly growing research conversation. The bibliography spans work on text-image relation propagation models, hierarchical visual prefixes for extraction, contrastive alignment methods, variational information bottlenecks for denoising, and mixup-based image augmentation for multimodal named entity recognition. DNMMA synthesizes ideas from these threads into a coherent whole, and its dual focus, mitigating noise in both modalities rather than treating the image as the sole culprit, marks a conceptual shift in how the community frames the problem. The authors declare no competing financial interests and report that no funding was received for the preparation of the manuscript, and the datasets underpinning the evaluation are openly available, which should make the results straightforward for other groups to scrutinize and extend.</p>
<p>For the broader public, the work hints at a future in which the flood of mixed text-and-image content online can be automatically organized into structured knowledge rather than simply archived. Better relation extraction means better answers to questions like who attended an event, which companies are linked to which products, or how communities relate to public figures, all derived from the informal, image-laden posts that dominate contemporary communication. It also raises the familiar stakes of automated social media analysis: the same machinery that builds knowledge graphs could be applied to monitoring, profiling, or misinformation detection, making the technical quality of these systems a matter of broad societal interest. With DNMMA, Zhang and Guo have shown that confronting the noise problem on both fronts, the written word and the accompanying picture, is not merely a cleanup exercise but a route to substantially more reliable machine understanding of the social web. The paper was received in February 2026, accepted in September 2026, and published on 1 October 2026 in volume 68 of Knowledge and Information Systems.</p>
<p><strong>Subject of Research:</strong> A dual-modality noise mitigation and multi-granularity augmentation framework for multimodal relation extraction in social media posts</p>
<p><strong>Article Title:</strong> Dnmma: dual-modality noise mitigation and multi-granularity augmentation for multimodal relation extraction in social media</p>
<p><strong>Article References:</strong> Zhang, Y., &amp; Guo, Y. (2026). Dnmma: dual-modality noise mitigation and multi-granularity augmentation for multimodal relation extraction in social media. <em>Knowledge and Information Systems, 68</em>(1), Article 271. <a href="https://doi.org/10.1007/s10115-026-02905-z" rel="noopener noreferrer">https://doi.org/10.1007/s10115-026-02905-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10115-026-02905-z" rel="noopener noreferrer">10.1007/s10115-026-02905-z</a></p>
<p><strong>Keywords:</strong> multimodal relation extraction, social media, knowledge graphs, noise mitigation, data augmentation, attention mechanism, text-image alignment, information extraction, machine learning, MNRE dataset, visual augmentation, natural language processing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221170</post-id>	</item>
	</channel>
</rss>
