<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>CMU-MOSEI &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/cmu-mosei/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 26 Sep 2026 02:38:43 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>CMU-MOSEI &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI network keeps reading emotions even when data streams go dark</title>
		<link>https://scienmag.com/new-ai-network-keeps-reading-emotions-even-when-data-streams-go-dark/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 02:38:43 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[affective computing]]></category>
		<category><![CDATA[AI resilience to data stream interruptions]]></category>
		<category><![CDATA[audiovisual emotion analysis]]></category>
		<category><![CDATA[CH-SIMS]]></category>
		<category><![CDATA[CMU-MOSEI]]></category>
		<category><![CDATA[CMU-MOSI]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[cross-modal fusion]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning robustness]]></category>
		<category><![CDATA[dual-path fusion network]]></category>
		<category><![CDATA[emotion detection from speech and visuals]]></category>
		<category><![CDATA[emotion inference in real-world scenarios]]></category>
		<category><![CDATA[Emotion recognition AI]]></category>
		<category><![CDATA[machine learning for incomplete data]]></category>
		<category><![CDATA[missing data in AI systems]]></category>
		<category><![CDATA[missing modality]]></category>
		<category><![CDATA[multi-sensor emotion decoding]]></category>
		<category><![CDATA[multimodal sentiment analysis]]></category>
		<category><![CDATA[robustness]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[variational autoencoder]]></category>
		<category><![CDATA[variational autoencoder models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216123</guid>

					<description><![CDATA[Researchers in China have developed a dual-path deep learning network that keeps multimodal sentiment analysis accurate even when text, audio or video inputs go missing.]]></description>
										<content:encoded><![CDATA[<p>Every time you speak, you communicate far more than words. The tilt of your eyebrows, the rhythm of your voice, the hesitation before a phrase—all of it carries emotional signal. Artificial intelligence systems built to decode these signals, known as multimodal sentiment analysis models, combine text, audio and visual data to judge whether a person sounds happy, angry or neutral. But in the real world, those data streams rarely arrive intact. A camera may fail in poor lighting, a microphone may drop out, and network delays can strip entire channels of information. A new study published in the International Journal of Machine Learning and Cybernetics tackles this fragility head-on with a deep learning architecture designed to stay accurate even when some of its inputs simply vanish.</p>
<p>The research, conducted by Jiaxu Li and Wei Liu of the School of Information and Electronic Engineering at Shandong Technology and Business University in Yantai, China, introduces a model called VMMD, short for Variational Autoencoder-Based Multi-Granularity Missing-Aware Dual-Path Fusion Network. Published on 17 September 2026, the work addresses what the authors describe as the dual challenges of accuracy and robustness in complex scenarios, where incomplete modality information degrades performance and undermines sentiment polarity inference.</p>
<p>The scale of the problem becomes clear when considering how these systems are deployed. Emotion-aware interfaces for virtual assistants, mental health screening tools and human-computer interaction platforms all rely on fusing information from multiple sensors. Earlier approaches to missing modality problems have included generative adversarial networks that attempt to impute missing views, modality translation techniques that map one channel onto another, and transformer-based feature reconstruction networks. Yet each strategy carries trade-offs: generative imputation can hallucinate plausible-looking but emotionally misleading features, while translation methods depend heavily on statistical correlations that may not hold across speakers or recording conditions.</p>
<p>At the heart of VMMD lies a component the authors call the Modal-Token Missing-Aware Module, or MT-MAM. This module generates what the researchers describe as robust proxy features—stand-in representations that can substitute for missing data. It does so by integrating core semantic information from each modality, guided by two distinct levels of modeling. At the modal level, an uncertainty weighting scheme assesses how reliable each modality is in a given instance, drawing on the principles of variational autoencoders, which encode data as probability distributions rather than fixed points and can therefore express how confident the model is about what it has learned. At the token level, the module applies missing-aware modeling of the dominant modality, typically the text stream, which in most sentiment benchmarks carries the strongest and most consistent emotional signal.</p>
<p>This dual granularity matters because emotions are expressed at different scales simultaneously. A single word can flip the polarity of an entire sentence, while a sustained facial expression shapes the interpretation of a whole video clip. By weighting whole modalities according to their uncertainty and then modeling token-level absence within the dominant modality, MT-MAM produces proxy features that preserve the core semantics needed for sentiment discrimination even when, for example, the video channel is completely unavailable. The variational autoencoder backbone allows the network to treat missingness as an uncertain condition to be reasoned about probabilistically, rather than a hard failure to be patched over with zeros or averages.</p>
<p>Proxy features alone, however, are not enough. The authors note that single proxy representations have limitations in sentiment discrimination, precisely because they compress an entire modality into a compact stand-in. To compensate, VMMD incorporates a second path built on the Transformer architecture, the same attention-driven design that underpins modern large language models. This second path integrates semantic clues from the non-dominant modalities—audio prosody, visual expressions, and any residual tokens that survived data loss—and supplies supplementary evidence to the fusion process. The dual-path design therefore runs two complementary streams of information in parallel: one robust and condensed, one rich and detailed, with each compensating for the other&#8217;s blind spots.</p>
<p>Running two paths creates its own subtle hazard, and it is here that the paper&#8217;s third innovation enters. Because the proxy features and the supplementary information are built at different granularities, their semantic representations can drift apart, becoming inconsistent in ways that confuse the final classifier. To address this, the researchers introduce the Cross-Path Semantic Consistency Constraint, or CSCC, a mechanism driven by contrastive learning. Contrastive learning, popularized in frameworks such as supervised contrastive loss, trains a model by pulling semantically similar representations closer together in an embedding space while pushing dissimilar ones apart. CSCC applies this idea across the two paths of the network, explicitly aligning the heterogeneous representations so that the proxy features and the supplementary clues tell a coherent story about the same emotional content.</p>
<p>The authors report that experiments on multiple datasets demonstrate the effectiveness and robustness of the proposed model. The experimental work, described in the paper&#8217;s contribution statement, included simulations of modality missing, data preprocessing on the widely used CMU-MOSI and CMU-MOSEI English benchmark corpora and the Chinese CH-SIMS dataset, model training and comprehensive result analysis. These benchmarks are the standard proving grounds for the field: CMU-MOSI contains short monologue video clips rated for sentiment intensity, CMU-MOSEI extends the format to thousands of speakers in the wild, and CH-SIMS provides fine-grained annotations of each modality separately, allowing researchers to test how well models handle disagreement between what is said, how it is said and how the speaker looks.</p>
<p>The broader significance of the work lies in a shift of philosophy. Rather than treating missing data as an anomaly to be repaired, VMMD treats uncertainty as a first-class property of the input, to be weighted, modeled and constrained throughout the fusion process. This mirrors a wider trend in machine learning toward reliable representation learning for incomplete multi-view data, an area that has seen growing attention as multimodal systems move out of the laboratory and into noisy deployment environments such as mobile devices, video conferencing platforms and social media monitoring.</p>
<p>For the field of affective computing, the implications extend to any application where machines must understand people through imperfect sensors. Depression screening tools that combine speech and facial analysis, sentiment monitoring of customer service calls, and social robotics all face the reality that modalities drop out unpredictably. An architecture that maintains accurate sentiment polarity inference under uncertain missingness, while explicitly reconciling the semantic tension between different levels of representation, offers a template for building systems that degrade gracefully rather than catastrophically. The research was supported by Shandong Technology and Business University, and the authors state that the data underlying the study will be made available on request.</p>
<p><strong>Subject of Research:</strong> Robust multimodal sentiment analysis under uncertain missing modalities using a dual-path fusion network</p>
<p><strong>Article Title:</strong> Dual-path fusion network for multimodal sentiment analysis under uncertain missing modalities</p>
<p><strong>Article References:</strong> Li, J., &amp; Liu, W. (2026). Dual-path fusion network for multimodal sentiment analysis under uncertain missing modalities. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 468. <a href="https://doi.org/10.1007/s13042-026-03312-0" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03312-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03312-0" rel="noopener noreferrer">10.1007/s13042-026-03312-0</a></p>
<p><strong>Keywords:</strong> multimodal sentiment analysis, missing modality, variational autoencoder, contrastive learning, Transformer, cross-modal fusion, deep learning, affective computing, robustness, CMU-MOSI, CMU-MOSEI, CH-SIMS</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216123</post-id>	</item>
		<item>
		<title>New AI Network Weighs Which Senses to Trust When Reading Human Emotion</title>
		<link>https://scienmag.com/new-ai-network-weighs-which-senses-to-trust-when-reading-human-emotion/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 18:30:12 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive gating]]></category>
		<category><![CDATA[affective computing]]></category>
		<category><![CDATA[AI emotion recognition]]></category>
		<category><![CDATA[CMU-MOSEI]]></category>
		<category><![CDATA[CMU-MOSI]]></category>
		<category><![CDATA[cross-modal transformer]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for human emotion detection]]></category>
		<category><![CDATA[emotion inference from speech and facial expressions]]></category>
		<category><![CDATA[emotion recognition]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[hierarchical disentanglement]]></category>
		<category><![CDATA[human-computer interaction applications]]></category>
		<category><![CDATA[mental health screening AI tools]]></category>
		<category><![CDATA[multimodal data fusion]]></category>
		<category><![CDATA[multimodal sentiment analysis]]></category>
		<category><![CDATA[multitask learning]]></category>
		<category><![CDATA[noisy and degraded signal handling]]></category>
		<category><![CDATA[reliability learning]]></category>
		<category><![CDATA[sensor reliability in sentiment analysis]]></category>
		<category><![CDATA[state-of-the-art AI benchmarks]]></category>
		<category><![CDATA[trust weighting in multimodal AI systems]]></category>
		<category><![CDATA[trustworthiness of communication channels]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197400</guid>

					<description><![CDATA[Researchers in China have developed a reliability-aware disentangled adaptive network that dynamically weighs the trustworthiness of language, facial, and acoustic signals to achieve superior multimodal sentiment analysis on the CMU-MOSI and CMU-MOSEI benchmarks.]]></description>
										<content:encoded><![CDATA[<p>When a person says they are fine but their voice trembles and their smile does not reach their eyes, human observers instinctively decide which signal to believe. Machines have historically been far worse at this judgment. A new study published in the International Journal of Data Science and Analytics introduces a deep learning architecture that explicitly learns how reliable each communication channel is before fusing them into a single sentiment prediction, and the approach delivers state-of-the-art results on two of the field&#8217;s most widely used benchmarks. The work, led by Jiahao Xu and colleagues at Jiangsu Ocean University in China, addresses one of the most persistent weaknesses in multimodal sentiment analysis: the tendency of models to treat every input stream as equally trustworthy, even when one of them is noisy, degraded, or actively misleading.</p>
<p>Multimodal sentiment analysis, often abbreviated MSA, is the branch of artificial intelligence that infers human affective states from heterogeneous signals such as spoken language, facial expressions, and acoustic cues. The promise of the field is considerable. Systems that can accurately read sentiment from video have obvious applications in human-computer interaction, mental health screening, customer service analytics, education technology, and social media monitoring. Yet the fundamental challenge is that real-world data is messy. A microphone may pick up background noise that distorts the prosody of a speaker&#8217;s voice. Poor lighting or a partially obscured face can render visual features unreliable. Sarcasm and irony create conflicts between what is said and how it is said. In most prior studies, the authors note, all modalities were assigned equal contributions to the final prediction, an assumption that neglects the inherent disparities in representational quality across channels and allows unreliable signals to contaminate the fused representation.</p>
<p>The proposed solution, called a reliability-aware disentangled adaptive network, consists of three cooperating components that dynamically modulate the contribution of each modality according to its information quality. The first component is a language-guided cross-modal transformer module. Transformers, the architecture family that underpins modern large language models, use attention mechanisms to weigh relationships between elements of a sequence. Here, the researchers employ a transformer to capture interactions between holistic perception and affective perception across modalities, using the linguistic stream as the guiding anchor. The module incorporates what the authors describe as adaptive hyper-learning, a mechanism that allows the model to adjust how it aligns semantics across channels during processing rather than relying on a fixed alignment scheme. This design directly targets the inadequacy of cross-modal semantic alignment, a well-known failure mode in which words, facial muscle movements, and vocal features, which unfold at different rates and carry information at different granularities, are forced into correspondence too rigidly.</p>
<p>The second component tackles a subtler problem: fusion-induced modality bias. When features from multiple channels are merged early in a network, dominant or noisy modalities can overwhelm the others, and the fused representation can entangle information that should remain separate. To mitigate this, the researchers designed a two-level coupled sentiment consistency hierarchical disentanglement module. At the fusion level, a shared encoder decomposes the fused representation into shared-semantic components, the parts of the signal that carry meaning common across all modalities. At the unimodal level, dimension-wise adaptive gating recalibrates each individual modality&#8217;s representation before it is decomposed into shared components and modality-specific private components. The gating mechanism operates on each feature dimension independently, allowing fine-grained control over which aspects of, say, the acoustic signal are amplified or suppressed. The result is a representation in which what is common to all channels and what is unique to each channel are kept distinct, reducing the risk that a strong but uninformative signal drowns out a weak but diagnostic one.</p>
<p>The third component is a reliability-aware multitask learning module that adaptively adjusts the weights assigned to each learning task during training. Multitask learning, in which a single network is trained to perform several related objectives simultaneously, is a standard technique in the field, but naive multitask setups suffer from task interference: gradients from one objective can degrade performance on another, particularly when some tasks are supervised by noisier signals. By learning to weight each task according to its reliability, the new architecture reduces interference from unreliable modalities and enhances the stability of the final prediction. This reliability awareness is the conceptual thread that ties the three modules together. Rather than adding reliability estimation as an afterthought, the network embeds it at the alignment stage, the disentanglement stage, and the training stage simultaneously.</p>
<p>The empirical evaluation was conducted on CMU-MOSI and CMU-MOSEI, the two canonical datasets for multimodal sentiment analysis. CMU-MOSI contains 2,199 short monologue video clips in which speakers express opinions on a range of topics, each annotated with sentiment intensity scores ranging from strongly negative to strongly positive. CMU-MOSEI is a far larger collection, drawn from more than 1,000 online speakers and nearly 24,000 annotated sentences, and it includes both continuous sentiment scores and discrete emotion classifications. Both datasets are publicly available on the Internet, which the authors note in their data availability statement, and they provide aligned text, visual, and acoustic features that have made them the de facto proving ground for fusion architectures over the past decade. The new model was tested in both regression and classification settings, and the authors report that extensive experiments validate its superior performance compared with existing methods on both tasks.</p>
<p>The significance of the reliability-aware framing extends beyond the benchmark numbers. A growing body of survey literature has highlighted the problem of low-quality data in multimodal machine learning generally, and recent comprehensive reviews of fusion methods have catalogued dozens of strategies, from tensor fusion networks and low-rank factorization to attention-based transformers and contrastive feature decomposition. Many of these approaches implicitly assume that all inputs are informative. By making reliability an explicit, learned quantity that modulates alignment, decomposition, and task weighting, the Chinese team&#8217;s architecture represents a shift toward what might be called quality-aware fusion, in which the network continuously asks not just what the data says but how much it should be trusted. This is particularly relevant for deployment scenarios, such as mental health monitoring or safety-critical human-robot interaction, where a single corrupted channel could push a system toward a confidently wrong conclusion.</p>
<p>The technical lineage of the work is also worth noting. Disentangled representation learning, the idea of separating shared and private factors within learned features, has been applied previously to multimodal sentiment analysis through frameworks such as modality-invariant and modality-specific representations, shared-private memory networks, and contrastive feature decomposition. The new study builds on this tradition but couples the disentanglement with sentiment consistency constraints at two levels and adds the adaptive gating and reliability-weighted multitask machinery on top. The language-guided cross-modal transformer likewise extends a line of research that began with the multimodal transformer for unaligned language sequences and continued through text-dominant perception networks that use linguistic context to structure cross-modal understanding. By combining these threads under a single reliability-aware objective, the authors have produced an architecture that is more than the sum of its parts.</p>
<p>Funded by the National Natural Science Foundation of China and several provincial and municipal research programs in Jiangsu Province, the research arrives at a moment when interest in affective computing is surging across both academia and industry. As voice assistants, video conferencing platforms, and embodied AI agents become ubiquitous, the ability to read emotional tone accurately and robustly from imperfect real-world signals will only grow in importance. The Jiangsu Ocean University team&#8217;s work suggests that the next generation of sentiment-aware systems will not simply listen harder; they will learn when to listen, when to discount, and when to let a more trustworthy channel carry the weight. For a field that has long wrestled with the gap between controlled laboratory benchmarks and the noisy reality of human communication, that lesson in calibrated trust may prove to be the most consequential contribution of all.</p>
<p><strong>Subject of Research:</strong> A reliability-aware deep learning architecture for multimodal sentiment analysis that dynamically modulates the contribution of language, visual, and acoustic modalities based on their information quality.</p>
<p><strong>Article Title:</strong> Reliability-aware disentangled adaptive network for multimodal sentiment analysis</p>
<p><strong>Article References:</strong> Xu, J., Zhao, X., Jia, L., Zhong, Z., &amp; Zhong, X. (2026). Reliability-aware disentangled adaptive network for multimodal sentiment analysis. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 298. <a href="https://doi.org/10.1007/s41060-026-01280-w" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01280-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01280-w" rel="noopener noreferrer">10.1007/s41060-026-01280-w</a></p>
<p><strong>Keywords:</strong> multimodal sentiment analysis, affective computing, cross-modal transformer, hierarchical disentanglement, reliability learning, multitask learning, CMU-MOSI, CMU-MOSEI, deep learning, emotion recognition, feature fusion, adaptive gating</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197400</post-id>	</item>
	</channel>
</rss>
