<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>multimodal human-computer interaction &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/multimodal-human-computer-interaction/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 01:27:37 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>multimodal human-computer interaction &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Robots, Virtual Worlds and Games: Mapping the Tech Reshaping Autism Therapy</title>
		<link>https://scienmag.com/robots-virtual-worlds-and-games-mapping-the-tech-reshaping-autism-therapy/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 01:27:37 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advancements and limitations in autism technology]]></category>
		<category><![CDATA[Assistive Technology]]></category>
		<category><![CDATA[augmented reality]]></category>
		<category><![CDATA[autism spectrum disorder]]></category>
		<category><![CDATA[Autism therapy technology]]></category>
		<category><![CDATA[closed-loop adaptation]]></category>
		<category><![CDATA[comparison challenges in autism studies]]></category>
		<category><![CDATA[emerging trends in autism therapy]]></category>
		<category><![CDATA[exergames]]></category>
		<category><![CDATA[extended reality]]></category>
		<category><![CDATA[gaps in autism technology research]]></category>
		<category><![CDATA[heterogeneity in autism research]]></category>
		<category><![CDATA[humanoid robots for autism]]></category>
		<category><![CDATA[immersive virtual worlds for autism]]></category>
		<category><![CDATA[multimodal human-computer interaction]]></category>
		<category><![CDATA[neurofeedback]]></category>
		<category><![CDATA[review of autism intervention methods]]></category>
		<category><![CDATA[scoping review]]></category>
		<category><![CDATA[scoping review of autism interventions]]></category>
		<category><![CDATA[social communication]]></category>
		<category><![CDATA[socially assistive robotics]]></category>
		<category><![CDATA[technology-based autism treatments]]></category>
		<category><![CDATA[virtual reality]]></category>
		<category><![CDATA[virtual reality in autism intervention]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224854</guid>

					<description><![CDATA[A new scoping review of 72 studies maps how extended reality, socially assistive robots, exergames and adaptive systems are being used in autism interventions, revealing heavy focus on social communication and major gaps in study design.]]></description>
										<content:encoded><![CDATA[<p>From humanoid robots that rehearse conversations with children to immersive virtual reality worlds that teach joint attention, technology-based interventions for autism spectrum disorder have exploded over the past several years. But with that growth has come a problem that has quietly plagued clinicians and researchers alike: the field has become so heterogeneous, so sprawling across devices, platforms and methods, that comparing one study to another has been nearly impossible. A new scoping review published in the Journal of Autism and Developmental Disorders attempts to bring order to this chaos, and its findings reveal both the remarkable promise and the sobering gaps of a research area that is advancing faster than it can be evaluated.</p>
<p>The review, led by Xiaowen Liu and Haoli Zhao of Xi&#8217;an University of Technology together with colleagues, analyzed 72 peer-reviewed empirical publications from 2019 to 2025. Following the established JBI and PRISMA-ScR methodological guidance, the team searched four major databases—Web of Science, Scopus, PubMed and IEEE Xplore—and supplemented the results with backward and forward citation searches. Rather than sorting the studies by the specific gadget involved, the authors took a more conceptually ambitious approach: they classified every study into one of four mutually exclusive pathways defined by the dominant interaction mechanism, a framework they argue moves beyond device-based categories to focus on what actually organizes the intervention.</p>
<p>The four pathways are extended reality, or XR, which encompasses virtual and augmented reality systems; socially assistive robotics, or SAR, in which physical robots mediate therapeutic interaction; exergames and embodied-interaction systems, or EXG, which engage the body through movement-based play; and hybrid or adaptive systems, or HYB, which combine multiple modalities or adjust their behavior dynamically. The distribution across these categories is telling. XR dominated the landscape with 25 publications, or 34.7 percent of the total, followed closely by socially assistive robotics with 23 publications, or 31.9 percent. Exergames accounted for 13 publications, or 18.1 percent, while hybrid and adaptive systems made up the remaining 11 publications, or 15.3 percent. The near-parity between XR and SAR suggests that the two most visible faces of assistive technology in autism research—digital worlds on one hand, embodied machines on the other—are developing in parallel rather than one displacing the other.</p>
<p>Just as striking is what these technologies are being asked to do. Social communication appeared as a target outcome in 49 of the 72 publications, a commanding 68.1 percent, and cognition or executive function featured in 32 publications, or 44.4 percent. At the other end of the spectrum, daily-living and safety-related skills appeared in only 5 publications, a mere 6.9 percent, and sensory processing in just 3, or 4.2 percent. The imbalance is significant because sensory processing differences and adaptive daily-living skills are core features of the autism profile with profound real-world consequences. The review&#8217;s authors point to this as a clear signal that the technology pipeline, whatever its strengths, is concentrating its firepower on a narrow band of outcomes while leaving other critical domains underexplored.</p>
<p>The technical sophistication on display across the included studies is considerable. In the XR pathway, systems range from fully immersive head-mounted displays used in randomized trials of psychological and behavioral intervention to augmented reality coloring books designed to teach children to focus on specific nonverbal social cues. Some platforms integrate multi-modal sensing to assess children while they train, fusing what the system presents with what it measures. Others embed cognitive behavioral therapy within immersive environments, as in a randomized feasibility trial that paired virtual reality with CBT to treat specific phobias in young people on the spectrum. The underlying logic is consistent: virtual environments allow clinicians to control, repeat and gradually scale social scenarios that would be unpredictable or overwhelming in the real world, providing a rehearsal space where difficulty can be tuned to the individual child.</p>
<p>The socially assistive robotics pathway tells a complementary story. Humanoid platforms such as the NAO robot have been deployed to train imitation skills using human action recognition, to support joint attention in comparative studies against typically developing peers, and even to assist minimally verbal children in communication-focused therapy. Researchers have built long-term personalization into in-home robots, tracked gaze behavior across months of in-home deployment, and run randomized controlled trials comparing robot-assisted pivotal response treatment against the same therapy delivered without robotic support. One recurring theme in this literature is predictability: robots can be engineered to behave with a consistency and simplicity that human interaction partners cannot always maintain, which may lower the social demands placed on autistic children and create a bridge toward human-human interaction. Yet studies have also documented how visual and hearing sensitivities affect robot-based training, a reminder that the sensory profile of each child constrains what any single technology can deliver.</p>
<p>The exergame and embodied-interaction pathway brings the body into the loop. Motion-tracking games and somatosensory systems have been used to train motor skills, executive function and postural balance, with randomized and crossover trials examining effects on inhibitory control and restricted and repetitive behaviors. This pathway reflects a growing recognition that motor, cognitive and socio-cognitive mechanisms are intertwined in autism, and that physical activity delivered through engaging game formats can target several of them simultaneously. Augmented reality game-based cognitive-motor training and virtual reality rehabilitation studies sit at the intersection of this pathway and the XR category, illustrating why the authors chose to classify studies by interaction mechanism rather than by surface technology: the same headset can serve fundamentally different intervention logics depending on how the interaction is structured.</p>
<p>Perhaps the most forward-looking findings concern the hybrid and adaptive systems pathway, and here the numbers reveal how early the field still is. Only 11 publications fell into this category, and just 5 studies—6.9 percent of the total—used physiologically or neurally driven closed-loop adaptation, meaning systems that read signals such as brain activity or heart rate variability and adjust their behavior in real time. These include brain-computer interface video games using neurofeedback to train attention, wearable EEG neurofeedback systems built on machine learning algorithms, and mobile augmented reality neurofeedback training games. Closed-loop adaptation represents the logical endpoint of personalized intervention: a system that senses when a child is dysregulated or disengaged and responds automatically. But at fewer than one in ten studies, it remains a frontier rather than a foundation, and the review&#8217;s authors explicitly call for greater transparency about how these adaptive mechanisms work, since an opaque algorithm that changes therapy on the fly is difficult to evaluate, replicate or trust.</p>
<p>The methodological picture that emerges from the review is one of a field in transition. Only 18 of the 72 publications, or 25 percent, used randomized designs, meaning the majority of the evidence base rests on weaker study architectures that are more vulnerable to bias and less able to establish causal effects. The authors also highlight the lack of longitudinal evaluation in real-world settings: many systems are validated in short laboratory sessions with researchers present, leaving open the question of whether gains persist and generalize to homes, classrooms and communities. Their recommendations are correspondingly concrete—strengthen pathway-specific reporting so that studies within each interaction mechanism can be compared, extend evaluation over longer time horizons in authentic environments, and open up the black box of adaptive systems.</p>
<p>What makes this review more than a bookkeeping exercise is the framework itself. By organizing the literature around interaction mechanisms rather than devices, the authors have created a shared vocabulary that connects system design to intervention goals and participant needs. A clinician choosing between a robot partner and a virtual scenario is not merely choosing hardware; they are choosing between fundamentally different interaction logics with different demands on the child. Mapping those logics makes it possible, for the first time, to see where the evidence is dense, where it is thin, and where the next generation of studies should aim. As multimodal technologies continue to migrate from research labs into clinics and homes, that kind of structural clarity may prove to be the field&#8217;s most valuable intervention yet.</p>
<p><strong>Subject of Research:</strong> Multimodal human–computer interaction interventions for children and adolescents with autism spectrum disorder</p>
<p><strong>Article Title:</strong> Multimodal Human–Computer Interaction Interventions for Children and Adolescents With Autism Spectrum Disorder: A Scoping Review</p>
<p><strong>Article References:</strong> Liu, X., Zhao, H., Zhang, W., Zhang, C., &amp; Liu, X. (2026). Multimodal Human–Computer Interaction Interventions for Children and Adolescents With Autism Spectrum Disorder: A Scoping Review. <em>Journal of Autism and Developmental Disorders</em>. <a href="https://doi.org/10.1007/s10803-026-07555-2" rel="noopener noreferrer">https://doi.org/10.1007/s10803-026-07555-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10803-026-07555-2" rel="noopener noreferrer">10.1007/s10803-026-07555-2</a></p>
<p><strong>Keywords:</strong> autism spectrum disorder, multimodal human-computer interaction, extended reality, socially assistive robotics, exergames, virtual reality, augmented reality, neurofeedback, social communication, scoping review, assistive technology, closed-loop adaptation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224854</post-id>	</item>
		<item>
		<title>Multimodal AI System Enhances Emotion Detection Accuracy</title>
		<link>https://scienmag.com/multimodal-ai-system-enhances-emotion-detection-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 19:20:24 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[deep learning for emotion classification]]></category>
		<category><![CDATA[emotion recognition accuracy improvement]]></category>
		<category><![CDATA[emotion-aware artificial intelligence systems]]></category>
		<category><![CDATA[facial expression and speech emotion analysis]]></category>
		<category><![CDATA[human emotion detection using AI]]></category>
		<category><![CDATA[IEMOCAP dataset emotion benchmark]]></category>
		<category><![CDATA[integrating speech and visual cues for emotion detection]]></category>
		<category><![CDATA[multi-source emotional signal processing]]></category>
		<category><![CDATA[multimodal emotion recognition]]></category>
		<category><![CDATA[multimodal human-computer interaction]]></category>
		<category><![CDATA[reliable multimodal emotion classification]]></category>
		<category><![CDATA[speech and visual emotion analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/multimodal-ai-system-enhances-emotion-detection-accuracy/</guid>

					<description><![CDATA[A new multimodal artificial intelligence system has demonstrated that combining speech, language, and visual information can substantially improve the recognition of human emotions. Developed by researchers Erwin Budi Setiawan and Arliyanna Nilla at Telkom University in Indonesia, the system achieved an accuracy of 86.48% when classifying five emotional states: angry, excited, frustrated, neutral, and sad. [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>A new multimodal artificial intelligence system has demonstrated that combining speech, language, and visual information can substantially improve the recognition of human emotions. Developed by researchers Erwin Budi Setiawan and Arliyanna Nilla at Telkom University in Indonesia, the system achieved an accuracy of 86.48% when classifying five emotional states: angry, excited, frustrated, neutral, and sad. The results, reported in the <em>Journal of Big Data</em>, suggest that emotion-aware machines may become more reliable when they are designed to interpret several kinds of human signals at the same time rather than depending on a single source of information.</p>
<p>Emotion detection is a central challenge in human–computer interaction because feelings are rarely expressed through one channel alone. A person’s words may sound neutral while their facial expression communicates frustration, or an apparently cheerful statement may be delivered in an angry tone. Systems that analyze only text, audio, or images can therefore miss important context. The researchers addressed this limitation by constructing a unified deep learning framework that processes multiple forms of data before making a final prediction. Their approach was tested using the IEMOCAP dataset, a widely used benchmark containing acted conversations in which participants express a range of emotional states through speech, language, and visual behavior.</p>
<p>The framework begins by converting audio into linguistic information through transcription. Rather than relying primarily on acoustic properties such as pitch, volume, or speaking speed, the system uses the words spoken in the audio recordings as a major source of emotional evidence. This design allows the language-processing component to examine the semantic meaning of an utterance. The transcription is then enriched with three complementary language technologies: RoBERTa, FastText, and the NRC Emotion Lexicon. Together, these tools provide the model with contextual, lexical, and emotion-related information, helping it distinguish between words whose emotional meaning can change depending on how they are used.</p>
<p>RoBERTa is a transformer-based language model designed to interpret words in relation to their surrounding context. This is important for emotion recognition because the same term can communicate different feelings in different sentences. FastText adds another layer by representing words and their subword components, which can help the system handle vocabulary variations and less common word forms. The NRC Emotion Lexicon contributes explicit links between words and emotional categories, giving the model a structured resource that connects language with affective meaning. The researchers’ hybrid enrichment scheme combines these approaches instead of treating them as competing alternatives, creating a richer textual representation before the information is sent to the fusion stage.</p>
<p>The visual component uses ResNet-18, a convolutional neural network architecture developed for extracting features from images. In this system, ResNet-18 converts visual input into numerical representations that capture patterns potentially associated with emotion, including facial configurations and other image-level cues. The network does not simply store a picture; it transforms visual information into a feature vector that can be compared and combined with representations derived from language. This is a crucial step in multimodal learning, because audio, text, and images are expressed in different mathematical forms. Feature extraction creates a common computational basis on which the separate signals can be integrated.</p>
<p>After the individual modalities have been processed, the system applies feature fusion to combine their representations. The fused features are passed to a multilayer perceptron, or MLP, which serves as the final classifier. An MLP is a feed-forward neural network made up of interconnected layers that learn how combinations of input features correspond to output categories. In this case, it learns patterns linking language-based evidence and visual cues to the five target emotions. The architecture is intended to compensate for weaknesses in any single modality. If the wording of a sentence is ambiguous, visual information may help; if an image is unclear, the textual content may provide stronger evidence.</p>
<p>According to the study, the multimodal configuration produced an accuracy of 86.48%, approximately 7.48 percentage points higher than the best unimodal model evaluated by the researchers. The system also achieved a macro F1-score of 0.8690. The F1-score combines precision, which measures how often a predicted category is correct, and recall, which measures how many examples of that category are successfully identified. Macro-averaging calculates these results across categories and then gives each class equal weight. This matters when a dataset contains imbalanced emotional categories, because a model could otherwise appear successful by performing well on common emotions while neglecting less frequently represented ones.</p>
<p>The researchers also conducted a statistical comparison between the multimodal architecture and a baseline model. The reported Z-value was 16.94, with significance stated at p &lt; 0.05, supporting the conclusion that the observed improvement was not merely the result of random variation under the study’s evaluation procedure. The finding strengthens the central claim that integrating three modalities can make a meaningful contribution to emotion classification. However, the system’s performance should still be interpreted within the boundaries of the IEMOCAP dataset and its specific experimental design. Emotion is influenced by culture, social setting, individual personality, and spontaneous behavior, all of which can be difficult to capture in controlled or curated data.</p>
<p>The study arrives as researchers and technology companies search for more natural forms of interaction between people and machines. Emotion-sensitive systems could eventually support conversational assistants, educational software, accessibility tools, mental-health interfaces, customer-service platforms, and social robots. A system capable of detecting frustration might adjust its explanations, while one recognizing confusion or sadness could alter the tone and pace of its responses. Yet these possibilities also raise important concerns about privacy, consent, bias, and misinterpretation. Emotional states cannot be measured directly from a face, voice, or sentence with absolute certainty, and automated predictions should not be treated as definitive judgments about a person’s inner experience.</p>
<p>By combining transformer-based language analysis, word-embedding technology, an emotion lexicon, convolutional image processing, and neural feature fusion, the Telkom University team presents a technically integrated route toward more capable affective computing. Its results indicate that machines can gain a clearer statistical picture of emotion when they examine several signals together. The next challenge will be to determine how well such systems perform on spontaneous, culturally diverse, real-world interactions and whether they can explain the reasoning behind their predictions. For now, the findings offer a compelling demonstration that the future of emotion-aware artificial intelligence may depend less on teaching machines to read one signal perfectly than on teaching them to interpret many imperfect signals together.</p>
<p><strong>Subject of Research</strong>: Multimodal deep learning for human emotion detection</p>
<p><strong>Article Title</strong>: A multimodal deep learning system for enhanced emotion detection</p>
<p><strong>Article References</strong>: Budi Setiawan, E., &amp; Nilla, A. “A multimodal deep learning system for enhanced emotion detection.” <em>Journal of Big Data</em> (2026).</p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: 10.1186/s40537-026-01476-8</p>
<p><strong>Keywords</strong>: Emotion detection, multimodal learning, deep learning, IEMOCAP, feature fusion, ResNet-18, RoBERTa, macro-averaging, statistical significance, affective computing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">181825</post-id>	</item>
	</channel>
</rss>
