<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>mixed reality &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/mixed-reality/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 13 Sep 2026 02:55:49 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>mixed reality &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Assistant That Reads Your Voice Makes Virtual Lab Work Feel More Human</title>
		<link>https://scienmag.com/ai-assistant-that-reads-your-voice-makes-virtual-lab-work-feel-more-human/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 02:55:49 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive AI for chemical procedure guidance]]></category>
		<category><![CDATA[affect-aware adaptation]]></category>
		<category><![CDATA[conversational agent]]></category>
		<category><![CDATA[enhancing human-computer interaction in extended reality]]></category>
		<category><![CDATA[extended reality]]></category>
		<category><![CDATA[GPT-4o-mini]]></category>
		<category><![CDATA[human-computer interaction]]></category>
		<category><![CDATA[humanized AI virtual lab assistants]]></category>
		<category><![CDATA[immersive mixed reality guided procedures]]></category>
		<category><![CDATA[immersive training]]></category>
		<category><![CDATA[large language model]]></category>
		<category><![CDATA[large language model-powered emotional AI]]></category>
		<category><![CDATA[mixed reality]]></category>
		<category><![CDATA[open-access research on emotion-aware conversational agents]]></category>
		<category><![CDATA[procedural guidance]]></category>
		<category><![CDATA[real-time emotional cue inference in conversational agents]]></category>
		<category><![CDATA[slow-down effects of emotional AI in VR]]></category>
		<category><![CDATA[speech emotion recognition]]></category>
		<category><![CDATA[supportiveness of emotion-sensitive AI in complex tasks]]></category>
		<category><![CDATA[System Usability Scale]]></category>
		<category><![CDATA[usability of emotion-aware virtual assistants]]></category>
		<category><![CDATA[user experience in mixed reality laboratories]]></category>
		<category><![CDATA[vocal prosody]]></category>
		<category><![CDATA[voice emotion recognition in mixed reality]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201060</guid>

					<description><![CDATA[A mixed-reality conversational agent that adapts its responses using emotional cues inferred from vocal prosody significantly improved perceived usability and engagement in a guided laboratory task, though users took longer to finish.]]></description>
										<content:encoded><![CDATA[<p>A conversational agent that listens not only to what users say but to how they say it has shown in a controlled experiment that it can make guided work in mixed reality feel noticeably more usable and supportive, even though it slows the task down. The finding comes from researchers at the University of Salerno, who built a mixed-reality laboratory assistant powered by a large language model and equipped it with a parallel pipeline that infers emotional cues from the prosody of a user&#8217;s voice. In a study of forty participants performing a guided chemical procedure while wearing a mixed-reality headset, those who interacted with the emotion-aware version of the agent rated the system significantly higher on perceived usability than those who received a flat, non-adaptive version.</p>
<p>The research, published open access in the Journal of Ambient Intelligence and Humanized Computing, addresses a gap that has grown as extended reality applications increasingly call on conversational agents for real-time guidance. Large language models have transformed what spoken assistants can do, replacing rigid dialogue trees with flexible, open-ended exchange. But in immersive, task-oriented environments, an agent must satisfy demands that go well beyond answering correctly: it must stay grounded in the task and the surrounding scene, respond within the tight timing of natural turn-taking, and offer help without breaking the user&#8217;s sense of presence. Menus, panels and controller commands pull attention away from the hands-on work; speech promises a more direct channel, particularly when assistance is needed mid-action rather than before or after it.</p>
<p>Whether an assistant should also read emotion from speech has remained an open question. Human voices carry information far beyond their linguistic content, and acoustic-prosodic features such as pitch, intensity, rhythm and pausing patterns offer partial evidence about a speaker&#8217;s affective state. In principle, an agent that detects frustration, hesitation or elevated effort could modulate its tone, phrasing and level of detail to better match the user&#8217;s needs. Previous work has examined conversational agents in extended reality, LLM-based dialogue and affect-aware systems largely in isolation or in pairs, but fully integrated evaluations combining spoken dialogue, language-model generation and real-time prosody-based adaptation in task-oriented immersive settings have been rare, and the empirical consequences of such adaptation for both experience and task execution were unclear.</p>
<p>To test the idea, the Salerno team designed a modular client-server system built around three layers. The client, developed in Unity 6 and running on standalone Meta Quest-class headsets, handles rendering, speech acquisition through the Meta Voice SDK, transcription via Wit.ai, gesture-based interaction through Meta&#8217;s hand-tracking SDKs, and voice playback through Meta Text-to-Speech. The backend, implemented as Python micro-services, hosts the computationally heavy work: a Dialog Orchestrator that queries GPT-4o-mini using the transcription and the retained dialogue history, and a speech emotion recognition model based on Whisper Large V3 that has been fine-tuned for emotion classification. Communication between the layers relies on JSON messages and streamed audio, with asynchronous exchanges deliberately chosen to avoid blocking operations and preserve conversational continuity.</p>
<p>The architectural elegance lies in running semantic and affective processing in parallel, so that emotional inference never delays a response. Raw audio is accumulated until roughly twenty-five seconds of speech are available, and the resulting affective estimate, complete with a confidence score, is stored in a short-lived memory valid for up to ninety seconds. When the orchestrator processes a dialogue request, it retrieves the most recent valid estimate if one exists; if not, the reply is generated without affective conditioning at all. Crucially, the emotional descriptor serves only as an additional conditioning signal. It shapes the wording, interpersonal tone, reassurance, response length and explanatory detail of the language model&#8217;s output, while leaving the procedural facts, the task objective, the agent&#8217;s voice and even its visual behavior untouched, so that affective conditioning is the only difference between the two experimental configurations.</p>
<p>The agent itself appears as a deliberately non-anthropomorphic symbolic sphere anchored in the workspace, whose brightness and surrounding animation signal whether it is idle, listening or speaking. Prior research on embodied agents suggests that less humanlike representations can strike a better balance between recognizability, expressiveness and comfort, especially when the system is not meant to emulate a person. Participants, all university students with little prior exposure to immersive technology and almost no laboratory experience, wore a Meta Quest 3 headset and carried out a nine-step chemistry procedure: placing a flask on a stand, measuring and transferring resorcinol, adding ethanol, heating the mixture, preparing a nitrating mixture from sulfuric and nitric acid, and observing the final color change. The chemistry framing was chosen not to study chemistry education but because such a procedure demands the coordination of physical manipulation, sequential decision making, spatial attention and continuous spoken dialogue that characterizes procedural assistance scenarios generally.</p>
<p>The results revealed a striking trade-off. On the System Usability Scale, the adaptive condition scored an average of 88.6 against 81.4 for the non-adaptive one, a statistically significant difference with a large effect size, and both ratings fall within what standardized guidelines describe as the excellent range. Yet participants guided by the empathetic agent took substantially longer to finish, averaging 693 seconds against 503 seconds, and engaged in more conversational turns, averaging 20.3 against 16.3, both differences statistically significant with large effects. Overall workload, measured with the NASA Task Load Index, showed no significant difference between conditions, though the subscales told a subtler story: the adaptive group reported lower physical and temporal demand but higher frustration, suggesting that the emotional adaptation softened the felt pressure of the task even as it lengthened it, and occasionally irritated users when responses seemed more elaborate than the moment required.</p>
<p>The qualitative interviews added texture to the numbers. Participants in the adaptive group frequently described the agent as more human, friendlier and more attentive, with remarks that the interaction felt natural, like dealing with someone trying to help rather than a device issuing commands, and that the agent&#8217;s manner made the procedure less stressful during its most complex phases. Others found the empathy excessive for a short task, noting that the agent sometimes explained too much or offered comfort when none was needed. The non-adaptive group, by contrast, described the interaction as simple, clear and direct, and offered notably fewer spontaneous comments, which the authors suggest may itself reflect lower engagement with an assistant perceived mainly as a functional tool.</p>
<p>The researchers are careful about interpretation. The longer completion times and extra turns cannot, on their own, distinguish productive support from unnecessary verbosity, and the authors argue the pattern is best read as a genuine trade-off between concise execution and a richer conversational experience, one whose desirability depends on context: extra dialogue may be an asset in education and training but a liability in time-critical procedures. They recommend treating affect-aware adaptation as a context-dependent design strategy rather than a default upgrade, triggering it selectively during hesitation, repeated clarification requests or errors, and regulating not just emotional tone but response length and detail. Limitations include the student-only sample, the single procedural domain, and reliance on vocal prosody alone; future work should explore multimodal affect recognition combining gestures, gaze, task progress and physiological signals, while addressing the privacy and transparency questions that such sensing inevitably raises. What the study establishes is that emotional attunement measurably reshapes how people experience machine guidance in immersive environments, making the interaction feel more supportive at the price of speed, and giving designers an evidence-based lever for deciding when that price is worth paying.</p>
<p><strong>Subject of Research:</strong> Affect-aware conversational adaptation using LLM-driven dialogue and prosody-based emotion recognition in mixed-reality procedural tasks.</p>
<p><strong>Article Title:</strong> Affect-aware conversational adaptation in mixed reality procedural tasks</p>
<p><strong>Article References:</strong> Cantone, A. A., Ercolino, M., Sebillo, M., &amp; Vitiello, G. (2026). Affect-aware conversational adaptation in mixed reality procedural tasks. <em>Journal of Ambient Intelligence and Humanized Computing</em>. <a href="https://doi.org/10.1007/s12652-026-05125-z" rel="noopener noreferrer">https://doi.org/10.1007/s12652-026-05125-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12652-026-05125-z" rel="noopener noreferrer">10.1007/s12652-026-05125-z</a></p>
<p><strong>Keywords:</strong> mixed reality, conversational agent, large language model, affect-aware adaptation, speech emotion recognition, extended reality, GPT-4o-mini, human-computer interaction, procedural guidance, System Usability Scale, vocal prosody, immersive training</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201060</post-id>	</item>
		<item>
		<title>New Real-Time Index Promises to Catch Drifting Virtual Anatomy During Brain Surgery</title>
		<link>https://scienmag.com/new-real-time-index-promises-to-catch-drifting-virtual-anatomy-during-brain-surgery/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 14:12:48 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[AR/MR surgical safety gaps]]></category>
		<category><![CDATA[augmented reality]]></category>
		<category><![CDATA[augmented reality surgical tools]]></category>
		<category><![CDATA[brain shift]]></category>
		<category><![CDATA[brain tumor resection precision]]></category>
		<category><![CDATA[continuous accuracy monitoring in surgery]]></category>
		<category><![CDATA[head-mounted display]]></category>
		<category><![CDATA[mixed reality]]></category>
		<category><![CDATA[mixed reality in brain surgery]]></category>
		<category><![CDATA[neuronavigation]]></category>
		<category><![CDATA[neurosurgery]]></category>
		<category><![CDATA[neurosurgery augmented reality]]></category>
		<category><![CDATA[preventing neurological deficits during surgery]]></category>
		<category><![CDATA[real-time]]></category>
		<category><![CDATA[real-time monitoring]]></category>
		<category><![CDATA[real-time surgical navigation]]></category>
		<category><![CDATA[real-time validation of virtual models]]></category>
		<category><![CDATA[registration accuracy]]></category>
		<category><![CDATA[Registration Quality Index]]></category>
		<category><![CDATA[registration quality index in AR]]></category>
		<category><![CDATA[safety in neurosurgical procedures]]></category>
		<category><![CDATA[surgical navigation]]></category>
		<category><![CDATA[target registration error]]></category>
		<category><![CDATA[virtual anatomy accuracy verification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=195179</guid>

					<description><![CDATA[Researchers have developed a real-time metric that continuously tracks registration stability in mixed reality neurosurgical navigation, correlating strongly with conventional accuracy measures.]]></description>
										<content:encoded><![CDATA[<p>Neurosurgeons working with mixed reality headsets may soon have a way to know—moment by moment—whether the virtual anatomy floating over their patient&#8217;s brain is still telling the truth. Researchers at the University of Salerno have developed and validated a new automated metric, the Registration Quality Index (RQI), that continuously quantifies how well a superimposed virtual anatomical model matches the physical surgical field in real time. Published in the International Journal of Computer Assisted Radiology and Surgery, the study offers a technical proof of concept for closing one of the most stubborn safety gaps in augmented and mixed reality (AR/MR) surgery: the fact that registration accuracy is typically verified once, at the start of a procedure, and then essentially taken on faith for everything that follows.</p>
<p>The stakes could hardly be higher. In neurosurgery, millimetric margins separate eloquent cortex, major blood vessels and white matter tracts from tumour tissue, and prior studies have documented that inadvertent injury to eloquent cortical regions during tumour resection produces permanent neurological deficits in a substantial proportion of patients. Errors exceeding two to three millimetres significantly increase the risk of vascular injury, incomplete resection and neurological deterioration, while deep targets such as the basal ganglia, thalamic nuclei and brainstem demand sub-millimetre precision. In deep brain stimulation surgery, targeting deviations of just one to two millimetres can reduce clinical benefit and provoke stimulation-induced adverse effects. Against these requirements, the errors reported for AR/MR neurosurgical applications—translational errors ranging from 0.62 to 6.93 mm and angular errors from 1.32 to 6.80 degrees depending on hardware and registration method—are sobering.</p>
<p>The problem, the researchers argue, is that the conventional yardsticks of navigation accuracy are fundamentally static. Fiducial Registration Error (FRE), the root mean square distance between corresponding markers after registration, and Target Registration Error (TRE), which measures accuracy at anatomical positions not used in the registration, are both discrete-point measurements computed at a single moment. Neither can capture the dynamic degradation of registration quality during an operation, which can be driven by tracking system drift from thermal effects, electromagnetic interference and line-of-sight occlusions; by patient micro-movements despite rigid cranial fixation; and by brain shift—the displacement of brain tissue caused by cerebrospinal fluid drainage, gravity and tumour resection, which can exceed 10 millimetres. Intraoperative MRI and ultrasound can partially compensate, but these modalities are specialised, interrupt the surgical workflow and are not universally available.</p>
<p>RQI attacks this gap with a deliberately lightweight computer vision pipeline that runs entirely from the headset&#8217;s own camera feed. The experimental platform paired a Varjo XR-3 head-mounted display—offering 2880 by 2720 pixels per eye at a 90 Hz refresh rate—with six HTC Vive Lighthouse base stations providing sub-millimetre optical tracking. A ceramic skull phantom, dimensioned to match adult human skull morphometry, served as the anatomical reference, and its virtual counterpart was reconstructed from computed tomography segmentation at 0.5 mm voxel resolution, imported into Unreal Engine 5.0.3 and rendered as a semi-transparent cyan overlay. Spatial registration between the virtual and physical coordinate systems was anchored to a printed 2D ArUco fiducial marker, detected by the headset&#8217;s stereo cameras at 30 Hz to compute the six-degree-of-freedom transformation in real time.</p>
<p>The computation itself proceeds in five stages on each passthrough frame. After frame acquisition at 30 Hz and a resolution of 320 by 160 pixels, the skull region is segmented from the background by adaptive thresholding, and Canny edge detection extracts the contours of both the physical phantom and the virtual overlay. Because the virtual model is rendered in cyan—a hue absent from the bone-white phantom—each detected edge pixel can be classified by examining its neighbourhood in the original RGB frame, yielding separate binary maps of physical and virtual boundaries. For every pixel on the physical boundary, the minimum Euclidean distance to the nearest virtual boundary pixel is then computed using OpenCV&#8217;s distance transform. Finally, RQI is expressed as the percentage of physical boundary pixels whose displacement exceeds a proximity threshold: a perfectly registered system scores 0 percent, while progressive misregistration drives the value toward 100 percent. Higher RQI, in other words, means worse alignment.</p>
<p>Choosing that proximity threshold was a careful balancing act. A systematic sensitivity analysis across thresholds from 2 to 12 pixels showed that values below 4 pixels produced unstable readings inflated by sub-pixel edge localisation noise, while values above 8 pixels compressed the RQI distribution toward zero and made the metric insensitive to clinically relevant degradation. The selected threshold of 6 pixels—roughly 8 millimetres at the 40 cm working distance, where one pixel corresponds to approximately 1.3 mm—maximised the Pearson correlation with TRE and was independently supported by receiver operating characteristic analysis, discriminating acceptable from degraded registration with a Youden index of 0.87, sensitivity of 0.93 and specificity of 0.94. The authors candidly note that because the threshold was tuned on the same dataset used to characterise the metric, the values carry an optimistic bias and must be confirmed on independent data.</p>
<p>The validation results were striking. Across fifteen experimental trials and seventy-five viewing configurations, RQI correlated strongly with both established metrics: r = 0.89 with FRE (95% CI 0.69–0.96, p &lt; 0.001) and r = 0.93 with TRE (95% CI 0.80–0.98, p &lt; 0.001), the latter meaning that 87 percent of the variance in target accuracy was explained by the surface-wide pixel misalignment. Regression analysis quantified the relationship in practical terms: each 1 percent increase in RQI corresponded to a 0.195 mm increase in FRE and a 0.244 mm increase in TRE. Spearman rank correlations confirmed the relationship was monotonic across the full experimental range—crucial, because for intraoperative monitoring, detecting that registration quality has degraded matters more than knowing the absolute error. Mean FRE was 2.95 ± 0.48 mm and mean TRE 3.78 ± 0.57 mm, while mean RQI was 3.28 ± 1.37 percent.</p>
<p>Speed and usability were also put to the test, with encouraging results. The entire image analysis pipeline averaged about 10 milliseconds per frame on a CPU alongside the GPU-bound rendering, achieving real-time computation at 30 Hz with no dropped frames and no disruption to workflow. A linear mixed effects model across the 75 viewing configurations revealed that viewing angle significantly influenced RQI (p &lt; 0.001, partial η² = 0.49): displacement was lowest at the frontal view and rose at oblique, lateral, posterior and contralateral perspectives, a geometrically predictable effect since errors perpendicular to the viewing axis project maximally. The authors suggest lateral views may serve as a more sensitive early warning of degradation, and that angle-dependent normalisation could improve alert specificity. Three board-certified neurosurgeons rated the system on the System Usability Scale at a mean of 79.0 ± 4.7, well above the acceptability threshold of 68.</p>
<p>The team is equally forthright about limitations. Validation was confined to a rigid phantom under controlled laboratory conditions and cannot account for brain shift, soft tissue deformation, pulsatile brain motion or variable illumination. A particularly important caveat is brain sag—the gravity-driven displacement of the brain after dural opening—which can migrate deep structures by several millimetres while the exposed cortical surface remains comparatively stable; a surface-based metric could then report acceptable alignment even as the true target drifts. The interpretability ranges offered in the study, with band boundaries at 2, 4 and 6 percent RQI, are explicitly exploratory research guidance rather than validated clinical decision boundaries, and the authors stress that RQI is not yet validated for intraoperative clinical decision-making. The most important next step, they write, is a dedicated experiment applying translations and rotations of predetermined magnitude to the registration transform to establish the metric&#8217;s sensitivity, linearity, repeatability and failure modes, followed by validation under clinical conditions and integration with intraoperative imaging.</p>
<p>Even so, the implications extend well beyond the cranial theatre. The authors point out that the absence of continuous, marker-free registration monitoring is equally pertinent to spinal navigation, where vertebral mobility between preoperative imaging and the operative position is a recognised source of drift, and where AR/MR guidance and robotic platforms are increasingly used for pedicle screw placement requiring sub-millimetre accuracy. A surface-based stability index computed on exposed bony landmarks could provide intraoperative feedback in those workflows too. If subsequent clinical studies confirm the promise of this proof of concept, the humble act of continuously comparing a cyan wireframe to the anatomy beneath it could become a quiet but essential guardian—watching, frame by frame, that the surgeon&#8217;s virtual map never silently stops matching the territory.</p>
<p><strong>Subject of Research:</strong> A real-time metric for continuous monitoring of registration stability in mixed reality neurosurgical navigation</p>
<p><strong>Article Title:</strong> A real-time metric for quantifying registration stability in mixed reality neurosurgical navigation</p>
<p><strong>Article References:</strong> Fontana, C., De Notaris, M., Iaconetta, G., &amp; Cappetti, N. (2026). A real-time metric for quantifying registration stability in mixed reality neurosurgical navigation. <em>International Journal of Computer Assisted Radiology and Surgery</em>. <a href="https://doi.org/10.1007/s11548-026-03792-z" rel="noopener noreferrer">https://doi.org/10.1007/s11548-026-03792-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11548-026-03792-z" rel="noopener noreferrer">10.1007/s11548-026-03792-z</a></p>
<p><strong>Keywords:</strong> mixed reality, neurosurgery, neuronavigation, registration accuracy, augmented reality, Registration Quality Index, target registration error, brain shift, head-mounted display, real-time monitoring, surgical navigation, real-time</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">195179</post-id>	</item>
	</channel>
</rss>
