<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>smart home &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/smart-home/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 23:16:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>smart home &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Graph Networks Turn Scattered Smart Speakers Into Powerful Microphone Arrays</title>
		<link>https://scienmag.com/graph-networks-turn-scattered-smart-speakers-into-powerful-microphone-arrays/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:16:36 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[ad-hoc microphone arrays]]></category>
		<category><![CDATA[AI Flow]]></category>
		<category><![CDATA[ambient noise suppression in smart devices]]></category>
		<category><![CDATA[channel selection]]></category>
		<category><![CDATA[collaborative microphone array technology]]></category>
		<category><![CDATA[deep learning for sound source localization]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[equal error rate]]></category>
		<category><![CDATA[far-field speech processing]]></category>
		<category><![CDATA[far-field voice recognition challenges]]></category>
		<category><![CDATA[graph attention networks]]></category>
		<category><![CDATA[graph neural networks for speaker verification]]></category>
		<category><![CDATA[graph-based speech signal processing]]></category>
		<category><![CDATA[informed machine learning]]></category>
		<category><![CDATA[multi-channel audio]]></category>
		<category><![CDATA[multi-device speech processing]]></category>
		<category><![CDATA[multi-microphone collaboration algorithms]]></category>
		<category><![CDATA[noise reduction in smart home devices]]></category>
		<category><![CDATA[reverberation mitigation in voice recognition]]></category>
		<category><![CDATA[smart home]]></category>
		<category><![CDATA[Smart speaker microphone array enhancement]]></category>
		<category><![CDATA[spatial-temporal graph neural network]]></category>
		<category><![CDATA[speaker identification in noisy environments]]></category>
		<category><![CDATA[speaker verification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213107</guid>

					<description><![CDATA[Researchers have developed a spatial-temporal graph attention network that lets scattered smart devices collaborate as an ad-hoc microphone array, cutting far-field speaker verification error rates by up to 17.70 percent relative on real-world data.]]></description>
										<content:encoded><![CDATA[<p>Speaker verification, the technology that decides whether a voice belongs to a claimed identity, has quietly become one of the most important building blocks of modern life. It unlocks smartphones, guards bank accounts, and wakes up smart home hubs. Yet the technology has a persistent weakness: it works brilliantly when a microphone is close to the speaker&#8217;s mouth, and it falters badly when the voice must travel across a noisy, reverberant room. A new study published in the open-access journal Vicinagearth by Yijiang Chen, Chengdong Liang, Xiao-Lei Zhang and colleagues at Northwestern Polytechnical University and collaborating institutions tackles this far-field problem head-on, and its solution is as elegant as it is ambitious: treat every smart device in a room as a node in a graph, and let the devices learn to collaborate.</p>
<p>The core difficulty is physics. When a person speaks from across a room, the sound that reaches a distant microphone is attenuated, smeared by echoes bouncing off walls and furniture, and buried under background noise from fans, traffic, or televisions. Early speaker verification systems, dating back to the 1960s, relied on statistical models such as Gaussian mixture models with universal background models and later i-vectors. The deep learning era brought neural embeddings like x-vectors that dramatically improved accuracy, but the fundamental problem remained: a single distant microphone simply does not capture enough of the speaker&#8217;s characteristic vocal signature. Previous remedies included deep-learning speech enhancement front-ends that attempt to strip noise before verification, and domain adaptation techniques that treat noisy speech as a shifted version of clean speech. These help, but they leave a crucial resource untapped: spatial information.</p>
<p>Multi-channel approaches try to recover that spatial information by combining signals from several microphones. Fixed microphone arrays, the kind built into dedicated hardware, can apply beamforming algorithms that steer acoustic sensitivity toward the speaker. Researchers have combined neural beamforming with verification back-ends, fed multi-channel signals directly into convolutional networks, and encoded direction-of-arrival estimates into spatially aware speaker vectors. But fixed arrays have small apertures and fixed geometries. When the speaker is far away or moves around, even these sophisticated systems struggle. The truly interesting opportunity, the authors argue, lies in the devices people already own: smart speakers, phones, tablets, and other internet-connected gadgets scattered around a home or office, each with its own microphone.</p>
<p>This is the concept of the ad-hoc microphone array. Instead of a rigid cluster of microphones, an ad-hoc array is a loose federation of independently placed devices, each acting as an intelligent edge endpoint that can process audio locally. The architecture aligns with AI Flow, a decentralized artificial intelligence paradigm in which intelligent agents distributed across edge devices cooperate through coordinated computation and communication rather than shipping everything to a central server. The catch is that in an ad-hoc array, nobody controls where the devices sit. Some will be close to the speaker and capture clean audio; others will be far away, tucked behind furniture, or near a noise source, and their signals may actively harm the system. Earlier work on ad-hoc speaker verification used attention mechanisms to reweight all channels, but the new study makes a sharper observation: channels that are too noisy should not merely be down-weighted, they should be discarded.</p>
<p>The team&#8217;s framework, called a spatial-temporal graph attention network, or ST-GAT, reformulates the entire multi-channel problem as learning on a graph. In graph neural networks, data points become nodes and relationships between them become edges, described by an adjacency matrix. What makes this paper unusual is its treatment of time. Existing graph-based approaches to multi-channel audio modeled only the relationships between microphones, ignoring how information flows between successive time frames. The authors instead make each frame of each channel a node in a single static graph, so that both the spatial relationships between microphones and the temporal relationships between frames are captured in one structure. To their knowledge, this is the first time spatial-temporal data has been formulated as a static graph learning problem, a choice that is simpler, requires fewer parameters, and trains with standard graph neural network machinery.</p>
<p>There is a practical obstacle: the full adjacency matrix would be enormous. A ten-second recording from forty microphones, analyzed in ten-millisecond frames, produces a graph with forty thousand nodes and an adjacency matrix of forty thousand by forty thousand. The authors sidestep this by decomposing the aggregation into two successive modules. A temporal module builds a graph over the frames of each individual channel, and a spatial module builds a graph over the channels at each frame. They implemented the aggregation in two ways: a self-attention mechanism that uses the adjacency matrix as a mask over attention scores, and a true graph attention network in which each node attends to its neighbors using its own representation as the query. Both use multi-head attention, and the blocks are stacked to deepen the model.</p>
<p>The second innovation is a graph-based channel selection block that exploits prior knowledge during training. When the training data includes labeled positions of microphones, speakers, and noise sources, the authors construct an auxiliary adjacency matrix that connects only the most useful channels, for example the microphones closest to the speaker, or those with the best signal-to-noise ratio, while masking out nodes near a noise source or behind the speaker&#8217;s head. Crucially, this selection happens only during training. At test time, when the layout of devices and speakers is unknown, the network applies a fully connected graph and autonomously decides which channels to trust, because the attention parameters have learned the selection rule through backpropagation. The authors frame this as a form of informed machine learning, where handcrafted rules inject prior knowledge into the graph topology.</p>
<p>Training proceeds in two stages. First, a standard single-channel verification system, built from a residual convolutional network front-end, self-attentive pooling, and a classification layer, is trained on abundant single-channel speech. The frame-level feature extractor is then frozen, and the graph-based channel fusion modules are trained on spatial-temporal data from ad-hoc arrays. The evaluation was thorough: two simulated datasets, LibriSIMU-noise and LibriSIMU-reverb, generated with random room dimensions, reverberation times up to 1.2 seconds, and signal-to-noise ratios down to minus five decibels, plus two real-world corpora, Libri-adhoc40, a forty-node replayed array recorded in a highly reverberant office, and Hi-mia, a far-field text-dependent smart home dataset. Six representative baselines were compared, including oracle selection of the closest microphone, classical beamforming, energy-envelope channel selection, and attention-based multi-channel aggregation methods.</p>
<p>The results are striking. On the simulated datasets, the best proposed variant achieved a relative reduction in equal error rate of 15.39 percent compared with the strongest reference method, and on the real-world data the reduction reached 17.70 percent. Ablation studies showed that the auxiliary adjacency matrix brought especially large gains when combined with the ST-GAT backbone, cutting the error rate by roughly 15.65 percent relative on noisy simulated data and about 18.94 percent relative on the real Libri-adhoc40 corpus in the eight-channel scenario. The system remained robust across signal-to-noise ratios from minus five to twenty decibels and reverberation times up to 1.2 seconds, and it transferred to a different verification architecture, ECAPA-TDNN, with a further relative improvement of 12.3 percent over the best baseline. Analysis of the learned attention weights confirmed the mechanism: when trained with the distance-based or SNR-based auxiliary matrix, the channels receiving the highest attention weights were consistently the microphones physically closest to the speaker.</p>
<p>The implications reach well beyond the laboratory. As homes and offices fill with voice-capable edge devices, the ability to fuse their microphones into a virtual array, without centralizing raw audio and without knowing where the devices sit, points toward privacy-preserving, robust voice interfaces that work as well from across the room as they do at arm&#8217;s length. The authors are candid about open questions, notably that the adjacency matrix design improves the graph attention mechanism but not the plain self-attention variant, and they envision future extensions using signal-to-interference ratios for multi-speaker scenes and estimated SNR when distances are unknown. For now, the study demonstrates a compelling principle: when microphones learn to collaborate as a graph, the sum of many imperfect ears becomes a remarkably sharp listener.</p>
<p><strong>Subject of Research:</strong> Far-field speaker verification using spatial-temporal graph attention networks with ad-hoc microphone arrays</p>
<p><strong>Article Title:</strong> Edge-collaborative multi-channel speaker verification via spatial-temporal graph with ad-hoc microphone arrays</p>
<p><strong>Article References:</strong> Chen, Y., Liang, C., Chen, S., Feng, L., Zhu, B., Zhang, C., &amp; Zhang, X.-L. (2025). Edge-collaborative multi-channel speaker verification via spatial-temporal graph with ad-hoc microphone arrays. <em>Vicinagearth, 2</em>(1), Article 12. <a href="https://doi.org/10.1007/s44336-025-00023-y" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00023-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00023-y" rel="noopener noreferrer">10.1007/s44336-025-00023-y</a></p>
<p><strong>Keywords:</strong> speaker verification, ad-hoc microphone arrays, graph attention networks, far-field speech processing, edge computing, AI Flow, channel selection, spatial-temporal graph neural network, multi-channel audio, equal error rate, smart home, informed machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213107</post-id>	</item>
		<item>
		<title>Robots, Tags and Apps: New Review Maps the Technology Helping Us Find Lost Household Items</title>
		<link>https://scienmag.com/robots-tags-and-apps-new-review-maps-the-technology-helping-us-find-lost-household-items/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 19:14:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[aging in place]]></category>
		<category><![CDATA[Assistive Technology]]></category>
		<category><![CDATA[assistive technology for object location]]></category>
		<category><![CDATA[Bluetooth tags]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[dementia]]></category>
		<category><![CDATA[dementia care and daily living aids]]></category>
		<category><![CDATA[gaps in development of household object locator systems]]></category>
		<category><![CDATA[home robots]]></category>
		<category><![CDATA[impact of lost item detection on caregiver burden]]></category>
		<category><![CDATA[innovations in home object retrieval devices]]></category>
		<category><![CDATA[memory aids]]></category>
		<category><![CDATA[object retrieval]]></category>
		<category><![CDATA[privacy]]></category>
		<category><![CDATA[privacy concerns in home tracking technology]]></category>
		<category><![CDATA[real-world testing of household object tracking tools]]></category>
		<category><![CDATA[scoping review]]></category>
		<category><![CDATA[smart home]]></category>
		<category><![CDATA[smart home object tracking devices]]></category>
		<category><![CDATA[systematic review of assistive tech for cognitive decline]]></category>
		<category><![CDATA[technological solutions for elderly independence]]></category>
		<category><![CDATA[usability]]></category>
		<category><![CDATA[usability challenges in assistive object location systems]]></category>
		<category><![CDATA[wearable and app-based object finders]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197756</guid>

					<description><![CDATA[A new scoping review of 23 studies finds that robots, smart tags and wearable cameras can help people find misplaced household objects, but warns that usability, privacy and real-world testing remain critically underdeveloped.]]></description>
										<content:encoded><![CDATA[<p>Misplacing your keys, glasses, wallet or phone is one of those small domestic frustrations that almost everyone knows intimately. But for millions of older adults, and especially for people living with dementia, losing track of everyday objects is far more than an annoyance. It can erode independence, trigger distress, burden family caregivers, and even serve as an early warning sign of cognitive decline. A new scoping review published in the Journal of Ambient Intelligence and Humanized Computing has now mapped, for the first time in a systematic way, the landscape of technologies designed to help people locate misplaced objects in their own homes — and the picture it paints is one of enormous promise undermined by glaring gaps in usability, privacy and real-world testing.</p>
<p>The review, led by Bing Ye of the University of Toronto&#8217;s Department of Occupational Science and Occupational Therapy, together with James McDonough, Cynthia Chui and Alex Mihailidis, and spanning institutions including the KITE Toronto Rehabilitation Institute, McMaster University and University Health Network, set out to answer a deceptively simple question: what technologies exist to help people find things they have lost at home, and how well do they actually work for the people who need them? An information specialist conducted a broad search across seven major databases — MEDLINE, Embase, Web of Science Core Collection, Compendex, Inspec, IEEE Xplore and the ACM Digital Library — ultimately including 23 records in the review.</p>
<p>What the team found is that the field is dominated by one approach above all others: robot-assisted technology. Home robots equipped with cameras and computer vision systems can be asked, in natural language, to search a living space for a missing item, recognize it visually, and either retrieve it or guide the user to it. Some systems go further, building what researchers call episodic memory models — computational analogues of the way humans remember where they last saw an object. A companion robot that observes its user placing a phone on a kitchen counter, for example, can later answer the question &#8220;where is my phone?&#8221; by recalling that stored observation. Other robotic platforms combine semantic localization with conversational skills, allowing an older adult to simply ask for help and receive a spoken answer rather than interacting with a screen.</p>
<p>Beyond robots, the review catalogued a diverse ecosystem of tagging and sensing technologies. Bluetooth Low Energy tags attached to frequently misplaced items allow a robot or smartphone to home in on the object&#8217;s radio signal even when it is hidden from view inside a drawer or under a cushion. Radio-frequency identification and Zigbee-based systems offer similar functionality with passive or low-power tags. Acoustic approaches such as ChirpTracker use sound signals and a single smartphone to pinpoint the precise location of a tagged object, while HyperEar demonstrated that indoor remote object finding can be achieved with audio cues alone. Wearable camera systems like GO-Finder take a different tack entirely: instead of requiring users to tag their belongings, the device passively observes hand-held object interactions and builds a searchable visual history of where items were last seen.</p>
<p>The way these systems communicate their findings to users emerged as a crucial design dimension. The most common retrieval feedback modality was visual, but the most common combination paired visual and audio feedback — for instance, a robot that points to or photographs the found object while announcing its location verbally. Multimodal feedback matters because users vary widely in sensory ability, cognitive status and personal preference, and the review&#8217;s usability findings confirmed that people strongly prefer systems offering multiple feedback channels and personalization options. Performance accuracy also proved central: a retrieval system that fails to find the object, or worse, reports a wrong location, quickly loses the trust of its user.</p>
<p>Yet the review&#8217;s most striking conclusions concern what the technology has not yet achieved. Retrieval technology, the authors conclude, remains in its early phase of development, which signals rich research opportunities but also means most systems never leave the laboratory. Usability and privacy, the two factors most likely to determine whether real people adopt these devices, are simply not yet researchers&#8217; priorities. The usability evidence that does exist points to the importance of training support — older adults need structured onboarding to gain confidence with new devices — and to the value of designs that accommodate prior experience and age-related changes in perception and cognition. The authors argue pointedly that usability education should begin early in university curricula, so that the next generation of engineers and designers internalizes the critical role of usability before they ever ship a product.</p>
<p>Privacy looms as an equally urgent concern. Many of the most capable retrieval systems rely on continuous camera monitoring of the home environment, raising obvious questions about surveillance, data storage and consent — particularly sensitive for people with dementia who may not be able to give informed permission. The review identified suggestions for privacy preservation drawn from the literature, including a focus on protecting algorithms rather than raw data, distributed approaches to artificial intelligence that keep processing on local edge devices, and careful attention to the documented privacy and security concerns that shape whether older adults accept Internet of Things technologies in their homes at all. User-centered design processes such as the co-conception approach used in the TROUVE project, which developed a tracking device for older adults with cognitive impairment alongside its intended users, offer one model for reconciling capability with acceptability.</p>
<p>The review also issues a clear methodological challenge to the field: longitudinal studies conducted in real homes. Most existing evaluations are short, laboratory-based demonstrations with small samples, which cannot reveal how systems perform over weeks and months in cluttered, dynamically changing domestic environments, nor how users&#8217; skills, trust and frustration evolve over time. Moving emerging technologies beyond the lab and into the real world is described as essential. The authors further highlight the essential role of policymakers, arguing that collaboration between researchers and policymakers during technology design and development — not after products are finished — is needed to address equitable access to assistive technology, a concern echoed by global health organizations and by Canadian policy research on access to assistive devices.</p>
<p>The stakes of this research are grounded in a substantial clinical literature. Misplacing objects is documented as one of the earliest and most prevalent symptoms of Alzheimer&#8217;s disease, and studies of people with mild to moderate Alzheimer&#8217;s have characterized the symptom in detail through clinical trials and online tracking tools. Research on normal aging shows that memory for item-location associations declines with age even in healthy adults, and surveys of psychogeriatric nursing home residents and their caregivers confirm that losing items is a persistent daily challenge. Caregiver burden, hoarding and hiding behaviors, and the emotional distress associated with lost belongings all amplify the human cost. Because the population over 65 is growing rapidly worldwide, technologies that preserve independence for people with memory deficits have enormous potential public health value.</p>
<p>Ultimately, the review is both a progress report and a call to arms. It demonstrates that the technical building blocks — object recognition, indoor localization, human-robot interaction, wearable sensing and voice interfaces powered by large language models — have advanced remarkably. Systems can now find keys hidden in drawers, answer spoken questions about an object&#8217;s whereabouts, and learn from observing daily routines. What remains missing is the human-centered engineering that turns these demonstrations into dependable daily companions: rigorous usability evaluation, privacy by design, long-term home trials, and policy frameworks that ensure the benefits reach the older adults and people with dementia who stand to gain the most. The authors advocate sustained attention and ongoing research in this field, arguing that the seemingly mundane act of helping someone find their glasses may prove to be one of the most meaningful applications of ambient intelligence in the aging home.</p>
<p><strong>Subject of Research:</strong> Assistive technologies for locating misplaced household objects, particularly for older adults and people with dementia</p>
<p><strong>Article Title:</strong> Help me find it!: a scoping review on technology assisting in locating misplaced objects in a home</p>
<p><strong>Article References:</strong> Ye, B., McDonough, J., Chui, C., &amp; Mihailidis, A. (2026). Help me find it!: a scoping review on technology assisting in locating misplaced objects in a home. <em>Journal of Ambient Intelligence and Humanized Computing</em>. <a href="https://doi.org/10.1007/s12652-026-05122-2" rel="noopener noreferrer">https://doi.org/10.1007/s12652-026-05122-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12652-026-05122-2" rel="noopener noreferrer">10.1007/s12652-026-05122-2</a></p>
<p><strong>Keywords:</strong> assistive technology, object retrieval, dementia, home robots, Bluetooth tags, usability, privacy, aging in place, computer vision, memory aids, scoping review, smart home</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197756</post-id>	</item>
	</channel>
</rss>
