<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>machine learning for speech progress &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/machine-learning-for-speech-progress/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 18:17:41 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>machine learning for speech progress &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak</title>
		<link>https://scienmag.com/smart-ring-and-ai-automate-speech-therapy-data-for-millions-who-cannot-speak/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 18:17:41 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in speech therapy]]></category>
		<category><![CDATA[AI-assisted communication intervention]]></category>
		<category><![CDATA[attribution]]></category>
		<category><![CDATA[augmentative and alternative communication]]></category>
		<category><![CDATA[Augmentative and alternative communication (AAC)]]></category>
		<category><![CDATA[automated therapy assessment]]></category>
		<category><![CDATA[clinical metrics]]></category>
		<category><![CDATA[convolutional neural network]]></category>
		<category><![CDATA[digital health innovation in speech therapy]]></category>
		<category><![CDATA[ECAPA-TDNN]]></category>
		<category><![CDATA[inertial measurement unit]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for speech progress]]></category>
		<category><![CDATA[sensor-based data collection for speech therapy]]></category>
		<category><![CDATA[smart ring]]></category>
		<category><![CDATA[smart ring for speech data collection]]></category>
		<category><![CDATA[speech therapy automation]]></category>
		<category><![CDATA[speech therapy data analysis]]></category>
		<category><![CDATA[speech-generating devices]]></category>
		<category><![CDATA[speech-language pathology]]></category>
		<category><![CDATA[speech-language pathology technology]]></category>
		<category><![CDATA[voice activity detection]]></category>
		<category><![CDATA[wearable sensing]]></category>
		<category><![CDATA[wearable technology for communication]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217966</guid>

					<description><![CDATA[Researchers have developed a smart ring and machine learning system that automatically attributes AAC device speech to the correct user with 94.81 percent accuracy, eliminating the need for video recording and removing all hardware burden from the AAC user.]]></description>
										<content:encoded><![CDATA[<p>For the roughly 97 million people worldwide who rely on augmentative and alternative communication (AAC) devices to speak, the technology itself is only half the story. The other half is the painstaking work of therapy: teachers and caregivers model language on the user&#8217;s speech-generating tablet, pressing symbols to demonstrate vocabulary and syntax while the user learns to follow along. Measuring whether that intervention actually works has long required something almost absurdly laborious — video cameras recording entire sessions, and human annotators cross-referencing hours of footage against device logs to determine who pressed which symbol. A new study published in Machine Learning with Applications proposes a radically simpler approach: a smart ring worn only by the communication partner, paired with a smartphone microphone and a machine learning pipeline that automates the entire data collection, attribution, and analysis process.</p>
<p>The research team, led by Jiayu Lu and colleagues at the University of New Hampshire, identified a fundamental bottleneck in AAC clinical practice. Speech-language pathologists need quantitative metrics such as selection rate and type-token ratio to assess intervention progress, but calculating these metrics requires knowing whether each audio output from the device was triggered by the AAC user or by the communication partner who is modeling language. A national survey of school-based speech-language pathologists cited in the study found that the time required for transcription and attribution was one of the major barriers to routine use of language sample analysis. Without large-scale, accurately attributed data, evidence-based AAC intervention remains more aspiration than reality.</p>
<p>Previous attempts at automation have stumbled on this attribution problem. Automated data logging protocols such as the Language Activity Monitor record every device activation with a timestamp, and automatic speech recognition systems have been explored for transcription, but neither can determine who triggered the activation. Computer-vision solutions, such as a smart-glasses system that tracks hand gestures with egocentric cameras, face their own obstacles: continuous video recording raises privacy and consent concerns in homes and classrooms, hand occlusion can hide the user&#8217;s gestures, changing lighting conditions degrade performance, and requiring the AAC user to wear additional hardware can be uncomfortable or distracting for people with sensory sensitivities.</p>
<p>The smart ring system sidesteps all of these issues with an elegant asymmetry. Only the communication partner wears hardware — a 3D-printed ring housing two circuit boards, a 6-axis inertial measurement unit sampling acceleration and gyroscope data at 120 Hz, an ESP32-H2 microcontroller transmitting via Bluetooth Low Energy, and a 40 mAh battery. The AAC user wears nothing at all, eliminating any physical or cognitive burden. Meanwhile, the smartphone&#8217;s microphone records the ambient audio at 44.1 kHz. The attribution logic exploits a simple temporal fact: every AAC audio output is triggered by a screen tap. If a recognized AAC audio event coincides with a detected tap from a ring-wearing partner, the audio belongs to that partner; if not, it belongs to the user.</p>
<p>The audio processing pipeline works in two stages. First, a voice activity detection module resamples incoming audio to 16 kHz, applies an 80 Hz high-pass filter to remove low-frequency rumble, and uses a real-time detection model to isolate speech segments, retaining only those longer than 150 milliseconds and padding each segment by 50 milliseconds to preserve initial and final phonemes. Second, each isolated segment is passed through a pre-trained ECAPA-TDNN speaker-recognition network that converts it into a 192-dimensional embedding vector capturing its acoustic characteristics. A support vector machine then performs the binary classification: AAC device audio versus everything else, including human speech. Because the system only needs to identify device-generated audio rather than transcribe individual voices, the task is far more tractable than full speech recognition.</p>
<p>Tap detection posed a subtler engineering challenge. The team discovered that the delay between a screen tap and the resulting audio is not fixed — it follows a bimodal distribution, with one cluster averaging 0.30 seconds and another 0.61 seconds, likely reflecting whether the device retrieves audio from memory or from disk. Rather than assuming a fixed offset, the algorithm searches a window spanning 0.1 to 1.0 seconds before each detected audio onset, looking for the peak in jerk — the rate of change of acceleration. Jerk is the key discriminator: smooth arm movements produce broad acceleration curves, while a physical tap against a screen creates an almost instantaneous collision signature. Once the jerk peak and an adjacent local minimum are located, a 20-sample window of IMU data is extracted and fed into a one-dimensional convolutional neural network that classifies it as tap or non-tap.</p>
<p>The results from a pilot study with 18 healthy adults simulating AAC interactions were striking. Voice activity detection caught 98.89 percent of speech events. The audio classifier achieved 99.83 percent accuracy on the training dataset and 99.64 percent on the test dataset under rigorous leave-one-subject-out and leave-one-group-out validation schemes designed to prevent data leakage. The full cascaded pipeline — voice detection, audio classification, tap detection, and attribution — maintained a cumulative end-to-end accuracy of 94.81 percent across 1,620 test events. Notably, a model trained only on normal tap patterns performed comparably to one trained on light, normal, and firm taps, and errors were distributed almost evenly between the two possible directions of misattribution, meaning the system neither systematically inflates nor deflates the user&#8217;s measured performance.</p>
<p>Beyond attribution, the system automatically computes three clinically meaningful metrics. Selection rate measures the user&#8217;s true information transfer speed in bits per second, isolated from the partner&#8217;s modeling inputs that artificially inflate conventional log-based counts. A newly proposed modeling balance index captures the percentage of total device activations attributable to the user, allowing clinicians to track how the balance of use shifts across sessions and adjust the amount of partner modeling accordingly. And a user-specific type-token ratio measures genuine lexical diversity by filtering out the advanced vocabulary that teachers model during sessions — a correction that prevents artificial inflation of the user&#8217;s apparent vocabulary breadth. Mean absolute percentage errors for these metrics ranged from 6.66 to 7.39 percent, though the authors caution that with only six test groups, these intervals represent uncertainty estimates from a pilot dataset rather than clinical-grade measurement accuracy.</p>
<p>The design also delivers a remarkable computational efficiency gain. Because tap detection is triggered only when an AAC audio event is detected, and each activation processes just a single 20-sample window, the system analyzed only a small fraction of the continuously recorded IMU stream — a 93.1 percent reduction in computational load compared with a traditional sliding-window approach. This efficiency matters for real-world deployment on consumer smartphones, where battery life and processing overhead constrain what wearable sensing systems can realistically sustain throughout a full therapy day.</p>
<p>Significant limitations remain before the system can leave the laboratory. The pilot involved healthy adults in a controlled environment with a vocabulary of only ten isolated words from a single device and synthesized voice; real interventions involve overlapping speech, variable tapping forces, ambient noise, phrases and sentences, and multiple devices. When caregivers speak over the device — as they often do, verbally confirming a word precisely as they activate it — the audio classifier may struggle, and the authors propose blind source separation algorithms and smartphone multi-microphone beamforming as future solutions. Only two communication partners were tested, and confidence-based conflict resolution was exercised in just four cases. Still, the core demonstration stands: a privacy-preserving, video-free, user-unburdened sensing platform that attributes every device activation with over 94 percent accuracy. If future work validates it in authentic clinical settings, the humble smart ring could transform AAC therapy from a data-starved discipline into a data-rich one, finally enabling the large-scale, evidence-based intervention research that millions of users with complex communication needs have been waiting for.</p>
<p><strong>Subject of Research:</strong> A smart-ring-based multimodal sensing and machine learning system for automated data collection, attribution, and analysis in augmentative and alternative communication intervention.</p>
<p><strong>Article Title:</strong> Machine learning-driven multimodal sensing for automated data collection, attribution, and analysis in augmentative and alternative communication</p>
<p><strong>Article References:</strong> Lu, J., Ghoreishi, N., Chen, S.-H. K., &amp; Chen, D. (2026). Machine learning-driven multimodal sensing for automated data collection, attribution, and analysis in augmentative and alternative communication. <em>Machine Learning with Applications, 26</em>, Article 101022. <a href="https://doi.org/10.1016/j.mlwa.2026.101022" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101022</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101022" rel="noopener noreferrer">10.1016/j.mlwa.2026.101022</a></p>
<p><strong>Keywords:</strong> augmentative and alternative communication, smart ring, wearable sensing, machine learning, inertial measurement unit, voice activity detection, convolutional neural network, speech-generating devices, attribution, speech-language pathology, ECAPA-TDNN, clinical metrics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217966</post-id>	</item>
	</channel>
</rss>
