Wednesday, September 30, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak

September 30, 2026
in Technology and Engineering
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 5 mins read
0
Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak

Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak

Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

For the roughly 97 million people worldwide who rely on augmentative and alternative communication (AAC) devices to speak, the technology itself is only half the story. The other half is the painstaking work of therapy: teachers and caregivers model language on the user’s speech-generating tablet, pressing symbols to demonstrate vocabulary and syntax while the user learns to follow along. Measuring whether that intervention actually works has long required something almost absurdly laborious — video cameras recording entire sessions, and human annotators cross-referencing hours of footage against device logs to determine who pressed which symbol. A new study published in Machine Learning with Applications proposes a radically simpler approach: a smart ring worn only by the communication partner, paired with a smartphone microphone and a machine learning pipeline that automates the entire data collection, attribution, and analysis process.

The research team, led by Jiayu Lu and colleagues at the University of New Hampshire, identified a fundamental bottleneck in AAC clinical practice. Speech-language pathologists need quantitative metrics such as selection rate and type-token ratio to assess intervention progress, but calculating these metrics requires knowing whether each audio output from the device was triggered by the AAC user or by the communication partner who is modeling language. A national survey of school-based speech-language pathologists cited in the study found that the time required for transcription and attribution was one of the major barriers to routine use of language sample analysis. Without large-scale, accurately attributed data, evidence-based AAC intervention remains more aspiration than reality.

Previous attempts at automation have stumbled on this attribution problem. Automated data logging protocols such as the Language Activity Monitor record every device activation with a timestamp, and automatic speech recognition systems have been explored for transcription, but neither can determine who triggered the activation. Computer-vision solutions, such as a smart-glasses system that tracks hand gestures with egocentric cameras, face their own obstacles: continuous video recording raises privacy and consent concerns in homes and classrooms, hand occlusion can hide the user’s gestures, changing lighting conditions degrade performance, and requiring the AAC user to wear additional hardware can be uncomfortable or distracting for people with sensory sensitivities.

The smart ring system sidesteps all of these issues with an elegant asymmetry. Only the communication partner wears hardware — a 3D-printed ring housing two circuit boards, a 6-axis inertial measurement unit sampling acceleration and gyroscope data at 120 Hz, an ESP32-H2 microcontroller transmitting via Bluetooth Low Energy, and a 40 mAh battery. The AAC user wears nothing at all, eliminating any physical or cognitive burden. Meanwhile, the smartphone’s microphone records the ambient audio at 44.1 kHz. The attribution logic exploits a simple temporal fact: every AAC audio output is triggered by a screen tap. If a recognized AAC audio event coincides with a detected tap from a ring-wearing partner, the audio belongs to that partner; if not, it belongs to the user.

The audio processing pipeline works in two stages. First, a voice activity detection module resamples incoming audio to 16 kHz, applies an 80 Hz high-pass filter to remove low-frequency rumble, and uses a real-time detection model to isolate speech segments, retaining only those longer than 150 milliseconds and padding each segment by 50 milliseconds to preserve initial and final phonemes. Second, each isolated segment is passed through a pre-trained ECAPA-TDNN speaker-recognition network that converts it into a 192-dimensional embedding vector capturing its acoustic characteristics. A support vector machine then performs the binary classification: AAC device audio versus everything else, including human speech. Because the system only needs to identify device-generated audio rather than transcribe individual voices, the task is far more tractable than full speech recognition.

Tap detection posed a subtler engineering challenge. The team discovered that the delay between a screen tap and the resulting audio is not fixed — it follows a bimodal distribution, with one cluster averaging 0.30 seconds and another 0.61 seconds, likely reflecting whether the device retrieves audio from memory or from disk. Rather than assuming a fixed offset, the algorithm searches a window spanning 0.1 to 1.0 seconds before each detected audio onset, looking for the peak in jerk — the rate of change of acceleration. Jerk is the key discriminator: smooth arm movements produce broad acceleration curves, while a physical tap against a screen creates an almost instantaneous collision signature. Once the jerk peak and an adjacent local minimum are located, a 20-sample window of IMU data is extracted and fed into a one-dimensional convolutional neural network that classifies it as tap or non-tap.

The results from a pilot study with 18 healthy adults simulating AAC interactions were striking. Voice activity detection caught 98.89 percent of speech events. The audio classifier achieved 99.83 percent accuracy on the training dataset and 99.64 percent on the test dataset under rigorous leave-one-subject-out and leave-one-group-out validation schemes designed to prevent data leakage. The full cascaded pipeline — voice detection, audio classification, tap detection, and attribution — maintained a cumulative end-to-end accuracy of 94.81 percent across 1,620 test events. Notably, a model trained only on normal tap patterns performed comparably to one trained on light, normal, and firm taps, and errors were distributed almost evenly between the two possible directions of misattribution, meaning the system neither systematically inflates nor deflates the user’s measured performance.

Beyond attribution, the system automatically computes three clinically meaningful metrics. Selection rate measures the user’s true information transfer speed in bits per second, isolated from the partner’s modeling inputs that artificially inflate conventional log-based counts. A newly proposed modeling balance index captures the percentage of total device activations attributable to the user, allowing clinicians to track how the balance of use shifts across sessions and adjust the amount of partner modeling accordingly. And a user-specific type-token ratio measures genuine lexical diversity by filtering out the advanced vocabulary that teachers model during sessions — a correction that prevents artificial inflation of the user’s apparent vocabulary breadth. Mean absolute percentage errors for these metrics ranged from 6.66 to 7.39 percent, though the authors caution that with only six test groups, these intervals represent uncertainty estimates from a pilot dataset rather than clinical-grade measurement accuracy.

The design also delivers a remarkable computational efficiency gain. Because tap detection is triggered only when an AAC audio event is detected, and each activation processes just a single 20-sample window, the system analyzed only a small fraction of the continuously recorded IMU stream — a 93.1 percent reduction in computational load compared with a traditional sliding-window approach. This efficiency matters for real-world deployment on consumer smartphones, where battery life and processing overhead constrain what wearable sensing systems can realistically sustain throughout a full therapy day.

Significant limitations remain before the system can leave the laboratory. The pilot involved healthy adults in a controlled environment with a vocabulary of only ten isolated words from a single device and synthesized voice; real interventions involve overlapping speech, variable tapping forces, ambient noise, phrases and sentences, and multiple devices. When caregivers speak over the device — as they often do, verbally confirming a word precisely as they activate it — the audio classifier may struggle, and the authors propose blind source separation algorithms and smartphone multi-microphone beamforming as future solutions. Only two communication partners were tested, and confidence-based conflict resolution was exercised in just four cases. Still, the core demonstration stands: a privacy-preserving, video-free, user-unburdened sensing platform that attributes every device activation with over 94 percent accuracy. If future work validates it in authentic clinical settings, the humble smart ring could transform AAC therapy from a data-starved discipline into a data-rich one, finally enabling the large-scale, evidence-based intervention research that millions of users with complex communication needs have been waiting for.

Subject of Research: A smart-ring-based multimodal sensing and machine learning system for automated data collection, attribution, and analysis in augmentative and alternative communication intervention.

Article Title: Machine learning-driven multimodal sensing for automated data collection, attribution, and analysis in augmentative and alternative communication

Article References: Lu, J., Ghoreishi, N., Chen, S.-H. K., & Chen, D. (2026). Machine learning-driven multimodal sensing for automated data collection, attribution, and analysis in augmentative and alternative communication. Machine Learning with Applications, 26, Article 101022. https://doi.org/10.1016/j.mlwa.2026.101022

Image Credits: AI Generated

DOI: 10.1016/j.mlwa.2026.101022

Keywords: augmentative and alternative communication, smart ring, wearable sensing, machine learning, inertial measurement unit, voice activity detection, convolutional neural network, speech-generating devices, attribution, speech-language pathology, ECAPA-TDNN, clinical metrics

Cite Scienmag News

Blake Davidson. (September 30, 2026). Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak. Scienmag. https://scienmag.com/smart-ring-and-ai-automate-speech-therapy-data-for-millions-who-cannot-speak/

Blake Davidson. "Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak." Scienmag, 30 September 2026, https://scienmag.com/smart-ring-and-ai-automate-speech-therapy-data-for-millions-who-cannot-speak/. Accessed 30 September 2026.

Blake Davidson. "Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak." Scienmag. September 30, 2026. https://scienmag.com/smart-ring-and-ai-automate-speech-therapy-data-for-millions-who-cannot-speak/

Tags: AI in speech therapyAI-assisted communication interventionattributionaugmentative and alternative communicationAugmentative and alternative communication (AAC)automated therapy assessmentclinical metricsconvolutional neural networkdigital health innovation in speech therapyECAPA-TDNNinertial measurement unitMachine learningmachine learning for speech progresssensor-based data collection for speech therapysmart ringsmart ring for speech data collectionspeech therapy automationspeech therapy data analysisspeech-generating devicesspeech-language pathologyspeech-language pathology technologyvoice activity detectionwearable sensingwearable technology for communication
Share26Tweet16
Previous Post

Electron Beam Mutagenesis Unlocks New Yield Potential in Field Corn

Next Post

AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug

Related Posts

AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug
Technology and Engineering

AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug

September 30, 2026
AI Agent Team Teaches Itself to Hack Binaries and Write Working Exploits
Technology and Engineering

AI Agent Team Teaches Itself to Hack Binaries and Write Working Exploits

September 30, 2026
Seven-Qubit Quantum Channel Teleports Four-Qubit Cluster States Under Supervision, Even in Noise
Technology and Engineering

Seven-Qubit Quantum Channel Teleports Four-Qubit Cluster States Under Supervision, Even in Noise

September 30, 2026
Photonic Lanterns Explained: Why a Simple Fiber Device Tames Fading in Free-Space Optical Links
Technology and Engineering

Photonic Lanterns Explained: Why a Simple Fiber Device Tames Fading in Free-Space Optical Links

September 30, 2026
New Dual-Stream Training Strategy Sharpens 3D LiDAR Segmentation for Autonomous Driving
Technology and Engineering

New Dual-Stream Training Strategy Sharpens 3D LiDAR Segmentation for Autonomous Driving

September 30, 2026
New AI Model Learns to Pick Its Own Evidence in the Fight Against Fake News
Technology and Engineering

New AI Model Learns to Pick Its Own Evidence in the Fight Against Fake News

September 30, 2026
Next Post
AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug

AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • AI Chatbots Redesign a Plant Molecule to Out-Bind a Gout Drug
  • Smart Ring and AI Automate Speech Therapy Data for Millions Who Cannot Speak
  • Electron Beam Mutagenesis Unlocks New Yield Potential in Field Corn
  • Older Donors, Warmer Organs: How the US Organ Recovery System Is Being Pushed to Its Limits

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading