<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>automatic speech recognition &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/automatic-speech-recognition/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 23:47:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>automatic speech recognition &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Speech Recognition and AI Language Models Face a Critical Test in Serving People With Language Impairments</title>
		<link>https://scienmag.com/speech-recognition-and-ai-language-models-face-a-critical-test-in-serving-people-with-language-impairments/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 23:47:38 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[accessibility]]></category>
		<category><![CDATA[accessible communication technology for neurological conditions]]></category>
		<category><![CDATA[AI bias]]></category>
		<category><![CDATA[aphasia]]></category>
		<category><![CDATA[Assistive Technology]]></category>
		<category><![CDATA[automatic speech recognition]]></category>
		<category><![CDATA[challenges in recognizing atypical prosody and mis]]></category>
		<category><![CDATA[communication disorders]]></category>
		<category><![CDATA[development of inclusive AI speech recognition systems]]></category>
		<category><![CDATA[evaluation of speech recognition systems for traumatic brain injury patients]]></category>
		<category><![CDATA[future of assistive communication tools for language-impaired individuals]]></category>
		<category><![CDATA[impact of disfluent speech on automatic transcription]]></category>
		<category><![CDATA[inclusive AI]]></category>
		<category><![CDATA[language impairments]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[limitations of voice-driven technology for people with speech disorders]]></category>
		<category><![CDATA[Nature Communications.]]></category>
		<category><![CDATA[performance of AI language models in speech transcription for impaired speech]]></category>
		<category><![CDATA[speech recognition accuracy]]></category>
		<category><![CDATA[speech recognition accuracy in aphasia and neurodegenerative diseases]]></category>
		<category><![CDATA[speech recognition challenges for language impairments]]></category>
		<category><![CDATA[word error rate]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=208867</guid>

					<description><![CDATA[New research in Nature Communications evaluates how well automatic speech recognition and large language models serve people with language impairments, revealing significant performance gaps and pathways to more inclusive assistive technology.]]></description>
										<content:encoded><![CDATA[<p>For millions of people living with language impairments, the promise of voice-driven technology has always been tantalizing yet frustratingly out of reach. Automatic speech recognition systems now transcribe everyday conversations with remarkable accuracy for typical speakers, and large language models can compose, summarize, and paraphrase text with fluency that would have seemed impossible a decade ago. But a growing body of research is asking a pointed question: do these tools actually work for the people who might benefit from them most? A new study published in Nature Communications examines exactly that, systematically assessing how well automatic speech recognition and large language models perform for individuals with language impairments, and the findings carry significant implications for the future of accessible communication technology.</p>
<p>The stakes could hardly be higher. Language impairments arising from aphasia after stroke, developmental language disorders, traumatic brain injury, neurodegenerative conditions such as Parkinson&#8217;s disease, and other neurological conditions affect communication in ways that standard speech technology was never designed to handle. Disfluent speech, word-finding pauses, mispronunciations, grammatical breakdowns, and atypical prosody can all scramble the acoustic and statistical patterns that modern recognition systems rely on. When a person with aphasia says a word haltingly or produces a neologism in place of the intended target, a recognition system trained predominantly on fluent, typical adult speech may simply fail, and that failure cascades downstream into every application that depends on accurate transcription.</p>
<p>The architecture of modern speech recognition helps explain why. State-of-the-art systems, including end-to-end neural models trained on tens of thousands of hours of audio, learn to map sound sequences to text by exploiting statistical regularities in their training data. Those regularities include not just phonetics but also the linguistic content of the speech itself. A recognizer hearing a garbled or incomplete utterance leans heavily on language-model priors to guess what was said, essentially autocorrecting toward plausible fluent speech. For typical speakers this bias improves accuracy, but for individuals with language impairments it can systematically distort what they actually said, replacing their intended words with the model&#8217;s own statistical expectations and effectively silencing their voice in favor of the algorithm&#8217;s prediction.</p>
<p>The research team evaluated how this plays out empirically by testing recognition systems on speech produced by individuals with language impairments and comparing performance against typical speech benchmarks. The results reveal a substantial performance gap. Word error rates climb steeply on impaired speech, and the errors are not randomly distributed: content words, which carry the semantic heart of a message, are disproportionately misrecognized or dropped, while function words are preserved. That asymmetry matters enormously, because a transcript that keeps the grammatical scaffolding but loses the meaningful content is nearly useless both for human readers and for any downstream language model asked to interpret, expand, or respond to the speaker&#8217;s intent.</p>
<p>On top of the transcription layer, the study examines large language models as assistive partners, the idea being that a person with aphasia might produce a fragmented utterance, which the recognizer transcribes imperfectly, and which the language model then attempts to repair, expand, or convert into a well-formed communicative act such as a text message or an email. In principle this pipeline could restore independence for people who struggle with everyday written and spoken communication. In practice, the researchers find that the pipeline inherits and sometimes amplifies the weaknesses of each stage. When the recognizer drops a content word, the language model cannot know what is missing, so it fluently completes the sentence with the wrong meaning, producing output that looks polished but betrays the speaker&#8217;s actual intent.</p>
<p>This problem of confident error propagation is one of the most consequential findings. Large language models are trained to produce coherent, plausible text, and they apply that objective whether or not the input transcript was accurate. The studies show that models asked to repair impaired-speech transcripts will happily generate grammatical, natural-sounding sentences that diverge from what the speaker meant. For assistive communication, such fluent fabrication is arguably worse than a raw, broken transcript, because conversation partners and caregivers may assume the polished output is trustworthy. The researchers emphasize that any deployed system must therefore build in transparency about uncertainty, letting users verify and correct the interpretation rather than presenting the model&#8217;s guess as fact.</p>
<p>Encouragingly, the work also identifies pathways toward improvement. Recognition accuracy for impaired speech improves substantially when systems are adapted, whether through fine-tuning on disordered speech data, personalizing acoustic models to an individual speaker&#8217;s voice and error patterns, or injecting contextual information about the communicative setting. Similarly, language models perform better as assistive aids when they are constrained, prompted with information about the speaker&#8217;s typical vocabulary and communication goals, or paired with interactive correction loops in which the user can confirm or reject candidate interpretations. These findings suggest that the technology is not fundamentally unsuited to this population, but that off-the-shelf deployment is. Deliberate, user-centered engineering is required to close the gap.</p>
<p>The study also raises important questions about evaluation methodology in the field. Benchmarks that dominate speech technology research contain almost no disordered speech, so headline accuracy figures say little about performance for this population. The authors argue for including individuals with language impairments in dataset collection, reporting performance stratified by speaker characteristics and impairment severity, and evaluating assistive pipelines end to end rather than optimizing transcription accuracy in isolation. A system that achieves slightly worse word error rates but preserves meaning and supports successful repair might serve users far better than one optimized purely for a conventional benchmark metric. Aligning evaluation with real communicative outcomes is presented as a necessary step for the field.</p>
<p>Beyond the technical conclusions, the research lands at a moment of intense public debate about artificial intelligence and accessibility. Voice interfaces are becoming the default way people interact with phones, homes, vehicles, and services, and large language models are being embedded in virtually every communication tool. If these systems systematically fail people with language impairments, the accessibility divide will widen even as technology advances, turning everyday tasks that others take for granted into new barriers. Conversely, if the gaps identified here are addressed through inclusive data, careful system design, and genuine involvement of affected communities, the same technologies could deliver on their long-promised potential: restoring voice, autonomy, and connection to people whose communication abilities have been compromised by injury or disease. The study&#8217;s message is ultimately one of cautious optimism grounded in rigor. The tools are powerful, the gaps are measurable, and the solutions are identifiable. What remains is the commitment to build speech and language AI not just for the typical speaker, but for the full diversity of human communication, ensuring that the next generation of voice technology leaves no voice behind.</p>
<p><strong>Subject of Research:</strong> Evaluation of automatic speech recognition and large language models for assisting individuals with language impairments</p>
<p><strong>Article Title:</strong> Assessing the use of automatic speech recognition and large language models for individuals with language impairments</p>
<p><strong>Article References:</strong> Xu, G., Yu, H., Wei, L., Liu, Y., Liu, D., Xu, C., Li, J., Abbasi, A., Xiong, J., Yu, X., Zheng, Z., Shi, Y., &amp; Qin, R. (2026). Assessing the use of automatic speech recognition and large language models for individuals with language impairments. <em>Nature Communications</em>. <a href="https://doi.org/10.1038/s41467-026-76677-z" rel="noopener noreferrer">https://doi.org/10.1038/s41467-026-76677-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41467-026-76677-z" rel="noopener noreferrer">10.1038/s41467-026-76677-z</a></p>
<p><strong>Keywords:</strong> automatic speech recognition, large language models, language impairments, aphasia, accessibility, assistive technology, speech recognition accuracy, word error rate, communication disorders, AI bias, inclusive AI, Nature Communications</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">208867</post-id>	</item>
		<item>
		<title>New Bilingual Speech Dataset Takes Aim at AI&#8217;s Weakest Spot: Code-Switching</title>
		<link>https://scienmag.com/new-bilingual-speech-dataset-takes-aim-at-ais-weakest-spot-code-switching/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:35:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automatic speech recognition]]></category>
		<category><![CDATA[automatic speech recognition challenges]]></category>
		<category><![CDATA[bilingual speech]]></category>
		<category><![CDATA[bilingual speech recognition]]></category>
		<category><![CDATA[code-switching]]></category>
		<category><![CDATA[code-switching dataset]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[DOTA-ME-CS corpus]]></category>
		<category><![CDATA[improving machine understanding of code-switching]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Mandarin]]></category>
		<category><![CDATA[Mandarin-English code-switching]]></category>
		<category><![CDATA[multilingual natural language processing]]></category>
		<category><![CDATA[multilingual speech processing]]></category>
		<category><![CDATA[open-source language datasets]]></category>
		<category><![CDATA[Paraformer]]></category>
		<category><![CDATA[phonetics]]></category>
		<category><![CDATA[SenseVoice]]></category>
		<category><![CDATA[speech dataset]]></category>
		<category><![CDATA[speech dataset for bilingual speakers]]></category>
		<category><![CDATA[speech recognition for code-switching]]></category>
		<category><![CDATA[transformer-based ASR models]]></category>
		<category><![CDATA[voice conversion]]></category>
		<category><![CDATA[Whisper]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203051</guid>

					<description><![CDATA[A new open dataset of 9300 Mandarin-English code-switched speech recordings, enhanced with AI-generated noise, speed and timbre changes, exposes how badly today's speech recognition models fail at bilingual conversation.]]></description>
										<content:encoded><![CDATA[<p>When bilingual speakers chat with one another, they rarely stay inside a single language. A sentence that begins in Mandarin may slip mid-phrase into English and back again, a behaviour linguists call code-switching. It is one of the most natural things multilingual people do, and one of the most unnatural things for machines to understand. Automatic speech recognition (ASR) systems, even the most powerful transformer-based models now in wide use, tend to stumble exactly at the point where one language hands off to another. A new openly available corpus called DOTA-ME-CS, short for Daily Oriented Text Audio Mandarin-English Code-Switching dataset, has been created to give researchers the fuel they need to close that gap.</p>
<p>The dataset, described in the Journal of Ambient Intelligence and Humanized Computing, contains 18.54 hours of audio spanning 9300 recordings produced by 34 bilingual participants, all of them fluent in both Mandarin and English. Unlike many earlier corpora, every single utterance in the collection involves code-switching. That design choice matters. Older resources such as SEAME, which stretches across roughly 190 hours, and TALCS, which covers 587 hours, contain substantial proportions of monolingual speech, meaning their effective supply of genuinely code-switched material is far smaller than their total length suggests. Some of those datasets, including the ASRU and TALCS corpora, are no longer publicly accessible at all, leaving the field with a shortage of usable, openly available benchmarks.</p>
<p>The construction of DOTA-ME-CS follows an unusual pipeline that blends large language model generation with human recording. The team used GPT-4o with carefully engineered prompts to produce scripted sentences across ten everyday scenarios: education, entertainment, environmental protection, food, health, home, life, pets, travel and work. Each prompt required the model to produce sentences that mimic daily conversational style, contain more English words than Mandarin words, and include at least one Mandarin word, with a dominant language assigned to every sentence. The authors justify this topic-anchored approach with a probabilistic argument: when a specific category is given, the probability of generating a relevant, high-quality sentence is higher than when the model is left to roam across all possible topics, which also reduces hidden cultural bias in the resulting scripts.</p>
<p>Human evaluators checked the generated scripts for grammatical problems and confirmed the presence of genuine switching, and a post-hoc naturalness study asked eight bilingual raters to score 100 randomly sampled stimuli on a five-point Likert scale. The average rating came in at 4.12 with a standard deviation of 0.61, and the median was 4.0, indicating that bilingual listeners generally perceive the generated sentences as natural, though the authors acknowledge that occasional awkwardness is an inherent limitation of LLM-based text generation. Bilingual volunteers, mostly college students with academic backgrounds in China and the United Kingdom, including roughly eighteen participants based at Imperial College London, then recorded the scripts on their own laptops or smartphones as 16-bit WAV files in quiet indoor settings. Recordings that failed basic quality checks were rejected and re-recorded, and accepted files were peak-normalised to 3 dBFS for consistent loudness.</p>
<p>The participants&#8217; linguistic profiles were deliberately diverse. Twelve reported English dominance, eighteen reported Mandarin dominance and four identified as balanced bilinguals, and the mix of speakers from China and the United Kingdom ensures that the corpus spans multiple varieties of English and second-language accents. An average of 2.22 switching points per utterance, combined with a broad part-of-speech distribution across both languages, means the recordings capture switching at varied syntactic positions and grammatical categories rather than concentrating it in a single predictable spot. The dataset also includes longer recordings of roughly 10 to 15 seconds and about 100 words, an intentional response to the weakness current ASR models show on extended speech, where dependencies and critical information can be lost.</p>
<p>What sets the corpus apart most sharply is its AI-driven augmentation. Because human recordings were captured in quiet conditions at a normal pace, the researchers used the Librosa audio library to modify a random subset of clips. Playback speed was adjusted to 0.75x, 0.5x, 1.25x, 1.5x and 2x, with 200 recordings modified at each setting. Five categories of background noise, drawn from highways, war, natural sounds, white noise and playground environments, were added at 200 recordings per type, with the noise pitch scaled to 0.8x to account for the Lombard effect, the well-documented tendency of speakers to raise their vocal intensity in noisy surroundings. In addition, five AI-generated timbres, two male and three female, replaced the original voices in 200 recordings each, simulating both voice conversion deepfake scenarios and privacy-preserving speech processing conditions.</p>
<p>The accompanying data analysis is unusually thorough. Phoneme distributions catalogued with the International Phonetic Alphabet reveal that Mandarin contributes more multisyllabic pronunciations, tones, aspirated consonants and palatalised sounds, while English favours single-syllable vowels; shared features such as the sounds t and a suggest cues that future switching-detection models could exploit. Physical measurements show a frame rate of 46.7 kHz and a sample width of 2.0 bytes, meeting professional audio standards, with an average maximum pitch of 3617.13 Hz and an average per-recording pitch of 535.04 Hz. Formant frequencies computed through Linear Predictive Coding yield averages of 804.02 Hz, 4419.15 Hz and 7549.48 Hz for F1, F2 and F3, pointing to mid-to-low and front vowels, reduced lip rounding and retroflex consonants. The measured speaking rate of 2.05 words per second, with a standard deviation of 0.49, sits squarely within the typical range for conversational read speech, supporting the corpus&#8217;s ecological validity.</p>
<p>To establish baseline performance, the team benchmarked three pre-trained ASR models on the full dataset without fine-tuning: Whisper large-v3, a 1.5-billion-parameter Transformer encoder-decoder trained on 680,000 hours of multilingual audio; SenseVoice Small, a lightweight hybrid CTC-attention model of roughly 200 million parameters that also performs language identification, emotion detection and audio event detection; and Paraformer, a non-autoregressive Transformer with 220 million parameters optimised for Mandarin and built on continuous integrate-and-fire alignment. Paraformer and SenseVoice achieved the best results, with Whisper slightly behind, and a word error rate of 0.177 alongside a character error rate of 0.176. Compared against the publicly available ASCEND corpus, the models produced higher error rates on DOTA-ME-CS, which the authors read not as a defect but as evidence that their dataset poses a more demanding and therefore more useful benchmark.</p>
<p>Case studies expose exactly where today&#8217;s systems fail. In one example embedding the Chinese noun 菠萝包, meaning pineapple bun, inside an English sentence, Paraformer and SenseVoice produced garbled outputs such as bullleball and boobao, while Whisper translated the phrase correctly but destroyed the mixed-language structure by refusing to preserve the code-switch itself, a translation bias that also appeared on noise-free recordings. At double playback speed, Paraformer&#8217;s output strayed entirely from the intended meaning, SenseVoice descended into repetitions such as Miamiami, and Whisper mangled Florida lifestyle into Forensic Lifestyle. Altered timbres degraded English recognition across the board, and all three models performed worst at 2.0x speed. Whisper proved the most resilient to background noise, while Paraformer was the most sensitive to voice changes.</p>
<p>The implications reach beyond engineering. Reliable code-switching recognition would make automated transcription, voice assistants and translation systems far more useful for the hundreds of millions of people who live their linguistic lives between two languages, reducing a bias that currently disadvantages non-monolingual speakers. The authors stress that the study received ethical approval from Imperial College London&#8217;s ethical board, that all participants gave informed consent, and that privacy protections were maintained throughout. They are candid about limitations: funding constrained the scale of the corpus, and scripted reading, while reducing transcription errors and ethical risks, sacrifices the spontaneity of natural conversation. Even so, their deliberate trade-off of quality and coverage over raw scale gives the community something it has lacked, a fully public, exclusively code-switching Mandarin-English benchmark with baseline scores, rigorous acoustic analysis and data, code and recordings all released through a public repository. As speech technology races toward multilingual ubiquity, datasets like this one may determine whether the next generation of voice interfaces can finally follow the way people actually talk.</p>
<p><strong>Subject of Research:</strong> A Mandarin-English code-switching speech dataset for advancing automatic speech recognition research.</p>
<p><strong>Article Title:</strong> DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset</p>
<p><strong>Article References:</strong> Li, Y., Wei, Z., Yu, H., Xue, J., Zhou, H., &amp; Schuller, B. W. (2026). DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset. <em>Journal of Ambient Intelligence and Humanized Computing</em>. <a href="https://doi.org/10.1007/s12652-026-05119-x" rel="noopener noreferrer">https://doi.org/10.1007/s12652-026-05119-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12652-026-05119-x" rel="noopener noreferrer">10.1007/s12652-026-05119-x</a></p>
<p><strong>Keywords:</strong> code-switching, automatic speech recognition, Mandarin, bilingual speech, speech dataset, Whisper, Paraformer, SenseVoice, data augmentation, phonetics, large language models, voice conversion</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203051</post-id>	</item>
		<item>
		<title>New Open-Source Toolkit Turns Any Speech Recording Into Acoustic-Phonetic Data</title>
		<link>https://scienmag.com/new-open-source-toolkit-turns-any-speech-recording-into-acoustic-phonetic-data/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:46:46 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[acoustic analysis of conversational speech]]></category>
		<category><![CDATA[acoustic-phonetic analysis]]></category>
		<category><![CDATA[automatic speech recognition]]></category>
		<category><![CDATA[Behavior Research Methods]]></category>
		<category><![CDATA[field-recorded speech analysis]]></category>
		<category><![CDATA[forced alignment]]></category>
		<category><![CDATA[fricative spectral moments]]></category>
		<category><![CDATA[natural speech data processing]]></category>
		<category><![CDATA[naturalistic speech]]></category>
		<category><![CDATA[open-source acoustic-phonetic analysis]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[phonetic alignment software]]></category>
		<category><![CDATA[phonetic measurement tools]]></category>
		<category><![CDATA[sociophonetics]]></category>
		<category><![CDATA[speaker diarization]]></category>
		<category><![CDATA[speaker-labeled speech data]]></category>
		<category><![CDATA[speech analysis toolkit]]></category>
		<category><![CDATA[speech data extraction from multimedia]]></category>
		<category><![CDATA[speech perception]]></category>
		<category><![CDATA[speech research automation]]></category>
		<category><![CDATA[speech transcription and alignment]]></category>
		<category><![CDATA[spontaneous speech analysis]]></category>
		<category><![CDATA[voice onset time]]></category>
		<category><![CDATA[vowel formants]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196547</guid>

					<description><![CDATA[An open-source pipeline called TAPA automates transcription, speaker diarization, forced alignment, and acoustic-phonetic measurement of naturalistic speech, validated against expert annotation on the 2016 U.S. presidential debate.]]></description>
										<content:encoded><![CDATA[<p>For decades, the speech sciences have faced an uncomfortable paradox: some of the most important questions about how humans actually talk can only be answered with natural, spontaneous speech, yet the data that researchers most easily collect and analyze is almost always carefully scripted laboratory audio. A newly published open-source pipeline called the Toolkit for Acoustic–Phonetic Analysis, or TAPA, promises to change that balance. In a paper in Behavior Research Methods, a team led by Ethan Kutlu of the University of Iowa describes a single integrated system that takes raw audio from almost any source — a YouTube video, a podcast, an interview recording made in the field — and converts it into speaker-labeled, phonetically aligned, acoustically measured data, all without requiring users to write custom code.</p>
<p>The motivation behind the toolkit is rooted in a long-standing tension within the field. Laboratory speech, produced in word-reading or sentence-repetition tasks, gives experimenters precise control over variables and supports repeatability, which is why it has anchored theoretical work in speech perception and production for generations. But an accumulating body of evidence shows that listeners and speakers dynamically shape one another in everyday conversation, adapting to accents, talkers, and social-linguistic associations in ways that rigid laboratory tasks cannot capture. The classic observer&#8217;s paradox, articulated by William Labov in 1972, compounds the problem: people who know they are being recorded often adjust their speech, consciously or not, making genuinely naturalistic samples scarce and difficult to obtain.</p>
<p>The logistical hurdles do not end there. Field recordings demand extensive preparation, community relationships, and careful attention to the researcher&#8217;s positionality. Reading tasks systematically exclude speakers — such as many heritage bilinguals — who are conversationally fluent but may lack literacy in the language being studied. Funding constraints make it increasingly difficult to build and share open-access speech corpora, a burden that falls hardest on graduate students, early-career scholars, and researchers working on minoritized language communities. The authors argue that these barriers do not merely slow research down; they distort the theoretical record, because dominant models of speech perception and production were built on idealized, read speech and often treat real-world linguistic diversity as noise rather than as the phenomenon itself.</p>
<p>TAPA&#8217;s answer is to stitch together a chain of proven open-source components into one Python pipeline that can be installed with a single command and run either locally or in Google Colab, where free GPU access removes the need for any local setup. The journey begins with a YouTube audio downloader built on yt-dlp, which pulls the audio stream from any standard video URL and converts it to an MP3 file; users can equally supply their own local audio. Transcription is handled by OpenAI&#8217;s Whisper, which produces word-level transcripts with precise start and end timestamps. Speaker diarization — deciding who spoke when — relies on two pretrained neural models: Silero VAD, which detects stretches of actual speech, and Resemblyzer, which maps each detected segment to a 256-dimensional voice embedding in which clips from the same speaker cluster together. Those embeddings are then grouped to assign every transcribed word to its speaker.</p>
<p>From there, the pipeline turns to fine-grained phonetic measurement. The Montreal Forced Aligner produces phone-level boundaries linking the audio to the transcript, using an American English acoustic model with a dictionary-based proportional-timing fallback when alignment is unavailable. Vowel formants — the resonant frequencies F1 and F2 that define vowel identity — are extracted with Praat-parselmouth, using the Burg algorithm over the central portion of each vowel after trimming the edges to suppress coarticulation, with the maximum formant ceiling adjusted adaptively to each speaker&#8217;s pitch. Stop consonant voice onset time, the interval between a stop&#8217;s burst release and the onset of voicing, is measured by Dr. VOT, a neural model that predicts burst and voicing boundaries directly from the waveform. Fricative spectral moments — center of gravity, standard deviation, skewness, and kurtosis — are computed over the middle 80 percent of each fricative interval. The output is a set of CSV or JSON files containing per-token measurements and per-speaker summaries.</p>
<p>To demonstrate what the pipeline can do, the team pointed it at a demanding piece of real-world audio: the 90-minute 2016 U.S. presidential debate between Hillary Clinton and Donald Trump, complete with audience noise, podium reverberation, overlapping talkers, and the performative speech style of political theater. From that single recording, TAPA extracted nearly 33,000 acoustic segments — 19,855 vowels, 6,547 stops, and 6,372 fricatives. The case study focused on the two main candidates, yielding more than 7,000 monophthong tokens each for vowel-space analysis, alongside thousands of stop and fricative tokens per speaker. Plotting the F1-by-F2 vowel spaces revealed that Clinton&#8217;s vowel categories occupied a larger acoustic area than Trump&#8217;s, and comparing both speakers&#8217; category means against the classic Hillenbrand reference norms showed the expected centralization of vowels in continuous speech relative to citation-form productions.</p>
<p>Crucially, the team did not simply trust the automated output. A phonetically trained research assistant hand-coded stratified samples of vowels, stops, and fricatives from the debate audio, blind to the pipeline&#8217;s results, and the two sets of measurements were compared token by token. Vowel formants emerged as the pipeline&#8217;s strongest suit: agreement with expert measurements was high, with correlations of 0.89 for F1 and 0.87 for F2, mean absolute errors of 37 and 125 hertz respectively, and no meaningful bias for F1. Fricative measurements showed a more differentiated picture. Spectral standard deviation agreed strongly with hand coding overall, and center of gravity tracked the human reference closely for the sibilant fricatives — correlations of 0.87 for /s/ and 0.94 for /ʃ/ — while the spectrally indistinct non-sibilants /f/ and /θ/ agreed far more weakly, a pattern consistent with long-standing findings that these sounds are hard to classify from spectral moments alone. Clinton&#8217;s /s/ center of gravity sat roughly 900 hertz above Trump&#8217;s, echoing documented patterns of gender-linked variation in sibilant production.</p>
<p>Stop consonants told a more cautionary tale. Although the aggregate voice-onset-time means for voiceless stops /p/, /t/, and /k/ landed squarely in the expected American English long-lag range of roughly 69 to 83 milliseconds, the per-token picture was far less reassuring. The Dr. VOT model, which was trained on isolated single-word productions with stop classes artificially rebalanced to a 50:50 positive-to-negative ratio, mislabeled 45 to 61 percent of voiceless onset stops as prevoiced — an outcome incompatible with English phonology — and its per-token magnitudes correlated essentially not at all with expert measurements (r = −0.04), diverging by 50 to 80 milliseconds on average. The authors attribute this to a training–deployment mismatch: continuous, noisy debate audio with a natural 95:5 ratio of positive to negative stops violates both of the model&#8217;s core training assumptions. Rather than hide the problem, they report it transparently, omit voiced stops from the case study entirely, and recommend that aggregate means be treated as descriptive rather than per-token measurements.</p>
<p>That candor reflects the paper&#8217;s broader stance on automation in science. The authors are explicit that TAPA is meant to aid researchers, not replace them: acoustic training must continue, research must be deployed and interpreted by humans, and the energy costs of training and fine-tuning large models warrant critical scrutiny. The pipeline inherits the biases of its components — the default configuration is American English-centric, diarization degrades on heavily overlapping speech, and forced alignment stumbles on out-of-vocabulary names — so the team recommends that every substantive deployment include a small hand-coded validation subsample from the researcher&#8217;s own corpus, and that publications report which component produced each measurement. With those safeguards in place, TAPA offers something the field has lacked: a free, modular, inspectable path from a web video or field recording to speaker-level, phone-level acoustic data at a scale that would once have required years of manual annotation, potentially opening naturalistic speech research to communities and questions that laboratory methods have long left on the periphery.</p>
<p><strong>Subject of Research:</strong> An open-source computational pipeline for automated acoustic-phonetic analysis of naturalistic multi-speaker speech recordings</p>
<p><strong>Article Title:</strong> Toolkit for acoustic–phonetic analysis of naturalistic speech data</p>
<p><strong>Article References:</strong> Kutlu, E., Peters, E., Tapanes, C., Chandio, S., &amp; Khalid, O. (2026). Toolkit for acoustic–phonetic analysis of naturalistic speech data. <em>Behavior Research Methods, 58</em>(10), Article 289. <a href="https://doi.org/10.3758/s13428-026-03174-y" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03174-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03174-y" rel="noopener noreferrer">10.3758/s13428-026-03174-y</a></p>
<p><strong>Keywords:</strong> acoustic-phonetic analysis, naturalistic speech, speaker diarization, forced alignment, vowel formants, voice onset time, fricative spectral moments, open-source software, speech perception, sociophonetics, automatic speech recognition, Behavior Research Methods</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196547</post-id>	</item>
	</channel>
</rss>
