<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>phonetics &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/phonetics/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 17:33:44 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>phonetics &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Personality traits fail to explain why listeners hear speech differently</title>
		<link>https://scienmag.com/personality-traits-fail-to-explain-why-listeners-hear-speech-differently/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 17:33:44 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[Autism Spectrum Quotient]]></category>
		<category><![CDATA[Bayesian modeling]]></category>
		<category><![CDATA[Big Five personality]]></category>
		<category><![CDATA[compensation for coarticulation]]></category>
		<category><![CDATA[context-dependent speech sound boundaries]]></category>
		<category><![CDATA[effects of neighboring sounds on perception]]></category>
		<category><![CDATA[experimental studies on speech sound recognition]]></category>
		<category><![CDATA[impact of acoustic context on speech interpretation]]></category>
		<category><![CDATA[individual differences]]></category>
		<category><![CDATA[individual differences in speech perception]]></category>
		<category><![CDATA[influence of coarticulation in speech]]></category>
		<category><![CDATA[listener disagreement on ambiguous sounds]]></category>
		<category><![CDATA[perception of fricatives in continuous speech]]></category>
		<category><![CDATA[perceptual categorization]]></category>
		<category><![CDATA[phonetics]]></category>
		<category><![CDATA[psychoacoustics]]></category>
		<category><![CDATA[psycholinguistics]]></category>
		<category><![CDATA[replication]]></category>
		<category><![CDATA[role of vowel context in sound categorization]]></category>
		<category><![CDATA[sound change]]></category>
		<category><![CDATA[speech perception]]></category>
		<category><![CDATA[speech perception variability]]></category>
		<category><![CDATA[speech sound categorization]]></category>
		<category><![CDATA[stability of speech perception traits]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=245237</guid>

					<description><![CDATA[A large replication study finds that personality questionnaires, including the Autism Spectrum Quotient and the Big Five, do not reliably predict how individual listeners compensate for coarticulation in speech perception.]]></description>
										<content:encoded><![CDATA[<p>When two people listen to the same ambiguous sound, they often disagree about what they heard. One listener categorizes a fricative as an &#8220;s&#8221; while another, hearing the identical waveform, hears &#8220;sh.&#8221; For decades, speech scientists have debated whether such disagreements reflect the fleeting circumstances of an experiment or stable, personal characteristics of the listeners themselves. A new study by John Kingston of the University of Massachusetts Amherst, published in Attention, Perception, &amp; Psychophysics, puts that question to a rigorous test and arrives at a sobering answer: perception, at least in this domain, does not appear to be predictably personal.</p>
<p>The phenomenon at the heart of the study is known as compensation for coarticulation. Speech sounds are produced in a continuous stream, and neighboring sounds blur into one another. When a fricative such as &#8220;s&#8221; or &#8220;sh&#8221; follows a vowel, listeners systematically shift the boundary between the two categories depending on the vowel they heard. After a vowel like &#8220;a,&#8221; which pushes the tongue body back and lowers the acoustic energy in the high-frequency range, listeners are more likely to categorize an ambiguous fricative as &#8220;sh&#8221;; after &#8220;i&#8221; or &#8220;u,&#8221; the boundary shifts the other way. This context effect, first documented systematically by Mann and Repp in 1980, is one of the most robust findings in speech perception, yet the size of the effect varies considerably from one listener to the next.</p>
<p>The obvious candidate for explaining that variation is personality. In 2010, researcher Alice Yu reported a striking result: neurotypical listeners who scored higher on the Autism Spectrum Quotient, a questionnaire measuring traits associated with the autism spectrum, compensated less for coarticulation. The finding resonated widely, in part because it connected a low-level perceptual effect to broader theories about how autistic traits shape cognitive style, including attention to detail and reduced sensitivity to context. It also fed into a growing literature on individual differences in phonology, suggesting that the seeds of sound change might lie in the perceptual quirks of particular listeners.</p>
<p>Kingston&#8217;s study set out to replicate that result with substantially more statistical power. The original work had relied on a modest sample, and the replication crisis in psychology has repeatedly shown that small samples can produce seductive correlations that evaporate under scrutiny. In the new AQ experiment, the number of participants was more than doubled, and the analysis was rebuilt from the ground up using Bayesian multilevel modeling, a framework that estimates full probability distributions over parameter values rather than single point estimates. Listeners completed the Autism Spectrum Quotient, including its subtests covering attention switching, attention to detail, communication, imagination, and social skills, before categorizing stimuli drawn from a synthetic continuum between &#8220;s&#8221; and &#8220;sh&#8221; embedded in different vowel contexts.</p>
<p>The outcome was unambiguous. Despite the larger sample and the more sensitive analysis, the specific dependencies that Yu had reported between autistic traits and compensation for coarticulation failed to replicate. Listeners still showed the classic context effects, and their responses still varied from person to person, but the variation did not track their questionnaire scores in the way the original study had suggested. A related study by Lai and colleagues in 2022 had already reported that perceptual compensation and lexical effects were uncorrelated with AQ responses, though a recording problem with their questionnaire had complicated the comparison. The new results strengthen the case that the original correlation was fragile.</p>
<p>But Kingston did not stop at a failed replication. A second experiment, the B5 study, asked a broader question: perhaps the wrong personality instrument had been used. The Big-Five questionnaire, which measures agreeableness, conscientiousness, emotional stability, extraversion, and intellect-imagination, offers a more comprehensive map of personality structure than the AQ, and some researchers have argued that autistic traits are not an independent personality dimension at all but a particular configuration of Big-Five characteristics. If compensation for coarticulation reflects personality in any form, the Big-Five framework should have been able to detect it.</p>
<p>Here the results were more intriguing, and more complicated. Listeners&#8217; categorization responses did differ as a function of their Big-Five trait profiles, and also as a function of their gender. Yet the size and even the direction of compensation varied across genders and across individual traits, with no single, coherent pattern emerging. Rather than revealing a clean mapping between a personality dimension and a perceptual strategy, the data showed a tangle of interactions whose interpretation is far from straightforward. The effects were real in the statistical sense, but they do not add up to a predictive theory of how a given person will hear a given sound.</p>
<p>Kingston draws a deliberately cautious conclusion. The answer to the question posed in the title, he suggests, is no, in one of two senses. Either listeners&#8217; performance is not predictable from their personality traits at all, or it is predictable in principle but the causal pathway remains elusive, with the observed correlations possibly reflecting confounds rather than genuine perceptual consequences of temperament. Distinguishing between these possibilities would require experimental designs that go beyond correlating questionnaire scores with perceptual judgments, perhaps manipulating state rather than trait, or measuring perception repeatedly within individuals across time.</p>
<p>The study also carries a broader methodological message. The original 2010 finding had been cited widely and woven into theoretical accounts of sound change, in which individual perceptual differences were proposed as the raw material on which language evolution acts. If those differences are not reliably tied to measurable personality traits, then theories of sound change may need to find their sources of variation elsewhere, perhaps in the acoustics of different speech communities or in the sheer stochasticity of perceptual processing. The failure also illustrates why large replications matter: a correlation that looks compelling in a small sample can dissolve when the sample doubles and the analysis becomes more rigorous.</p>
<p>In keeping with the open-science practices that have become standard in the field, Kingston has made the data, analysis code, stimuli, and even the script for generating the s-to-sh continuum freely available on the Open Science Framework, allowing any researcher to scrutinize, reuse, or extend the work. Whatever the ultimate verdict on whether perception is personal, the tools for settling the question are now in everyone&#8217;s hands. For now, the sounds we hear may differ from person to person, but the reasons for those differences remain stubbornly, fascinatingly opaque.</p>
<p><strong>Subject of Research:</strong> Individual differences in speech perception and compensation for coarticulation in relation to personality traits</p>
<p><strong>Article Title:</strong> Is perception personal?</p>
<p><strong>Article References:</strong> Kingston, J. (2026). Is perception personal?. <em>Attention, Perception, &amp;amp; Psychophysics, 88</em>(7), Article 200. <a href="https://doi.org/10.3758/s13414-026-03340-6" rel="noopener noreferrer">https://doi.org/10.3758/s13414-026-03340-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13414-026-03340-6" rel="noopener noreferrer">10.3758/s13414-026-03340-6</a></p>
<p><strong>Keywords:</strong> speech perception, compensation for coarticulation, Autism Spectrum Quotient, Big Five personality, individual differences, replication, Bayesian modeling, psychoacoustics, phonetics, perceptual categorization, sound change, psycholinguistics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">245237</post-id>	</item>
		<item>
		<title>New Bilingual Speech Dataset Takes Aim at AI&#8217;s Weakest Spot: Code-Switching</title>
		<link>https://scienmag.com/new-bilingual-speech-dataset-takes-aim-at-ais-weakest-spot-code-switching/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:35:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[automatic speech recognition]]></category>
		<category><![CDATA[automatic speech recognition challenges]]></category>
		<category><![CDATA[bilingual speech]]></category>
		<category><![CDATA[bilingual speech recognition]]></category>
		<category><![CDATA[code-switching]]></category>
		<category><![CDATA[code-switching dataset]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[DOTA-ME-CS corpus]]></category>
		<category><![CDATA[improving machine understanding of code-switching]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Mandarin]]></category>
		<category><![CDATA[Mandarin-English code-switching]]></category>
		<category><![CDATA[multilingual natural language processing]]></category>
		<category><![CDATA[multilingual speech processing]]></category>
		<category><![CDATA[open-source language datasets]]></category>
		<category><![CDATA[Paraformer]]></category>
		<category><![CDATA[phonetics]]></category>
		<category><![CDATA[SenseVoice]]></category>
		<category><![CDATA[speech dataset]]></category>
		<category><![CDATA[speech dataset for bilingual speakers]]></category>
		<category><![CDATA[speech recognition for code-switching]]></category>
		<category><![CDATA[transformer-based ASR models]]></category>
		<category><![CDATA[voice conversion]]></category>
		<category><![CDATA[Whisper]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203051</guid>

					<description><![CDATA[A new open dataset of 9300 Mandarin-English code-switched speech recordings, enhanced with AI-generated noise, speed and timbre changes, exposes how badly today's speech recognition models fail at bilingual conversation.]]></description>
										<content:encoded><![CDATA[<p>When bilingual speakers chat with one another, they rarely stay inside a single language. A sentence that begins in Mandarin may slip mid-phrase into English and back again, a behaviour linguists call code-switching. It is one of the most natural things multilingual people do, and one of the most unnatural things for machines to understand. Automatic speech recognition (ASR) systems, even the most powerful transformer-based models now in wide use, tend to stumble exactly at the point where one language hands off to another. A new openly available corpus called DOTA-ME-CS, short for Daily Oriented Text Audio Mandarin-English Code-Switching dataset, has been created to give researchers the fuel they need to close that gap.</p>
<p>The dataset, described in the Journal of Ambient Intelligence and Humanized Computing, contains 18.54 hours of audio spanning 9300 recordings produced by 34 bilingual participants, all of them fluent in both Mandarin and English. Unlike many earlier corpora, every single utterance in the collection involves code-switching. That design choice matters. Older resources such as SEAME, which stretches across roughly 190 hours, and TALCS, which covers 587 hours, contain substantial proportions of monolingual speech, meaning their effective supply of genuinely code-switched material is far smaller than their total length suggests. Some of those datasets, including the ASRU and TALCS corpora, are no longer publicly accessible at all, leaving the field with a shortage of usable, openly available benchmarks.</p>
<p>The construction of DOTA-ME-CS follows an unusual pipeline that blends large language model generation with human recording. The team used GPT-4o with carefully engineered prompts to produce scripted sentences across ten everyday scenarios: education, entertainment, environmental protection, food, health, home, life, pets, travel and work. Each prompt required the model to produce sentences that mimic daily conversational style, contain more English words than Mandarin words, and include at least one Mandarin word, with a dominant language assigned to every sentence. The authors justify this topic-anchored approach with a probabilistic argument: when a specific category is given, the probability of generating a relevant, high-quality sentence is higher than when the model is left to roam across all possible topics, which also reduces hidden cultural bias in the resulting scripts.</p>
<p>Human evaluators checked the generated scripts for grammatical problems and confirmed the presence of genuine switching, and a post-hoc naturalness study asked eight bilingual raters to score 100 randomly sampled stimuli on a five-point Likert scale. The average rating came in at 4.12 with a standard deviation of 0.61, and the median was 4.0, indicating that bilingual listeners generally perceive the generated sentences as natural, though the authors acknowledge that occasional awkwardness is an inherent limitation of LLM-based text generation. Bilingual volunteers, mostly college students with academic backgrounds in China and the United Kingdom, including roughly eighteen participants based at Imperial College London, then recorded the scripts on their own laptops or smartphones as 16-bit WAV files in quiet indoor settings. Recordings that failed basic quality checks were rejected and re-recorded, and accepted files were peak-normalised to 3 dBFS for consistent loudness.</p>
<p>The participants&#8217; linguistic profiles were deliberately diverse. Twelve reported English dominance, eighteen reported Mandarin dominance and four identified as balanced bilinguals, and the mix of speakers from China and the United Kingdom ensures that the corpus spans multiple varieties of English and second-language accents. An average of 2.22 switching points per utterance, combined with a broad part-of-speech distribution across both languages, means the recordings capture switching at varied syntactic positions and grammatical categories rather than concentrating it in a single predictable spot. The dataset also includes longer recordings of roughly 10 to 15 seconds and about 100 words, an intentional response to the weakness current ASR models show on extended speech, where dependencies and critical information can be lost.</p>
<p>What sets the corpus apart most sharply is its AI-driven augmentation. Because human recordings were captured in quiet conditions at a normal pace, the researchers used the Librosa audio library to modify a random subset of clips. Playback speed was adjusted to 0.75x, 0.5x, 1.25x, 1.5x and 2x, with 200 recordings modified at each setting. Five categories of background noise, drawn from highways, war, natural sounds, white noise and playground environments, were added at 200 recordings per type, with the noise pitch scaled to 0.8x to account for the Lombard effect, the well-documented tendency of speakers to raise their vocal intensity in noisy surroundings. In addition, five AI-generated timbres, two male and three female, replaced the original voices in 200 recordings each, simulating both voice conversion deepfake scenarios and privacy-preserving speech processing conditions.</p>
<p>The accompanying data analysis is unusually thorough. Phoneme distributions catalogued with the International Phonetic Alphabet reveal that Mandarin contributes more multisyllabic pronunciations, tones, aspirated consonants and palatalised sounds, while English favours single-syllable vowels; shared features such as the sounds t and a suggest cues that future switching-detection models could exploit. Physical measurements show a frame rate of 46.7 kHz and a sample width of 2.0 bytes, meeting professional audio standards, with an average maximum pitch of 3617.13 Hz and an average per-recording pitch of 535.04 Hz. Formant frequencies computed through Linear Predictive Coding yield averages of 804.02 Hz, 4419.15 Hz and 7549.48 Hz for F1, F2 and F3, pointing to mid-to-low and front vowels, reduced lip rounding and retroflex consonants. The measured speaking rate of 2.05 words per second, with a standard deviation of 0.49, sits squarely within the typical range for conversational read speech, supporting the corpus&#8217;s ecological validity.</p>
<p>To establish baseline performance, the team benchmarked three pre-trained ASR models on the full dataset without fine-tuning: Whisper large-v3, a 1.5-billion-parameter Transformer encoder-decoder trained on 680,000 hours of multilingual audio; SenseVoice Small, a lightweight hybrid CTC-attention model of roughly 200 million parameters that also performs language identification, emotion detection and audio event detection; and Paraformer, a non-autoregressive Transformer with 220 million parameters optimised for Mandarin and built on continuous integrate-and-fire alignment. Paraformer and SenseVoice achieved the best results, with Whisper slightly behind, and a word error rate of 0.177 alongside a character error rate of 0.176. Compared against the publicly available ASCEND corpus, the models produced higher error rates on DOTA-ME-CS, which the authors read not as a defect but as evidence that their dataset poses a more demanding and therefore more useful benchmark.</p>
<p>Case studies expose exactly where today&#8217;s systems fail. In one example embedding the Chinese noun 菠萝包, meaning pineapple bun, inside an English sentence, Paraformer and SenseVoice produced garbled outputs such as bullleball and boobao, while Whisper translated the phrase correctly but destroyed the mixed-language structure by refusing to preserve the code-switch itself, a translation bias that also appeared on noise-free recordings. At double playback speed, Paraformer&#8217;s output strayed entirely from the intended meaning, SenseVoice descended into repetitions such as Miamiami, and Whisper mangled Florida lifestyle into Forensic Lifestyle. Altered timbres degraded English recognition across the board, and all three models performed worst at 2.0x speed. Whisper proved the most resilient to background noise, while Paraformer was the most sensitive to voice changes.</p>
<p>The implications reach beyond engineering. Reliable code-switching recognition would make automated transcription, voice assistants and translation systems far more useful for the hundreds of millions of people who live their linguistic lives between two languages, reducing a bias that currently disadvantages non-monolingual speakers. The authors stress that the study received ethical approval from Imperial College London&#8217;s ethical board, that all participants gave informed consent, and that privacy protections were maintained throughout. They are candid about limitations: funding constrained the scale of the corpus, and scripted reading, while reducing transcription errors and ethical risks, sacrifices the spontaneity of natural conversation. Even so, their deliberate trade-off of quality and coverage over raw scale gives the community something it has lacked, a fully public, exclusively code-switching Mandarin-English benchmark with baseline scores, rigorous acoustic analysis and data, code and recordings all released through a public repository. As speech technology races toward multilingual ubiquity, datasets like this one may determine whether the next generation of voice interfaces can finally follow the way people actually talk.</p>
<p><strong>Subject of Research:</strong> A Mandarin-English code-switching speech dataset for advancing automatic speech recognition research.</p>
<p><strong>Article Title:</strong> DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset</p>
<p><strong>Article References:</strong> Li, Y., Wei, Z., Yu, H., Xue, J., Zhou, H., &amp; Schuller, B. W. (2026). DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset. <em>Journal of Ambient Intelligence and Humanized Computing</em>. <a href="https://doi.org/10.1007/s12652-026-05119-x" rel="noopener noreferrer">https://doi.org/10.1007/s12652-026-05119-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12652-026-05119-x" rel="noopener noreferrer">10.1007/s12652-026-05119-x</a></p>
<p><strong>Keywords:</strong> code-switching, automatic speech recognition, Mandarin, bilingual speech, speech dataset, Whisper, Paraformer, SenseVoice, data augmentation, phonetics, large language models, voice conversion</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203051</post-id>	</item>
	</channel>
</rss>
