<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>prosody &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/prosody/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 17:41:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>prosody &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Voice Clones Can Now Copy the Human Sound of Confidence and Doubt</title>
		<link>https://scienmag.com/ai-voice-clones-can-now-copy-the-human-sound-of-confidence-and-doubt/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 17:41:38 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[acoustic analysis]]></category>
		<category><![CDATA[acoustic features of confidence and doubt]]></category>
		<category><![CDATA[AI voice cloning]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Current Psychology]]></category>
		<category><![CDATA[eGeMAPS]]></category>
		<category><![CDATA[experimental design in voice perception studies]]></category>
		<category><![CDATA[fundamental frequency]]></category>
		<category><![CDATA[human expression of confidence and doubt]]></category>
		<category><![CDATA[Human-AI Interaction]]></category>
		<category><![CDATA[human-AI voice interaction research]]></category>
		<category><![CDATA[implications for conversational AI development]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[prosody]]></category>
		<category><![CDATA[psychological experiments with voice stimuli]]></category>
		<category><![CDATA[spectral flux]]></category>
		<category><![CDATA[speech emotion recognition]]></category>
		<category><![CDATA[speech prosody analysis]]></category>
		<category><![CDATA[speech synthesis]]></category>
		<category><![CDATA[synthetic voice emotional mimicry]]></category>
		<category><![CDATA[vocal confidence]]></category>
		<category><![CDATA[voice cloning]]></category>
		<category><![CDATA[voice cloning fidelity]]></category>
		<category><![CDATA[voice identity versus emotional expression]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197043</guid>

					<description><![CDATA[A new study shows that AI voice cloning systems preserve not only speaker identity but also the human-specific prosodic cues of confidence and doubt.]]></description>
										<content:encoded><![CDATA[<p>Voice-cloning artificial intelligence has long been judged by how faithfully it reproduces the identity of a speaker: the recognizable timbre, pitch range, and resonance that make a voice sound like one particular person rather than another. A new study published in Current Psychology by Wenjun Chen and Xiaoming Jiang of Shanghai International Studies University, with Chen also affiliated with McGill University, asks a subtler and arguably more consequential question. Can AI voice clones reproduce not just who a speaker is, but how that speaker sounds when expressing something human-specific, such as confidence or doubt? The answer, based on a rigorous acoustic analysis of thousands of utterances, is largely yes, and the finding has immediate implications for how psychologists build experiments on human-AI voice interaction.</p>
<p>The motivation for the study stems from a methodological bottleneck. As voice assistants, synthetic narrators, and conversational agents become embedded in daily life, researchers increasingly need experimental stimuli in which speaker identity and prosodic style can be manipulated independently. If a scientist wants to test how listeners react to a confident versus a doubtful statement, the comparison is confounded when different human speakers deliver the confident and doubtful versions, because listeners may respond to the person rather than the prosody. Voice cloning promises a solution: train a model on one speaker&#8217;s voice, then generate new sentences carrying a target emotional or attitudinal tone. But this only works if the cloning system genuinely transfers prosody rather than flattening it into a generic synthetic delivery.</p>
<p>To test this, the researchers recruited ten native Mandarin speakers who each produced thirty sentences, fifteen drawn from geography statements and fifteen from trivia statements, in three distinct intonations: confident, doubtful, and neutral. This design yielded a rich corpus of human recordings spanning different sentence contents and prosodic intentions. For each speaker and each prosody, the team trained separate AI voice clones using either the geography recordings or the trivia recordings as training material. Each clone was then asked to generate all thirty sentences, producing a total of 2,700 utterances for analysis: 900 human recordings, 900 AI-generated utterances from geography-trained clones, and 900 from trivia-trained clones. Crucially, the training and generation sets were crossed, so a clone trained only on confident geography sentences still had to produce doubtful trivia sentences, forcing the system to generalize prosodic style to entirely novel content.</p>
<p>The analytical strategy combined machine learning classification with traditional statistical modeling. The researchers extracted 88 acoustic features using the extended Geneva Minimalistic Acoustic Parameter Set, or eGeMAPS, a standardized toolkit widely used in voice research and affective computing. They then trained gradient-boosted tree classifiers, a machine learning approach implemented through the XGBoost framework, to distinguish confident from doubtful prosody within each source of speech. Within a single source, whether human or AI, the classifiers achieved high accuracy, averaging 0.85, indicating that the acoustic signatures of confidence and doubt are robust and detectable in both natural and synthetic voices. More importantly, when classifiers trained on human speech were tested on AI speech, and vice versa, accuracy remained substantially above chance, averaging 0.65. This cross-source generalization demonstrates that the prosodic cues of confidence and doubt are encoded in a shared acoustic currency that transcends the boundary between human and machine voices.</p>
<p>Among the 88 features, one emerged as a consistently dominant contributor: spectral flux, a measure of how rapidly the frequency content of the speech signal changes over time. Spectral flux captures the crispness and dynamism of articulation, and confident speech, in both humans and their AI clones, exhibited higher spectral flux than doubtful speech. This makes intuitive sense. A confident speaker articulates with decisive energy, producing sharper spectral transitions, whereas a hesitant speaker tends toward softer, less defined acoustic edges. Notably, spectral flux showed no effect of speaker sex, which the authors interpret as evidence that it indexes prosodic style rather than stable anatomical differences between speakers. In other words, it is a marker of how something is said, not of who is saying it, which is precisely the kind of feature a stimulus designer would want to manipulate independently of speaker identity.</p>
<p>Linear mixed-effects models, a statistical framework that accounts for the nested structure of the data with multiple speakers, sentences, and prosodies, confirmed a coherent acoustic profile of vocal confidence. Across human and AI speakers alike, confident prosody was associated with higher spectral flux, a longer estimated vocal tract length, and lower fundamental frequency, the acoustic correlate of pitch, compared with doubtful prosody. The estimated vocal tract length, derived from formant frequencies, reflects how speakers shape their resonating cavities; a longer apparent vocal tract projects a larger, more authoritative body, echoing a well-documented literature on vocal size exaggeration in humans. That AI clones reproduced this constellation of cues, including the anatomical-sounding shift in apparent vocal tract size, suggests that modern cloning systems capture not only surface acoustics but the embodied gestalt of an attitudinal vocal stance.</p>
<p>The study also probed how prosodic information unfolds in time. Using time-resolved decoding of fundamental frequency, the researchers found that the distinction between confident and doubtful prosody was most reliably recovered in the late windows of each utterance, across human and AI sources alike. This late-emerging pattern aligns with the intuition that speakers often commit to an attitude as a sentence progresses, with final intonational contours sealing the impression of certainty or hesitation. However, a subtle asymmetry appeared: human speakers distributed their prosodic markers more broadly across the utterance, weaving cues of confidence and doubt throughout the signal, whereas the AI clones concentrated their prosodic differentiation more narrowly toward the end. This temporal signature may represent one of the remaining acoustic fingerprints distinguishing synthetic from natural expressive speech.</p>
<p>Indeed, the authors are careful to note that despite the impressive prosodic fidelity, AI-generated speech still forms a distinguishable acoustic population relative to human speech. The cross-source classification accuracy of 0.65, while well above chance, falls short of the 0.85 achieved within sources, indicating that machine and human voices are similar but not acoustically identical. This finding resonates with recent work showing that voice clones sound realistic but not yet hyperrealistic, and that human listeners and neural systems can, under some conditions, separate deepfake from genuine speaker identity. For the immediate purposes of psychological research, however, the residual human-machine gap may even be an advantage, since it allows researchers to verify that their synthetic stimuli behave acoustically as intended while retaining a measurable distinction from natural recordings.</p>
<p>The broader significance of the study lies in its methodological contribution. By demonstrating that voice cloning preserves prosodic style alongside speaker identity, Chen and Jiang provide empirical license for a new generation of controlled experiments on human-AI voice interaction. Researchers can now, with appropriate validation, generate stimulus sets in which the same voice delivers the same content with systematically varied confidence, doubt, or neutrality, isolating the causal impact of prosody on listener trust, memory, persuasion, and social perception. As AI voices become conversation partners, teachers, and companions, understanding how their vocal expressions of certainty are produced, perceived, and potentially distinguished from human ones becomes a scientific priority. This study shows that the tools for doing that science rigorously are already within reach, and that the line between human and machine vocal expression, while still detectable, is growing finer by the year.</p>
<p><strong>Subject of Research:</strong> Whether AI voice cloning systems can reproduce human prosodic expressions of confidence and doubt</p>
<p><strong>Article Title:</strong> Voice-cloning artificial-intelligence speakers can also mimic human-specific vocal expression</p>
<p><strong>Article References:</strong> Chen, W., &amp; Jiang, X. (2026). Voice-cloning artificial-intelligence speakers can also mimic human-specific vocal expression. <em>Current Psychology, 45</em>(17), Article 1490. <a href="https://doi.org/10.1007/s12144-026-09991-w" rel="noopener noreferrer">https://doi.org/10.1007/s12144-026-09991-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12144-026-09991-w" rel="noopener noreferrer">10.1007/s12144-026-09991-w</a></p>
<p><strong>Keywords:</strong> voice cloning, artificial intelligence, prosody, speech synthesis, vocal confidence, spectral flux, machine learning, human-AI interaction, acoustic analysis, eGeMAPS, fundamental frequency, Current Psychology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197043</post-id>	</item>
	</channel>
</rss>
