<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>speech synthesis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/speech-synthesis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 20:25:31 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>speech synthesis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Cloud-Edge AI System Translates Speech While Protecting Speaker Identity</title>
		<link>https://scienmag.com/cloud-edge-ai-system-translates-speech-while-protecting-speaker-identity/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 20:25:31 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[adaptive computational architecture for multilingual speech translation]]></category>
		<category><![CDATA[AI Flow]]></category>
		<category><![CDATA[AI-driven voice anonymization techniques]]></category>
		<category><![CDATA[Cloud-Edge AI speech translation]]></category>
		<category><![CDATA[cloud-edge collaboration]]></category>
		<category><![CDATA[cloud-edge collaboration in speech translation systems]]></category>
		<category><![CDATA[collaborative cloud-edge AI for speech-to-speech translation]]></category>
		<category><![CDATA[early exit]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[heterogeneous device support for speech translation]]></category>
		<category><![CDATA[machine translation]]></category>
		<category><![CDATA[multilingual speech translation without raw voice transfer]]></category>
		<category><![CDATA[neural network model scalability for edge devices]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[privacy]]></category>
		<category><![CDATA[privacy and efficiency in real-time speech translation]]></category>
		<category><![CDATA[privacy-aware voice data processing in AI translation]]></category>
		<category><![CDATA[resource-constrained inference]]></category>
		<category><![CDATA[secure voice data transmission in AI translation pipelines]]></category>
		<category><![CDATA[speaker identity]]></category>
		<category><![CDATA[speaker privacy preservation in neural translation systems]]></category>
		<category><![CDATA[speech synthesis]]></category>
		<category><![CDATA[speech translation]]></category>
		<category><![CDATA[voice preservation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198320</guid>

					<description><![CDATA[A new cloud-edge collaborative framework adapts speech-to-speech translation to any device's compute budget while preserving speaker identity through a privacy-preserving retrieval system.]]></description>
										<content:encoded><![CDATA[<p>Speech-to-speech translation has long promised a world in which language barriers simply dissolve: a traveler speaks Japanese into a phone and their companion hears fluent English, a business negotiator conducts a multilingual conference call without an interpreter, and a filmmaker dubs content across dozens of languages. Modern neural systems have grown remarkably good at translating the words themselves. Yet two stubborn problems have limited real-world deployment. Most research teams release only a single model size, forcing every device—from a smartwatch to a data center—to run the same network regardless of its computing budget, and the standard trick for preserving a speaker&#8217;s distinctive voice requires feeding sensitive acoustic data directly into the translation pipeline, raising uncomfortable privacy questions. A new study published in the journal Vicinagearth tackles both problems at once with a cloud-edge collaborative architecture that adapts its own computational effort to the hardware it runs on while never transmitting the speaker&#8217;s raw voice across the network.</p>
<p>The research, led by Boyu Zhu and Xiao-Lei Zhang of Northwestern Polytechnical University together with Rujin Chen and Chi Zhang of the Institute of Artificial Intelligence at China Telecom, builds on the AI Flow framework, which reconceives the network&#8217;s job as transmitting intelligence flows rather than raw information flows. In this design, edge devices, edge servers, and cloud servers share an inference pipeline, each contributing the compute it has available. Applied to speech translation, the system splits into three cooperating modules: a speech-to-text translation module distributed between the sender&#8217;s device and the cloud, a voice preservation module running entirely on the sender&#8217;s device, and a text-to-speech synthesis module on the receiver&#8217;s device. The result is a pipeline that translates source speech into target-language text, retrieves a compact identity reference, and regenerates natural target speech that carries the original speaker&#8217;s acoustic character.</p>
<p>The central technical innovation is a family of early-exit heads attached to the translation backbone, a strategy borrowed from efficient inference research on models such as HuBERT and Whisper. Rather than forcing every input through all decoder layers, an early-exit system evaluates its own uncertainty after each layer and stops when confidence is high enough. In the traditional version of this strategy, the team computes the Shannon entropy of the predicted token distribution after each decoder layer; if entropy falls below a fixed threshold, computation terminates immediately, saving the cost of the remaining layers. Static evaluations at layers 16, 18, 20, and 24 showed that usable translations emerge well before the full model completes its pass, meaning the same trained network can serve phones, laptops, and servers with very different resources simply by deciding where to stop.</p>
<p>But a simple entropy threshold is blunt: easy sentences exit too late and hard ones exit too early, degrading translation quality. The team&#8217;s solution is a small language model that learns to predict, for each input, which exit layer will produce the best result. Training this predictor required difficulty labels, which the authors generated with a teacher-guided classifier built on the GPT-4o API. The large language model scored each training sample on a one-to-ten difficulty scale, focusing on the rarity and frequency of long-tail vocabulary, structural divergence between source and target languages, and ambiguity or context dependence. Because difficulty judgments can shift systematically across language pairs, the researchers calibrated the scores statistically—standardizing within each pair and applying quantile mapping to a common reference distribution—so that a score of seven in French-to-English means the same thing as a seven in Polish-to-German.</p>
<p>Once the difficulty labels existed, the team trained a Flan-T5-based small model that takes the output of an automatic speech recognition pass and predicts the optimal early-exit layer for the translation model. During collaborative inference, the large translation model sits in the cloud while a pruned small model runs on the sender&#8217;s edge device, both fed by a lightweight convolutional preprocessing module that converts raw waveforms into compact feature tensors instead of streaming audio. The small model applies an entropy test: if it exits confidently, it raises a flag and the receiver synthesizes speech immediately from the local translation. If not, the sender waits up to a bounded number of seconds for the cloud result, falling back to the local output if the network stalls—guaranteeing bounded latency under real network conditions.</p>
<p>The second major contribution addresses voice preservation without privacy leakage. Conventional expressive S2ST systems pass speaker-related acoustic embeddings into the translation model itself, exposing biometric voice information to the cloud. The new system instead performs retrieval-based voice preservation. On the sender&#8217;s device, a lightweight pipeline extracts three kinds of features from the input speech: a speaker embedding generated by a RawNet3-based verification model, an emotion category from the emotion2vec+ representation, and a speaking rate computed as syllables per unit of voiced duration using energy-based peak detection. These features drive a hierarchical search through a large multilingual acoustic database built from the M3PDB dataset, which contains multiple speakers per language with samples spanning varied emotions and speaking rates.</p>
<p>The hierarchical matching narrows candidates in stages: first the speaker embedding locates the most acoustically similar speaker, then emotion category must match exactly, and finally the speech-rate similarity selects the reference sample whose cadence most closely mirrors the input. Only the identifier of that winning reference—never the voice itself—is transmitted to the receiver. The receiving device looks up the corresponding speech tokens, extracted with the tokenizer from IndexTTS2, and conditions its synthesis on both the translated text and this token, reconstructing target speech that echoes the source speaker&#8217;s timbre, emotion, and pacing without any raw acoustic data ever leaving the sender.</p>
<p>Experimental validation spanned two fronts. On the CoVoST2 French translation benchmark, evaluated with BLEU scores on Whisper-Large-V2, the SLM-assisted early-exit strategy outperformed entropy-, confidence-, and cosine-similarity-based early-exit baselines, indicating that predicting the optimal exit depth preserves translation quality better than uncertainty thresholds alone. For voice preservation, the team constructed demanding test sets from LibriTTS samples corrupted with noise at signal-to-noise ratios of negative five, five, and fifteen decibels, drawing noise from AudioSet, FreeSound, WHAM!, and FSD50K sources, with reverberation added probabilistically using simulated room impulse responses. Across these conditions, the retrieval-based method achieved lower word error rates and higher UTMOS naturalness scores than a baseline that transmitted the source speech directly, with the advantage widening as noise increased—a striking demonstration that a compact retrieved reference can be more robust than the original recording.</p>
<p>Cross-lingual generalization tests on the VoxPopuli corpus covered sixteen European languages, including English, German, French, Spanish, Polish, Italian, Romanian, Hungarian, Czech, Dutch, Finnish, Croatian, Slovak, Slovene, Estonian, and Lithuanian. Word error rates dropped for fourteen of the sixteen languages relative to the direct-transmission baseline, while speaking-rate consistency held nearly constant at the looser tolerance threshold and emotion consistency showed no substantial overall difference. The authors note that although objective speaker-similarity scores dipped below the baseline in noisy conditions, human listeners are typically far less sensitive to such differences than automated metrics suggest, so the perceptual gap is likely smaller than the numbers imply.</p>
<p>The broader significance of this work lies in what it reframes rather than what it merely accelerates. By treating model size as a deployment decision rather than a design constraint, the early-exit architecture lets one trained network flexibly serve the full spectrum of hardware, from battery-constrained earbuds to cloud GPUs. By relocating speaker identity from the model&#8217;s input to a retrieval lookup, it decouples expressive fidelity from biometric exposure, addressing a concern that will only grow as voice interfaces proliferate. And by showing that a few transmitted identifiers can outperform full audio transmission under noise, the study makes a compelling case that the future of real-time speech translation may rest less on bigger models than on smarter collaboration between the devices we carry and the servers that await our hardest questions. The framework, the authors conclude, offers a practical path toward deploying speech translation where it matters most: out in the noisy, bandwidth-limited, privacy-sensitive real world.</p>
<p><strong>Subject of Research:</strong> A cloud-edge collaborative speech-to-speech translation system with early-exit inference and retrieval-based voice preservation.</p>
<p><strong>Article Title:</strong> Speech to speech translation system based on cloud-edge collaboration</p>
<p><strong>Article References:</strong> Zhu, B., Chen, R., Zhang, C., &amp; Zhang, X.-L. (2026). Speech to speech translation system based on cloud-edge collaboration. <em>Vicinagearth, 3</em>(1), Article 4. <a href="https://doi.org/10.1007/s44336-026-00033-4" rel="noopener noreferrer">https://doi.org/10.1007/s44336-026-00033-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-026-00033-4" rel="noopener noreferrer">10.1007/s44336-026-00033-4</a></p>
<p><strong>Keywords:</strong> speech translation, cloud-edge collaboration, early exit, voice preservation, privacy, speech synthesis, neural networks, edge computing, machine translation, speaker identity, AI Flow, resource-constrained inference</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198320</post-id>	</item>
		<item>
		<title>AI Voice Clones Can Now Copy the Human Sound of Confidence and Doubt</title>
		<link>https://scienmag.com/ai-voice-clones-can-now-copy-the-human-sound-of-confidence-and-doubt/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 17:41:38 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[acoustic analysis]]></category>
		<category><![CDATA[acoustic features of confidence and doubt]]></category>
		<category><![CDATA[AI voice cloning]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Current Psychology]]></category>
		<category><![CDATA[eGeMAPS]]></category>
		<category><![CDATA[experimental design in voice perception studies]]></category>
		<category><![CDATA[fundamental frequency]]></category>
		<category><![CDATA[human expression of confidence and doubt]]></category>
		<category><![CDATA[Human-AI Interaction]]></category>
		<category><![CDATA[human-AI voice interaction research]]></category>
		<category><![CDATA[implications for conversational AI development]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[prosody]]></category>
		<category><![CDATA[psychological experiments with voice stimuli]]></category>
		<category><![CDATA[spectral flux]]></category>
		<category><![CDATA[speech emotion recognition]]></category>
		<category><![CDATA[speech prosody analysis]]></category>
		<category><![CDATA[speech synthesis]]></category>
		<category><![CDATA[synthetic voice emotional mimicry]]></category>
		<category><![CDATA[vocal confidence]]></category>
		<category><![CDATA[voice cloning]]></category>
		<category><![CDATA[voice cloning fidelity]]></category>
		<category><![CDATA[voice identity versus emotional expression]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197043</guid>

					<description><![CDATA[A new study shows that AI voice cloning systems preserve not only speaker identity but also the human-specific prosodic cues of confidence and doubt.]]></description>
										<content:encoded><![CDATA[<p>Voice-cloning artificial intelligence has long been judged by how faithfully it reproduces the identity of a speaker: the recognizable timbre, pitch range, and resonance that make a voice sound like one particular person rather than another. A new study published in Current Psychology by Wenjun Chen and Xiaoming Jiang of Shanghai International Studies University, with Chen also affiliated with McGill University, asks a subtler and arguably more consequential question. Can AI voice clones reproduce not just who a speaker is, but how that speaker sounds when expressing something human-specific, such as confidence or doubt? The answer, based on a rigorous acoustic analysis of thousands of utterances, is largely yes, and the finding has immediate implications for how psychologists build experiments on human-AI voice interaction.</p>
<p>The motivation for the study stems from a methodological bottleneck. As voice assistants, synthetic narrators, and conversational agents become embedded in daily life, researchers increasingly need experimental stimuli in which speaker identity and prosodic style can be manipulated independently. If a scientist wants to test how listeners react to a confident versus a doubtful statement, the comparison is confounded when different human speakers deliver the confident and doubtful versions, because listeners may respond to the person rather than the prosody. Voice cloning promises a solution: train a model on one speaker&#8217;s voice, then generate new sentences carrying a target emotional or attitudinal tone. But this only works if the cloning system genuinely transfers prosody rather than flattening it into a generic synthetic delivery.</p>
<p>To test this, the researchers recruited ten native Mandarin speakers who each produced thirty sentences, fifteen drawn from geography statements and fifteen from trivia statements, in three distinct intonations: confident, doubtful, and neutral. This design yielded a rich corpus of human recordings spanning different sentence contents and prosodic intentions. For each speaker and each prosody, the team trained separate AI voice clones using either the geography recordings or the trivia recordings as training material. Each clone was then asked to generate all thirty sentences, producing a total of 2,700 utterances for analysis: 900 human recordings, 900 AI-generated utterances from geography-trained clones, and 900 from trivia-trained clones. Crucially, the training and generation sets were crossed, so a clone trained only on confident geography sentences still had to produce doubtful trivia sentences, forcing the system to generalize prosodic style to entirely novel content.</p>
<p>The analytical strategy combined machine learning classification with traditional statistical modeling. The researchers extracted 88 acoustic features using the extended Geneva Minimalistic Acoustic Parameter Set, or eGeMAPS, a standardized toolkit widely used in voice research and affective computing. They then trained gradient-boosted tree classifiers, a machine learning approach implemented through the XGBoost framework, to distinguish confident from doubtful prosody within each source of speech. Within a single source, whether human or AI, the classifiers achieved high accuracy, averaging 0.85, indicating that the acoustic signatures of confidence and doubt are robust and detectable in both natural and synthetic voices. More importantly, when classifiers trained on human speech were tested on AI speech, and vice versa, accuracy remained substantially above chance, averaging 0.65. This cross-source generalization demonstrates that the prosodic cues of confidence and doubt are encoded in a shared acoustic currency that transcends the boundary between human and machine voices.</p>
<p>Among the 88 features, one emerged as a consistently dominant contributor: spectral flux, a measure of how rapidly the frequency content of the speech signal changes over time. Spectral flux captures the crispness and dynamism of articulation, and confident speech, in both humans and their AI clones, exhibited higher spectral flux than doubtful speech. This makes intuitive sense. A confident speaker articulates with decisive energy, producing sharper spectral transitions, whereas a hesitant speaker tends toward softer, less defined acoustic edges. Notably, spectral flux showed no effect of speaker sex, which the authors interpret as evidence that it indexes prosodic style rather than stable anatomical differences between speakers. In other words, it is a marker of how something is said, not of who is saying it, which is precisely the kind of feature a stimulus designer would want to manipulate independently of speaker identity.</p>
<p>Linear mixed-effects models, a statistical framework that accounts for the nested structure of the data with multiple speakers, sentences, and prosodies, confirmed a coherent acoustic profile of vocal confidence. Across human and AI speakers alike, confident prosody was associated with higher spectral flux, a longer estimated vocal tract length, and lower fundamental frequency, the acoustic correlate of pitch, compared with doubtful prosody. The estimated vocal tract length, derived from formant frequencies, reflects how speakers shape their resonating cavities; a longer apparent vocal tract projects a larger, more authoritative body, echoing a well-documented literature on vocal size exaggeration in humans. That AI clones reproduced this constellation of cues, including the anatomical-sounding shift in apparent vocal tract size, suggests that modern cloning systems capture not only surface acoustics but the embodied gestalt of an attitudinal vocal stance.</p>
<p>The study also probed how prosodic information unfolds in time. Using time-resolved decoding of fundamental frequency, the researchers found that the distinction between confident and doubtful prosody was most reliably recovered in the late windows of each utterance, across human and AI sources alike. This late-emerging pattern aligns with the intuition that speakers often commit to an attitude as a sentence progresses, with final intonational contours sealing the impression of certainty or hesitation. However, a subtle asymmetry appeared: human speakers distributed their prosodic markers more broadly across the utterance, weaving cues of confidence and doubt throughout the signal, whereas the AI clones concentrated their prosodic differentiation more narrowly toward the end. This temporal signature may represent one of the remaining acoustic fingerprints distinguishing synthetic from natural expressive speech.</p>
<p>Indeed, the authors are careful to note that despite the impressive prosodic fidelity, AI-generated speech still forms a distinguishable acoustic population relative to human speech. The cross-source classification accuracy of 0.65, while well above chance, falls short of the 0.85 achieved within sources, indicating that machine and human voices are similar but not acoustically identical. This finding resonates with recent work showing that voice clones sound realistic but not yet hyperrealistic, and that human listeners and neural systems can, under some conditions, separate deepfake from genuine speaker identity. For the immediate purposes of psychological research, however, the residual human-machine gap may even be an advantage, since it allows researchers to verify that their synthetic stimuli behave acoustically as intended while retaining a measurable distinction from natural recordings.</p>
<p>The broader significance of the study lies in its methodological contribution. By demonstrating that voice cloning preserves prosodic style alongside speaker identity, Chen and Jiang provide empirical license for a new generation of controlled experiments on human-AI voice interaction. Researchers can now, with appropriate validation, generate stimulus sets in which the same voice delivers the same content with systematically varied confidence, doubt, or neutrality, isolating the causal impact of prosody on listener trust, memory, persuasion, and social perception. As AI voices become conversation partners, teachers, and companions, understanding how their vocal expressions of certainty are produced, perceived, and potentially distinguished from human ones becomes a scientific priority. This study shows that the tools for doing that science rigorously are already within reach, and that the line between human and machine vocal expression, while still detectable, is growing finer by the year.</p>
<p><strong>Subject of Research:</strong> Whether AI voice cloning systems can reproduce human prosodic expressions of confidence and doubt</p>
<p><strong>Article Title:</strong> Voice-cloning artificial-intelligence speakers can also mimic human-specific vocal expression</p>
<p><strong>Article References:</strong> Chen, W., &amp; Jiang, X. (2026). Voice-cloning artificial-intelligence speakers can also mimic human-specific vocal expression. <em>Current Psychology, 45</em>(17), Article 1490. <a href="https://doi.org/10.1007/s12144-026-09991-w" rel="noopener noreferrer">https://doi.org/10.1007/s12144-026-09991-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12144-026-09991-w" rel="noopener noreferrer">10.1007/s12144-026-09991-w</a></p>
<p><strong>Keywords:</strong> voice cloning, artificial intelligence, prosody, speech synthesis, vocal confidence, spectral flux, machine learning, human-AI interaction, acoustic analysis, eGeMAPS, fundamental frequency, Current Psychology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197043</post-id>	</item>
	</channel>
</rss>
