<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>acoustic analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/acoustic-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 21:22:42 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>acoustic analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Chimpanzee Food Calls Reveal Surprisingly Detailed Secrets About Hidden Food</title>
		<link>https://scienmag.com/chimpanzee-food-calls-reveal-surprisingly-detailed-secrets-about-hidden-food/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:22:42 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[acoustic analysis]]></category>
		<category><![CDATA[acoustic specificity]]></category>
		<category><![CDATA[animal alarm and food calls]]></category>
		<category><![CDATA[animal behavior research]]></category>
		<category><![CDATA[animal cognition]]></category>
		<category><![CDATA[animal cognition studies]]></category>
		<category><![CDATA[animal communication]]></category>
		<category><![CDATA[Chimpanzee food calls]]></category>
		<category><![CDATA[chimpanzees]]></category>
		<category><![CDATA[comparative psychology of communication]]></category>
		<category><![CDATA[food calls]]></category>
		<category><![CDATA[food discovery vocalizations]]></category>
		<category><![CDATA[foraging behaviour]]></category>
		<category><![CDATA[functional reference]]></category>
		<category><![CDATA[language evolution]]></category>
		<category><![CDATA[non-human language precursors]]></category>
		<category><![CDATA[playback experiment]]></category>
		<category><![CDATA[primate communication]]></category>
		<category><![CDATA[primatology]]></category>
		<category><![CDATA[referential signalling]]></category>
		<category><![CDATA[referential signals in animals]]></category>
		<category><![CDATA[rough grunts]]></category>
		<category><![CDATA[vocal communication]]></category>
		<category><![CDATA[vocalization analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202748</guid>

					<description><![CDATA[A new playback study shows chimpanzee food calls encode both the value and the specific identity of foods, and that listeners can decode that fine-grained information.]]></description>
										<content:encoded><![CDATA[<p>In the dense forests and captive colonies where chimpanzees live out their daily lives, one of the most common sounds is a burst of rough, throaty grunting when an individual discovers food. For decades, researchers regarded these so-called food calls as little more than emotional leakage — an involuntary expression of excitement that said something about how much the caller wanted the food, but little else. A new playback experiment, published in the journal Animal Cognition, upends that comfortable assumption. The study, led by Stuart K. Watson of the University of Zurich together with colleagues including Katie E. Slocombe of the University of York, Klaus Zuberbühler and Josep Call of the University of St Andrews, and Simon W. Townsend of the University of Zurich, demonstrates that chimpanzee food calls carry a degree of acoustic specificity that has, until now, been documented almost exclusively in alarm call systems. Listeners, the findings show, can extract not only how valuable a food is, but which particular food the caller has found.</p>
<p>Functionally referential signals — vocalizations that reliably inform listeners about specific external events or objects — have long fascinated comparative psychologists because they blur the boundary between animal communication and human language. Vervet monkeys emit distinct alarm calls for leopards, eagles and snakes, and recipients respond appropriately to each. Prairie dogs embed information about predator features in their calls, and some birds encode details about food quality. But the question of just how fine-grained these references can be has remained open. Most documented referential systems distinguish broad categories: predator versus non-predator, high-value versus low-value food. Whether a signal can pick out individual food types within a category — whether a chimp&#8217;s grunt for bread sounds meaningfully different from its grunt for mango — had never been empirically tested for chimpanzee food calls.</p>
<p>Chimpanzees produce rough grunts when finding and eating food, and earlier work had established that the acoustic structure of these calls varies with the preference value of the food that elicits them. Grunts given to highly preferred foods differ systematically from those given to less desirable items, and subsequent research showed that listening chimpanzees can use this variation to guide their own foraging decisions. That established a meaningful baseline: the calls encode relative preference. What remained unknown was whether the specificity stopped there. If two foods are both highly preferred — say, bread and mango — do the grunts they elicit sound the same to a listener, or do they carry information distinguishing one food from the other? And critically, does that fine-grained information actually matter to receivers, or is it merely acoustic noise that no chimpanzee attends to?</p>
<p>To answer these questions, the research team, whose work was supported by the Biotechnology and Biological Sciences Research Council and by the Swiss National Science Foundation through the NCCR Evolving Language programme, combined two complementary approaches. The first was a careful acoustic analysis of food calls recorded from chimpanzees at the Wolfgang Köhler Primate Research Centre in Leipzig, Germany, where calls elicited by foods of matched preference values could be compared. The second was a playback experiment — a technique in which recorded vocalizations are broadcast to an animal while researchers observe how it responds — allowing the team to test whether the acoustic distinctions they detected were actually meaningful to a listening chimpanzee rather than incidental by-products of arousal.</p>
<p>The acoustic analysis produced a strikingly layered picture. The calls differed between foods of high and low preference value, replicating and extending the earlier findings that grunts track how much a caller likes what it has found. But the analysis went further: significant acoustic differences also emerged between individual food types within each of those preference categories. In other words, grunts elicited by different high-value foods could be statistically discriminated from one another, and so could grunts elicited by different low-value foods. The calls were not simply graded along a single dimension of excitement. Instead, they appeared to encode at least two distinct dimensions of information simultaneously — the general value of the food and its specific identity — a combination that suggests a far richer encoding system than anyone had demonstrated in chimpanzee feeding vocalizations.</p>
<p>Statistically detectable differences, however, do not automatically translate into communicative relevance. An acoustic signature can reflect physiological differences in arousal, jaw position, or food texture in the mouth without any listener ever extracting information from it. This is where the playback design proved decisive. The researchers broadcast recorded food calls to a chimpanzee subject and measured the responses, focusing on whether the animal&#8217;s behaviour indicated that it had interpreted the calls as information about what lay hidden at the playback location. By pairing calls recorded in one context with search opportunities in another, the experiment isolated the informational content of the vocalization itself from any direct sensory experience of the food.</p>
<p>The results were remarkable on all three fronts. The listening chimpanzee discriminated between calls elicited by high-value and low-value foods, confirming once again that preference information is accessible to receivers. More importantly, the subject also discriminated between calls elicited by individual types of high-value food, and between calls elicited by individual types of low-value food. The fine-grained distinctions that showed up in the acoustic analysis were not lost on the audience: the receiver behaved as though it knew not just that something good had been found, but what kind of good thing it was. This represents, the authors write, an unprecedented degree of specificity in the referential food calls of our closest living relatives — a level of detail previously identified only in alarm call systems, not in feeding vocalizations.</p>
<p>The theoretical implications ripple outward in several directions. For theories of language evolution, the finding narrows the perceived gap between primate vocal communication and human speech. One cornerstone of linguistic meaning is the ability of a signal to pick out specific referents rather than broad emotional states, and this study shows that chimpanzee calls can do something analogous with food types, not just food categories. It also challenges the long-standing division between &#8216;referential&#8217; alarm calls and &#8217;emotional&#8217; food calls: if food calls encode specific identities in a way receivers use, the referential-emotional dichotomy looks increasingly like a human-imposed simplification rather than a real feature of chimpanzee communication. The study adds to a growing body of evidence that the vocal repertoire of great apes is far more structured than the behaviorist traditions of the twentieth century assumed.</p>
<p>The methodology deserves attention too. Combining acoustic analysis with playback experiments is the gold standard for demonstrating functional reference, because each approach compensates for the other&#8217;s weaknesses. Acoustic analysis can reveal structure that receivers ignore; playback can reveal comprehension of differences the analysis missed. By requiring both — statistically discriminable calls and demonstrably different receiver responses — the researchers built a case that is difficult to dismiss. The work was approved by the School of Psychology Ethics committee at the University of St Andrews, and the authors declare no competing interests. The study was published open access, with a citable, permanent DOI, meaning anyone can examine the evidence in full.</p>
<p>Questions naturally remain. The playback tested a single subject, so the generality of comprehension across individuals and communities will need confirmation. It is also unclear how the specificity arises — whether callers intentionally vary their grunts to inform others, whether the acoustic differences emerge as side effects of eating different foods, or whether listeners have simply learned to exploit regularities in the sounds. None of this diminishes the central result: chimpanzees listening to a distant group member&#8217;s grunts can apparently learn both how desirable a hidden food is and what that food actually is. For a species separated from our own lineage by roughly six million years, that is a strikingly sophisticated channel of information, and it suggests that the seeds of referential specificity run far deeper in the primate lineage than the study of alarm calls alone had ever revealed.</p>
<p><strong>Subject of Research:</strong> Referential specificity of chimpanzee food calls tested through acoustic analysis and playback experiments</p>
<p><strong>Article Title:</strong> Chimpanzee food calls provide information about the value and type of food: a playback study</p>
<p><strong>Article References:</strong> Watson, S. K., Kaller, T., Cheng, L., Wathan, J., Townsend, S. W., Zuberbühler, K., Call, J., &amp; Slocombe, K. E. (2026). Chimpanzee food calls provide information about the value and type of food: a playback study. <em>Animal Cognition</em>. <a href="https://doi.org/10.1007/s10071-026-02105-w" rel="noopener noreferrer">https://doi.org/10.1007/s10071-026-02105-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10071-026-02105-w" rel="noopener noreferrer">10.1007/s10071-026-02105-w</a></p>
<p><strong>Keywords:</strong> chimpanzees, food calls, functional reference, rough grunts, playback experiment, vocal communication, animal cognition, primatology, language evolution, acoustic analysis, foraging behaviour, referential signalling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202748</post-id>	</item>
		<item>
		<title>AI Voice Clones Can Now Copy the Human Sound of Confidence and Doubt</title>
		<link>https://scienmag.com/ai-voice-clones-can-now-copy-the-human-sound-of-confidence-and-doubt/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 17:41:38 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[acoustic analysis]]></category>
		<category><![CDATA[acoustic features of confidence and doubt]]></category>
		<category><![CDATA[AI voice cloning]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Current Psychology]]></category>
		<category><![CDATA[eGeMAPS]]></category>
		<category><![CDATA[experimental design in voice perception studies]]></category>
		<category><![CDATA[fundamental frequency]]></category>
		<category><![CDATA[human expression of confidence and doubt]]></category>
		<category><![CDATA[Human-AI Interaction]]></category>
		<category><![CDATA[human-AI voice interaction research]]></category>
		<category><![CDATA[implications for conversational AI development]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[prosody]]></category>
		<category><![CDATA[psychological experiments with voice stimuli]]></category>
		<category><![CDATA[spectral flux]]></category>
		<category><![CDATA[speech emotion recognition]]></category>
		<category><![CDATA[speech prosody analysis]]></category>
		<category><![CDATA[speech synthesis]]></category>
		<category><![CDATA[synthetic voice emotional mimicry]]></category>
		<category><![CDATA[vocal confidence]]></category>
		<category><![CDATA[voice cloning]]></category>
		<category><![CDATA[voice cloning fidelity]]></category>
		<category><![CDATA[voice identity versus emotional expression]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197043</guid>

					<description><![CDATA[A new study shows that AI voice cloning systems preserve not only speaker identity but also the human-specific prosodic cues of confidence and doubt.]]></description>
										<content:encoded><![CDATA[<p>Voice-cloning artificial intelligence has long been judged by how faithfully it reproduces the identity of a speaker: the recognizable timbre, pitch range, and resonance that make a voice sound like one particular person rather than another. A new study published in Current Psychology by Wenjun Chen and Xiaoming Jiang of Shanghai International Studies University, with Chen also affiliated with McGill University, asks a subtler and arguably more consequential question. Can AI voice clones reproduce not just who a speaker is, but how that speaker sounds when expressing something human-specific, such as confidence or doubt? The answer, based on a rigorous acoustic analysis of thousands of utterances, is largely yes, and the finding has immediate implications for how psychologists build experiments on human-AI voice interaction.</p>
<p>The motivation for the study stems from a methodological bottleneck. As voice assistants, synthetic narrators, and conversational agents become embedded in daily life, researchers increasingly need experimental stimuli in which speaker identity and prosodic style can be manipulated independently. If a scientist wants to test how listeners react to a confident versus a doubtful statement, the comparison is confounded when different human speakers deliver the confident and doubtful versions, because listeners may respond to the person rather than the prosody. Voice cloning promises a solution: train a model on one speaker&#8217;s voice, then generate new sentences carrying a target emotional or attitudinal tone. But this only works if the cloning system genuinely transfers prosody rather than flattening it into a generic synthetic delivery.</p>
<p>To test this, the researchers recruited ten native Mandarin speakers who each produced thirty sentences, fifteen drawn from geography statements and fifteen from trivia statements, in three distinct intonations: confident, doubtful, and neutral. This design yielded a rich corpus of human recordings spanning different sentence contents and prosodic intentions. For each speaker and each prosody, the team trained separate AI voice clones using either the geography recordings or the trivia recordings as training material. Each clone was then asked to generate all thirty sentences, producing a total of 2,700 utterances for analysis: 900 human recordings, 900 AI-generated utterances from geography-trained clones, and 900 from trivia-trained clones. Crucially, the training and generation sets were crossed, so a clone trained only on confident geography sentences still had to produce doubtful trivia sentences, forcing the system to generalize prosodic style to entirely novel content.</p>
<p>The analytical strategy combined machine learning classification with traditional statistical modeling. The researchers extracted 88 acoustic features using the extended Geneva Minimalistic Acoustic Parameter Set, or eGeMAPS, a standardized toolkit widely used in voice research and affective computing. They then trained gradient-boosted tree classifiers, a machine learning approach implemented through the XGBoost framework, to distinguish confident from doubtful prosody within each source of speech. Within a single source, whether human or AI, the classifiers achieved high accuracy, averaging 0.85, indicating that the acoustic signatures of confidence and doubt are robust and detectable in both natural and synthetic voices. More importantly, when classifiers trained on human speech were tested on AI speech, and vice versa, accuracy remained substantially above chance, averaging 0.65. This cross-source generalization demonstrates that the prosodic cues of confidence and doubt are encoded in a shared acoustic currency that transcends the boundary between human and machine voices.</p>
<p>Among the 88 features, one emerged as a consistently dominant contributor: spectral flux, a measure of how rapidly the frequency content of the speech signal changes over time. Spectral flux captures the crispness and dynamism of articulation, and confident speech, in both humans and their AI clones, exhibited higher spectral flux than doubtful speech. This makes intuitive sense. A confident speaker articulates with decisive energy, producing sharper spectral transitions, whereas a hesitant speaker tends toward softer, less defined acoustic edges. Notably, spectral flux showed no effect of speaker sex, which the authors interpret as evidence that it indexes prosodic style rather than stable anatomical differences between speakers. In other words, it is a marker of how something is said, not of who is saying it, which is precisely the kind of feature a stimulus designer would want to manipulate independently of speaker identity.</p>
<p>Linear mixed-effects models, a statistical framework that accounts for the nested structure of the data with multiple speakers, sentences, and prosodies, confirmed a coherent acoustic profile of vocal confidence. Across human and AI speakers alike, confident prosody was associated with higher spectral flux, a longer estimated vocal tract length, and lower fundamental frequency, the acoustic correlate of pitch, compared with doubtful prosody. The estimated vocal tract length, derived from formant frequencies, reflects how speakers shape their resonating cavities; a longer apparent vocal tract projects a larger, more authoritative body, echoing a well-documented literature on vocal size exaggeration in humans. That AI clones reproduced this constellation of cues, including the anatomical-sounding shift in apparent vocal tract size, suggests that modern cloning systems capture not only surface acoustics but the embodied gestalt of an attitudinal vocal stance.</p>
<p>The study also probed how prosodic information unfolds in time. Using time-resolved decoding of fundamental frequency, the researchers found that the distinction between confident and doubtful prosody was most reliably recovered in the late windows of each utterance, across human and AI sources alike. This late-emerging pattern aligns with the intuition that speakers often commit to an attitude as a sentence progresses, with final intonational contours sealing the impression of certainty or hesitation. However, a subtle asymmetry appeared: human speakers distributed their prosodic markers more broadly across the utterance, weaving cues of confidence and doubt throughout the signal, whereas the AI clones concentrated their prosodic differentiation more narrowly toward the end. This temporal signature may represent one of the remaining acoustic fingerprints distinguishing synthetic from natural expressive speech.</p>
<p>Indeed, the authors are careful to note that despite the impressive prosodic fidelity, AI-generated speech still forms a distinguishable acoustic population relative to human speech. The cross-source classification accuracy of 0.65, while well above chance, falls short of the 0.85 achieved within sources, indicating that machine and human voices are similar but not acoustically identical. This finding resonates with recent work showing that voice clones sound realistic but not yet hyperrealistic, and that human listeners and neural systems can, under some conditions, separate deepfake from genuine speaker identity. For the immediate purposes of psychological research, however, the residual human-machine gap may even be an advantage, since it allows researchers to verify that their synthetic stimuli behave acoustically as intended while retaining a measurable distinction from natural recordings.</p>
<p>The broader significance of the study lies in its methodological contribution. By demonstrating that voice cloning preserves prosodic style alongside speaker identity, Chen and Jiang provide empirical license for a new generation of controlled experiments on human-AI voice interaction. Researchers can now, with appropriate validation, generate stimulus sets in which the same voice delivers the same content with systematically varied confidence, doubt, or neutrality, isolating the causal impact of prosody on listener trust, memory, persuasion, and social perception. As AI voices become conversation partners, teachers, and companions, understanding how their vocal expressions of certainty are produced, perceived, and potentially distinguished from human ones becomes a scientific priority. This study shows that the tools for doing that science rigorously are already within reach, and that the line between human and machine vocal expression, while still detectable, is growing finer by the year.</p>
<p><strong>Subject of Research:</strong> Whether AI voice cloning systems can reproduce human prosodic expressions of confidence and doubt</p>
<p><strong>Article Title:</strong> Voice-cloning artificial-intelligence speakers can also mimic human-specific vocal expression</p>
<p><strong>Article References:</strong> Chen, W., &amp; Jiang, X. (2026). Voice-cloning artificial-intelligence speakers can also mimic human-specific vocal expression. <em>Current Psychology, 45</em>(17), Article 1490. <a href="https://doi.org/10.1007/s12144-026-09991-w" rel="noopener noreferrer">https://doi.org/10.1007/s12144-026-09991-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12144-026-09991-w" rel="noopener noreferrer">10.1007/s12144-026-09991-w</a></p>
<p><strong>Keywords:</strong> voice cloning, artificial intelligence, prosody, speech synthesis, vocal confidence, spectral flux, machine learning, human-AI interaction, acoustic analysis, eGeMAPS, fundamental frequency, Current Psychology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197043</post-id>	</item>
	</channel>
</rss>
