Anyone who has spent an afternoon struggling to follow a colleague’s heavily accented English knows the strange magic that happens a few days later: suddenly, other accented speakers seem easier to understand. Psychologists call this perceptual adaptation, a form of implicit learning in which the auditory system recalibrates itself to the systematic distortions of non-native speech. Decades of research have established that when listeners are exposed to multiple second-language talkers, they often become better at understanding second-language speech in general, a phenomenon known as cross-talker generalization. But a new study published in the journal Attention, Perception, & Psychophysics delivers a sobering and fascinating message about the single-talker case: training on just one accented voice can indeed transfer to new voices, yet the two most obvious explanations for when and why that transfer happens both fail to hold up under scrutiny.
The study, led by Seung-Eun Kim of Yonsei University together with Matthew Goldrick, Joseph Keshet, and Ann R. Bradlow of Northwestern University and the Technion-Israel Institute of Technology, tackled a puzzle that has dogged the field for years. When native English listeners adapt to one non-native talker, does that adaptation help them understand a different non-native talker they have never heard before? The literature is inconsistent. Some experiments report robust generalization from a single training voice; others find that the learning stays stubbornly locked to the specific talker. The new work was designed to adjudicate between two competing theoretical accounts, and to do so with an unusually rigorous, data-driven methodology.
The first hypothesis centers on intelligibility. Perhaps, the researchers reasoned, generalization is driven by how easy the training talker is to understand. A highly intelligible talker might provide cleaner, more reliable evidence about how English sounds are realized in non-native speech, allowing the listener’s perceptual system to extract general rules rather than talker-specific quirks. Under this account, training with a clear, easily understood accented speaker should produce stronger benefits for understanding other accented speakers than training with a hard-to-understand one. The second hypothesis centers on acoustic similarity. Perhaps generalization depends on how closely the training talker resembles the test talker in the physical details of their speech. If the perceptual system tunes to specific acoustic properties, then a training voice that sounds similar to a test voice should yield better transfer than a dissimilar pairing. This account echoes earlier findings by Xie and Myers, who showed that acoustic similarity between talkers constrains how far foreign-accent adaptation spreads.
Testing these hypotheses properly required something the field has rarely had: precise, quantitative measurements of both intelligibility and acoustic similarity for a large pool of second-language talkers. The team drew on the ALLSSTAR archive, a Northwestern University repository of recordings from more than 100 speakers of English as a second language, each of whom had been assessed with empirical intelligibility measures. For the similarity dimension, the researchers turned to a cutting-edge computational tool: self-supervised speech models, neural networks of the same family as HuBERT that learn rich representations of speech sound by being trained on massive amounts of audio without labels. Distances between talkers in the representational space of such models have been shown in prior work by the same group to capture perceptually meaningful acoustic similarity, providing a principled way to measure how alike any two speakers sound to a machine that has learned the statistics of human speech.
From this large dataset, the researchers constructed twenty training-test talker pairs, deliberately engineered so that the pairs spanned a wide range of intelligibility levels and acoustic similarity values. Each pair was then evaluated by ten native English listeners in a sentence-in-noise transcription task. The design was an exposure-test format: listeners first heard sentences produced by the training talker, giving their perceptual system a chance to adapt, and then transcribed sentences from the test talker, a different accented speaker they had not encountered before. Accuracy on those test sentences, measured as the proportion of words correctly transcribed against the noise, served as the index of cross-talker generalization. The inclusion of background noise was not incidental. Noise is one of the most ecologically common and cognitively demanding listening conditions, and prior work has shown that understanding accented speech in noise is substantially harder than in quiet, making any adaptation benefit both more meaningful and more measurable.
The headline result was genuinely encouraging: most of the twenty talker pairs showed cross-talker generalization. In other words, brief exposure to a single second-language talker was frequently sufficient to improve listeners’ comprehension of a completely different second-language talker. This finding matters practically. It suggests that single-talker training paradigms, which are far cheaper and simpler to implement than multi-talker regimes, can still deliver real-world listening benefits, whether in language education, in preparing listeners for communication in multilingual workplaces, or in clinical contexts where listeners must adjust to unfamiliar voices.
But then comes the twist that gives the study its title. When the researchers asked whether the size of the generalization effect was predicted by the intelligibility of the training talker, the answer was no. When they asked whether it was predicted by the acoustic similarity between training and test talkers, again the answer was no. Neither variable reliably drove the magnitude of transfer across the twenty pairs. The effects, in statistical terms, refused to line up with either theoretical account. The researchers had planned a replication strategy, intending to test additional talker triplets after the initial set, but because no clear patterns emerged in the first batch, the planned follow-up analyses were not pursued, and the untested triplets’ data were deposited in the project’s Open Science Framework repository for other researchers to examine.
This null result is more provocative than it might first appear. In science, a clean failure of two leading hypotheses is a map of where not to dig, and it forces the field to consider what else could be carrying the load. One possibility is that generalization depends on properties of the listener rather than the talker: individual differences in working memory, attention, or prior experience with accented speech are known to shape how people cope with difficult listening conditions, and such listener-side factors could swamp talker-side predictors in samples of ten listeners per pair. Another possibility is that the relevant similarity is not global acoustic distance but something more structured and linguistic, such as overlap in specific error patterns, for example whether two talkers share the same substitutions of particular English vowels or consonants. Two speakers could be acoustically distant overall yet confusable in exactly the phonemic categories that matter for transcription, a hypothesis that coarse representational distances might miss.
The study also carries a methodological lesson that extends well beyond speech perception. The use of self-supervised model representations to quantify acoustic similarity represents a broader shift in the cognitive sciences, where large-scale neural models trained on natural data are increasingly used as measurement instruments for psychological constructs. The same approach has recently been used by other groups, including work presented by Jin, Zhu, and Jaeger, to predict generalization of adaptation across talkers from latent speech representations. Yet the current findings temper the enthusiasm: even a sophisticated, data-driven similarity metric grounded in state-of-the-art speech models did not predict human generalization in this paradigm. The representations may capture acoustic similarity faithfully, but what the human perceptual system extracts from a brief encounter with one accented voice may be organized along dimensions that neither overall intelligibility nor global acoustic distance adequately describes.
What remains is a robust empirical anchor and an open question. The empirical anchor is that single-talker training works more often than not: listeners who adapt to one second-language voice typically carry some of that benefit to new voices, even under the adverse condition of background noise. The open question is the mechanism. The authors are candid that the predictors remain elusive, and that candor is itself a contribution, redirecting the field away from tidy two-factor stories and toward richer models that incorporate listener variability, linguistic structure, training design, and the temporal dynamics of perceptual learning. For the millions of people who navigate multilingual environments daily, the practical takeaway stands: a little exposure goes a long way. For the scientists, the accent on the horizon is clear even if its coordinates are not. The data and materials from the study are publicly available, ensuring that the next generation of hypotheses about how the brain learns to understand unfamiliar speech can be tested against the same evidence.
Subject of Research: Cross-talker generalization of perceptual adaptation to second-language speech after single-talker training
Article Title: Generalization of perceptual adaptation to second-language speech in a single-talker training paradigm: Predictors remain elusive
Article References: Kim, S.-E., Goldrick, M., Keshet, J., & Bradlow, A. R. (2026). Generalization of perceptual adaptation to second-language speech in a single-talker training paradigm: Predictors remain elusive. Attention, Perception, & Psychophysics, 88(7), Article 192. https://doi.org/10.3758/s13414-026-03333-5
Image Credits: AI Generated
DOI: 10.3758/s13414-026-03333-5
Keywords: speech perception, perceptual learning, second-language speech, cross-talker generalization, single-talker training, speech intelligibility, acoustic similarity, self-supervised representations, accented speech, speech-in-noise, bilingualism, auditory adaptation
Cite Scienmag News
Glenn Wilkins. (September 26, 2026). One Voice Is Enough: How Listeners Learn to Understand Accented Speech, and Why Science Still Can’t Predict It. Scienmag. https://scienmag.com/one-voice-is-enough-how-listeners-learn-to-understand-accented-speech-and-why-science-still-cant-predict-it/
Glenn Wilkins. "One Voice Is Enough: How Listeners Learn to Understand Accented Speech, and Why Science Still Can’t Predict It." Scienmag, 26 September 2026, https://scienmag.com/one-voice-is-enough-how-listeners-learn-to-understand-accented-speech-and-why-science-still-cant-predict-it/. Accessed 26 September 2026.
Glenn Wilkins. "One Voice Is Enough: How Listeners Learn to Understand Accented Speech, and Why Science Still Can’t Predict It." Scienmag. September 26, 2026. https://scienmag.com/one-voice-is-enough-how-listeners-learn-to-understand-accented-speech-and-why-science-still-cant-predict-it/

