Sunday, September 20, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New Bilingual Speech Dataset Takes Aim at AI’s Weakest Spot: Code-Switching

September 20, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
New Bilingual Speech Dataset Takes Aim at AI’s Weakest Spot: Code-Switching

New Bilingual Speech Dataset Takes Aim at AI's Weakest Spot: Code-Switching

New Bilingual Speech Dataset Takes Aim at AI's Weakest Spot: Code-Switching

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

When bilingual speakers chat with one another, they rarely stay inside a single language. A sentence that begins in Mandarin may slip mid-phrase into English and back again, a behaviour linguists call code-switching. It is one of the most natural things multilingual people do, and one of the most unnatural things for machines to understand. Automatic speech recognition (ASR) systems, even the most powerful transformer-based models now in wide use, tend to stumble exactly at the point where one language hands off to another. A new openly available corpus called DOTA-ME-CS, short for Daily Oriented Text Audio Mandarin-English Code-Switching dataset, has been created to give researchers the fuel they need to close that gap.

The dataset, described in the Journal of Ambient Intelligence and Humanized Computing, contains 18.54 hours of audio spanning 9300 recordings produced by 34 bilingual participants, all of them fluent in both Mandarin and English. Unlike many earlier corpora, every single utterance in the collection involves code-switching. That design choice matters. Older resources such as SEAME, which stretches across roughly 190 hours, and TALCS, which covers 587 hours, contain substantial proportions of monolingual speech, meaning their effective supply of genuinely code-switched material is far smaller than their total length suggests. Some of those datasets, including the ASRU and TALCS corpora, are no longer publicly accessible at all, leaving the field with a shortage of usable, openly available benchmarks.

The construction of DOTA-ME-CS follows an unusual pipeline that blends large language model generation with human recording. The team used GPT-4o with carefully engineered prompts to produce scripted sentences across ten everyday scenarios: education, entertainment, environmental protection, food, health, home, life, pets, travel and work. Each prompt required the model to produce sentences that mimic daily conversational style, contain more English words than Mandarin words, and include at least one Mandarin word, with a dominant language assigned to every sentence. The authors justify this topic-anchored approach with a probabilistic argument: when a specific category is given, the probability of generating a relevant, high-quality sentence is higher than when the model is left to roam across all possible topics, which also reduces hidden cultural bias in the resulting scripts.

Human evaluators checked the generated scripts for grammatical problems and confirmed the presence of genuine switching, and a post-hoc naturalness study asked eight bilingual raters to score 100 randomly sampled stimuli on a five-point Likert scale. The average rating came in at 4.12 with a standard deviation of 0.61, and the median was 4.0, indicating that bilingual listeners generally perceive the generated sentences as natural, though the authors acknowledge that occasional awkwardness is an inherent limitation of LLM-based text generation. Bilingual volunteers, mostly college students with academic backgrounds in China and the United Kingdom, including roughly eighteen participants based at Imperial College London, then recorded the scripts on their own laptops or smartphones as 16-bit WAV files in quiet indoor settings. Recordings that failed basic quality checks were rejected and re-recorded, and accepted files were peak-normalised to 3 dBFS for consistent loudness.

The participants’ linguistic profiles were deliberately diverse. Twelve reported English dominance, eighteen reported Mandarin dominance and four identified as balanced bilinguals, and the mix of speakers from China and the United Kingdom ensures that the corpus spans multiple varieties of English and second-language accents. An average of 2.22 switching points per utterance, combined with a broad part-of-speech distribution across both languages, means the recordings capture switching at varied syntactic positions and grammatical categories rather than concentrating it in a single predictable spot. The dataset also includes longer recordings of roughly 10 to 15 seconds and about 100 words, an intentional response to the weakness current ASR models show on extended speech, where dependencies and critical information can be lost.

What sets the corpus apart most sharply is its AI-driven augmentation. Because human recordings were captured in quiet conditions at a normal pace, the researchers used the Librosa audio library to modify a random subset of clips. Playback speed was adjusted to 0.75x, 0.5x, 1.25x, 1.5x and 2x, with 200 recordings modified at each setting. Five categories of background noise, drawn from highways, war, natural sounds, white noise and playground environments, were added at 200 recordings per type, with the noise pitch scaled to 0.8x to account for the Lombard effect, the well-documented tendency of speakers to raise their vocal intensity in noisy surroundings. In addition, five AI-generated timbres, two male and three female, replaced the original voices in 200 recordings each, simulating both voice conversion deepfake scenarios and privacy-preserving speech processing conditions.

The accompanying data analysis is unusually thorough. Phoneme distributions catalogued with the International Phonetic Alphabet reveal that Mandarin contributes more multisyllabic pronunciations, tones, aspirated consonants and palatalised sounds, while English favours single-syllable vowels; shared features such as the sounds t and a suggest cues that future switching-detection models could exploit. Physical measurements show a frame rate of 46.7 kHz and a sample width of 2.0 bytes, meeting professional audio standards, with an average maximum pitch of 3617.13 Hz and an average per-recording pitch of 535.04 Hz. Formant frequencies computed through Linear Predictive Coding yield averages of 804.02 Hz, 4419.15 Hz and 7549.48 Hz for F1, F2 and F3, pointing to mid-to-low and front vowels, reduced lip rounding and retroflex consonants. The measured speaking rate of 2.05 words per second, with a standard deviation of 0.49, sits squarely within the typical range for conversational read speech, supporting the corpus’s ecological validity.

To establish baseline performance, the team benchmarked three pre-trained ASR models on the full dataset without fine-tuning: Whisper large-v3, a 1.5-billion-parameter Transformer encoder-decoder trained on 680,000 hours of multilingual audio; SenseVoice Small, a lightweight hybrid CTC-attention model of roughly 200 million parameters that also performs language identification, emotion detection and audio event detection; and Paraformer, a non-autoregressive Transformer with 220 million parameters optimised for Mandarin and built on continuous integrate-and-fire alignment. Paraformer and SenseVoice achieved the best results, with Whisper slightly behind, and a word error rate of 0.177 alongside a character error rate of 0.176. Compared against the publicly available ASCEND corpus, the models produced higher error rates on DOTA-ME-CS, which the authors read not as a defect but as evidence that their dataset poses a more demanding and therefore more useful benchmark.

Case studies expose exactly where today’s systems fail. In one example embedding the Chinese noun 菠萝包, meaning pineapple bun, inside an English sentence, Paraformer and SenseVoice produced garbled outputs such as bullleball and boobao, while Whisper translated the phrase correctly but destroyed the mixed-language structure by refusing to preserve the code-switch itself, a translation bias that also appeared on noise-free recordings. At double playback speed, Paraformer’s output strayed entirely from the intended meaning, SenseVoice descended into repetitions such as Miamiami, and Whisper mangled Florida lifestyle into Forensic Lifestyle. Altered timbres degraded English recognition across the board, and all three models performed worst at 2.0x speed. Whisper proved the most resilient to background noise, while Paraformer was the most sensitive to voice changes.

The implications reach beyond engineering. Reliable code-switching recognition would make automated transcription, voice assistants and translation systems far more useful for the hundreds of millions of people who live their linguistic lives between two languages, reducing a bias that currently disadvantages non-monolingual speakers. The authors stress that the study received ethical approval from Imperial College London’s ethical board, that all participants gave informed consent, and that privacy protections were maintained throughout. They are candid about limitations: funding constrained the scale of the corpus, and scripted reading, while reducing transcription errors and ethical risks, sacrifices the spontaneity of natural conversation. Even so, their deliberate trade-off of quality and coverage over raw scale gives the community something it has lacked, a fully public, exclusively code-switching Mandarin-English benchmark with baseline scores, rigorous acoustic analysis and data, code and recordings all released through a public repository. As speech technology races toward multilingual ubiquity, datasets like this one may determine whether the next generation of voice interfaces can finally follow the way people actually talk.

Subject of Research: A Mandarin-English code-switching speech dataset for advancing automatic speech recognition research.

Article Title: DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset

Article References: Li, Y., Wei, Z., Yu, H., Xue, J., Zhou, H., & Schuller, B. W. (2026). DOTA-ME-CS: daily oriented text audio-Mandarin English-Code switching dataset. Journal of Ambient Intelligence and Humanized Computing. https://doi.org/10.1007/s12652-026-05119-x

Image Credits: AI Generated

DOI: 10.1007/s12652-026-05119-x

Keywords: code-switching, automatic speech recognition, Mandarin, bilingual speech, speech dataset, Whisper, Paraformer, SenseVoice, data augmentation, phonetics, large language models, voice conversion

Cite Scienmag News

Denise Maddox. (September 20, 2026). New Bilingual Speech Dataset Takes Aim at AI’s Weakest Spot: Code-Switching. Scienmag. https://scienmag.com/new-bilingual-speech-dataset-takes-aim-at-ais-weakest-spot-code-switching/

Denise Maddox. "New Bilingual Speech Dataset Takes Aim at AI’s Weakest Spot: Code-Switching." Scienmag, 20 September 2026, https://scienmag.com/new-bilingual-speech-dataset-takes-aim-at-ais-weakest-spot-code-switching/. Accessed 20 September 2026.

Denise Maddox. "New Bilingual Speech Dataset Takes Aim at AI’s Weakest Spot: Code-Switching." Scienmag. September 20, 2026. https://scienmag.com/new-bilingual-speech-dataset-takes-aim-at-ais-weakest-spot-code-switching/

Tags: automatic speech recognitionautomatic speech recognition challengesbilingual speechbilingual speech recognitioncode-switchingcode-switching datasetdata augmentationDOTA-ME-CS corpusimproving machine understanding of code-switchinglarge language modelsMandarinMandarin-English code-switchingmultilingual natural language processingmultilingual speech processingopen-source language datasetsParaformerphoneticsSenseVoicespeech datasetspeech dataset for bilingual speakersspeech recognition for code-switchingtransformer-based ASR modelsvoice conversionWhisper
Share26Tweet16
Previous Post

Zinc-Doped Carbon Dots Shield Sperm Cells From Microplastic Damage

Next Post

Hair Follicles Mailed in a Kit Yield Stem Cells and Mini Brains

Related Posts

Zinc-Doped Carbon Dots Shield Sperm Cells From Microplastic Damage
Technology and Engineering

Zinc-Doped Carbon Dots Shield Sperm Cells From Microplastic Damage

September 20, 2026
New Optical Windows Let Scientists Watch the Living Mouse Brain for Weeks
Technology and Engineering

New Optical Windows Let Scientists Watch the Living Mouse Brain for Weeks

September 20, 2026
Laser Technique Maps Swelling Inside Organic Transistor Channels with Submicrometre Precision
Technology and Engineering

Laser Technique Maps Swelling Inside Organic Transistor Channels with Submicrometre Precision

September 20, 2026
Tiny Transformer Reads 95 Million Student Clicks in Minutes, Explains Its Predictions in Real Time
Technology and Engineering

Tiny Transformer Reads 95 Million Student Clicks in Minutes, Explains Its Predictions in Real Time

September 20, 2026
Gallium Nitride Transistor Probes Strip Light Artifacts From Optogenetic Brain Recordings
Technology and Engineering

Gallium Nitride Transistor Probes Strip Light Artifacts From Optogenetic Brain Recordings

September 20, 2026
AI Takes the Wheel in the Quest to Mass-Produce Atomically Thin Materials
Technology and Engineering

AI Takes the Wheel in the Quest to Mass-Produce Atomically Thin Materials

September 20, 2026
Next Post
Smarter Forecasting Could Ease Climate Adaptation Trade-Offs for African Herders

Smarter Forecasting Could Ease Climate Adaptation Trade-Offs for African Herders

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Female and Young Rats Absorb Far More Radioactive Iodine in the Thyroid Than Adult Males
  • Machine Learning Reveals What Drives Soil CO₂ Emissions in Semi-Arid India
  • Smarter Forecasting Could Ease Climate Adaptation Trade-Offs for African Herders
  • Hair Follicles Mailed in a Kit Yield Stem Cells and Mini Brains

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading