Wednesday, September 9, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Testing Transformer Models’ Emotion Recognition Across Languages and Cultures

September 9, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 6 mins read
0
Testing Transformer Models’ Emotion Recognition Across Languages and Cultures

Testing Transformer Models’ Emotion Recognition Across Languages and Cultures

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence systems that claim to understand human emotion are losing their grip the moment people start talking the way people actually talk. That is the central finding of a new study published in the journal Cognitive Computation, which put three of the world’s most widely used multilingual language models through a gauntlet of slang, idioms, sarcasm and half-Spanish, half-English internet speak. The results reveal consistent and sometimes dramatic drops in accuracy whenever emotional meaning depends on cultural context rather than dictionary definitions.

The research, led by William Villegas-Ch of Universidad de Las Américas in Ecuador together with colleagues from Universidad Internacional del Ecuador and Universidad Estatal de Milagro, focused on transformer-based models: multilingual BERT, known as mBERT, and XLM-R, a larger model built on the RoBERTa architecture. The team also tested a fine-tuned variant, XLM-R-FT, adapted to culturally diverse training data. These models sit behind countless applications, from customer-service chatbots to social media monitoring tools, and their perceived fluency in dozens of languages has made them the default choice for multilingual sentiment and emotion analysis.

The problem, the researchers argue, is that standard evaluations of these systems rely on structured, homogeneous datasets that bear little resemblance to real-world communication. Everyday emotional expression is messy. Speakers switch languages mid-sentence, lean on idioms whose meanings cannot be assembled from their parts, and convey feelings indirectly through irony and understatement. “Me hierve la sangre,” a Spanish phrase that literally translates to “my blood boils,” expresses anger through a figurative construction. “Estoy sad AF hoy” blends English and Spanish in a single clause. “Oh, great… just what I needed” communicates frustration through positive words. None of these patterns are well represented in the canonical corpora on which multilingual models are trained.

To measure how well the models cope, the team assembled an evaluation set of 4,892 instances drawn from public sources including the BBC Forum Dataset, the Brazilian offensive-comment corpus OFFCOMBR, Italian SENTIPOLC data and Twitter-based corpora. The set was balanced across four languages: English, Spanish, Portuguese and Italian. English served as a reference point because of its dominance in pre-training data, while the three Romance languages were chosen for their structural similarity paired with sharply different pragmatic and colloquial conventions around irony, attenuation and intensification. Each instance was categorized as idiomatic, code-switched, or indirect in structure.

Because the source datasets carried inconsistent labels, the researchers manually annotated a representative subset according to Ekman’s six basic emotions: joy, sadness, anger, fear, surprise and disgust. Native speakers with linguistic training performed the annotation, with two annotators independently labeling each sample. Agreement was required to reach a Cohen’s kappa of at least 0.8, and disagreements were resolved through adjudication or the samples were discarded, ensuring that the emotional ground truth was itself robust.

The experimental protocol was rigorous. All models ran on the HuggingFace Transformers library atop PyTorch, executed on NVIDIA A100 GPUs in a Linux computing cluster. Fine-tuning used the AdamW optimizer with a learning rate of 2 × 10⁻⁵, a batch size of 32, up to ten epochs, dropout regularization at 0.1 and early stopping when validation Macro-F1 stalled for three consecutive epochs. Evaluation used stratified five-fold cross-validation, with folds balanced jointly by language and emotion class so that no category was underrepresented in either training or testing.

The headline result is a pattern of consistent degradation on culturally marked input. When models moved from clean, monolingual sentences to code-switched or idiomatic ones, F1 scores fell by as much as 10 points and, in some configurations, between 13 and 18 points. In Spanish, mBERT dropped from an F1 of 0.82 on monolingual inputs to 0.67 on code-switched sentences and 0.66 on idiomatic expressions. Portuguese and Italian showed even larger declines. English, by contrast, retained average accuracy above 0.84 under clean conditions and degraded only modestly under noise.

The language gap tracks closely with how much of each language the models saw during pre-training. English, massively represented in training corpora, proved most stable. Spanish occupied an intermediate position. Portuguese and Italian, less richly represented, exhibited the largest drops and the widest variability, with Italian showing the greatest instability of all. The researchers computed performance deviations relative to English of −0.06 for Spanish, −0.09 for Portuguese and −0.12 for Italian, a gradient that reinforces what they describe as representation-driven bias: models are most reliable precisely where their training data was densest.

Specific emotional categories proved especially fragile. Joy and disgust showed the highest variability across languages and conditions. In Portuguese, accuracy on disgust fell from 0.82 on clean text to 0.69 on noisy text, while in Italian it reached the study’s lowest observed value at 0.67. The team also documented systematic confusion among anger, fear and sadness, with confusion rates of 0.25 between anger and fear and 0.20 between fear and sadness, suggesting that when emotional states are conveyed informally, the models’ decision boundaries between negative emotions blur badly.

Noise alone was enough to trip up the baseline models. When the researchers introduced perturbations mimicking real digital communication, emoji substitutions, social media abbreviations, minor spelling errors and emphatic reduplicated constructions, mBERT suffered accuracy losses of 0.13 in Portuguese and 0.15 in Italian. In one striking example, the Portuguese sentence “Fiquei muito fps com esso [disgusted face],” a deliberately corrupted expression of disgust, was classified as joy by mBERT, apparently because surface cues such as the emoji overwhelmed contextual meaning. The fine-tuned XLM-R-FT fared considerably better, keeping accuracy losses to within 0.05 across languages, but even it could not eliminate sensitivity to distortion.

Cultural interference from English emerged as a distinct failure mode. The researchers constructed adversarial examples using false cognates, expressions in Spanish, Portuguese or Italian that superficially resemble English words but carry different meanings. The Spanish phrase “estoy constipado,” which means “I have a cold” rather than what an English speaker might assume, could push the model toward incorrect negative emotional categories purely through lexical similarity. Across culturally ambiguous sentences, mBERT and XLM-R reached error rates of up to 25 percent in Italian and Spanish due to such interference, while the fine-tuned variant reduced this to below 15 percent.

Perhaps the most illuminating part of the study is its use of interpretability tools, LIME and SHAP, which attribute a model’s prediction to individual input tokens. These analyses showed that models frequently anchor their judgments on lexically salient words while ignoring context. In the Spanish expression “Estoy re quemado con esta vaina,” the model correctly weighted “quemado” as emotionally negative but also amplified the contribution of “vaina,” a context-dependent filler word with neutral semantic load. In Italian, multi-word idioms were decomposed into independent tokens, producing erroneous emotional assignments whenever the expression resisted compositional interpretation. By contrast, the Portuguese colloquialism “massa” was correctly associated with positive emotion, suggesting that models succeed mainly where colloquial usage happens to align with patterns in training data.

The performance gap between direct and indirect emotional expression proved remarkably uniform. All models showed F1 reductions of between 0.08 and 0.10 when moving from explicit statements of feeling to indirect or sarcastic formulations, with mBERT showing the largest swings. This, the authors conclude, reflects a bias toward explicit emotional patterns aligned with Anglophone linguistic structures, an Anglocentric tendency that persists even when the model is operating in another language.

Fine-tuning helps, but it is not a cure. XLM-R-FT outperformed both baselines across every emotional category, improving Macro-F1 by up to 0.14 points over mBERT in categories such as fear and disgust, and it limited degradation on idiomatic and code-switched input to less than 8 F1 points relative to monolingual sentences. Yet the gap never closed entirely. Adaptation, the study finds, mitigates sensitivity to linguistic variation without removing it, and the fine-tuned model retains whatever structural biases were baked into its original pre-training.

The authors are careful to note the study’s limits. Only three target languages were examined, without full coverage of dialectal and regional variation within each. Human annotation of the culturally sensitive subset introduces subjectivity despite consensus mechanisms, and interpretability methods such as LIME and SHAP rest on assumptions that may not fully capture a transformer’s internal computation. The evaluation also focused on classification, not generative models or interactive settings.

Even so, the implications reach well beyond the laboratory. Conversational agents, content moderation systems and social media analytics tools deployed in multilingual environments routinely encounter exactly the kind of language this study shows the models mishandling. A benchmark score earned on clean, structured data, the researchers warn, says little about behavior in the wild. Their proposed framework, combining controlled perturbations, cross-linguistic comparison and token-level interpretability, offers a way to test for those hidden fragilities before deployment. Future work, they argue, should focus on datasets that explicitly capture cultural variation across regions and registers, and on integrating external semantic resources that encode idiomatic and context-dependent meaning, so that the next generation of multilingual systems can finally read between the lines the way humans do.

Subject of Research: Evaluation of the semantic robustness and cultural adaptability of multilingual transformer models (mBERT, XLM-R and a fine-tuned XLM-R-FT) for emotion recognition across English, Spanish, Portuguese and Italian

Subject of Research: Technology and Engineering

Article Title: Multilingual Evaluation of Semantic Robustness and Cultural Adaptability in Transformer Models for Emotion Recognition

Article References: Villegas-Ch, W., Gutierrez, R., Mera-Navarrete, A., & Guevara-Reyes, R. (2026). Multilingual Evaluation of Semantic Robustness and Cultural Adaptability in Transformer Models for Emotion Recognition. Cognitive Computation, 18(1), Article 73. https://doi.org/10.1007/s12559-026-10621-7

Image Credits: AI Generated

DOI: 10.1007/s12559-026-10621-7

Keywords: Multilingual NLP, Emotion classification, Transformer models, Cultural bias, Semantic robustness, Code-switching, Idiomatic expressions, Fine-tuning, XLM-R, mBERT, LIME, SHAP

Cite Scienmag News

Denise Maddox. (September 9, 2026). Testing Transformer Models’ Emotion Recognition Across Languages and Cultures. Scienmag. https://scienmag.com/testing-transformer-models-emotion-recognition-across-languages-and-cultures/

Denise Maddox. "Testing Transformer Models’ Emotion Recognition Across Languages and Cultures." Scienmag, 9 September 2026, https://scienmag.com/testing-transformer-models-emotion-recognition-across-languages-and-cultures/. Accessed 9 September 2026.

Denise Maddox. "Testing Transformer Models’ Emotion Recognition Across Languages and Cultures." Scienmag. September 9, 2026. https://scienmag.com/testing-transformer-models-emotion-recognition-across-languages-and-cultures/

Tags: challenges of emotion detection across languageschallenges of sarcasm detection in multilingual NLPcross-cultural differences in emotion expression and AI interpretationcross-cultural sentiment analysiscultural context in AI emotion analysiscultural context in AI emotion understandingevaluation of emotion recognition accuracy across languagesfine-tuning transformer models for cultural diversityimpact of cultural nuances on AI sentiment analysisimpact of internet slang on AI emotion interpretationlimitations of current emotion recognition datasetslimitations of standard datasets in emotion AIlinguistic and cultural barriers in emotionmultilingual BERT and XLM-R performance in emotion tasksMultilingual emotion recognition in transformer modelsmultilingual sentiment analysis in social media monitoringsarcasm detection in multilingual modelsslang and idiom processing in natural language processingslang and internet speech understanding by language modelstransformer-based AI for social media sentiment analysistransformer-based models’ performance on informal languageunderstanding human emotions in diverse linguistic contexts
Share26Tweet16
Previous Post

Latent representations and SNOMED-CT mapping improve diagnosis classification in EMR data

Next Post

Outcome reward models improve LLM-based Text-to-SQL generation with GradeSQL

Related Posts

Outcome reward models improve LLM-based Text-to-SQL generation with GradeSQL
Technology and Engineering

Outcome reward models improve LLM-based Text-to-SQL generation with GradeSQL

September 9, 2026
Latent representations and SNOMED-CT mapping improve diagnosis classification in EMR data
Technology and Engineering

Latent representations and SNOMED-CT mapping improve diagnosis classification in EMR data

September 9, 2026
Drug-specific atrial fibrillation risk seen in coronary artery disease patients
Technology and Engineering

Drug-specific atrial fibrillation risk seen in coronary artery disease patients

September 9, 2026
Squeezed quadratures observed in degenerate optical parametric oscillator above threshold
Technology and Engineering

Squeezed quadratures observed in degenerate optical parametric oscillator above threshold

September 8, 2026
M2IND spotlights manufacturing hurdles facing modern industry
Technology and Engineering

M2IND spotlights manufacturing hurdles facing modern industry

September 8, 2026
Machine Learning Framework Assesses Roadway Vulnerability Using Aerial Imagery
Technology and Engineering

Machine Learning Framework Assesses Roadway Vulnerability Using Aerial Imagery

September 8, 2026
Next Post
Outcome reward models improve LLM-based Text-to-SQL generation with GradeSQL

Outcome reward models improve LLM-based Text-to-SQL generation with GradeSQL

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Outcome reward models improve LLM-based Text-to-SQL generation with GradeSQL
  • Testing Transformer Models’ Emotion Recognition Across Languages and Cultures
  • Latent representations and SNOMED-CT mapping improve diagnosis classification in EMR data
  • New Framework Measures Digital Maturity of Chinese Hospitals

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading