Sunday, October 11, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy

October 11, 2026
in Technology and Engineering
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 4 mins read
0
AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy

AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Grading spoken English has long been one of the most stubborn bottlenecks in language education. Human raters are expensive, slow, and notoriously inconsistent, with scores that can vary depending on the assessor’s mood, background, or expectations. Now a study published in Discover Artificial Intelligence describes a multimodal deep learning framework that automates oral English fluency assessment by listening to speech and reading its transcript at the same time, achieving 97.53% accuracy in classifying learners’ proficiency levels.

The framework, developed by Ping Zhang of the Shanghai University of Political Science and Law, rests on a simple but powerful insight: fluency is not just a property of sound or of words, but of both together. Most existing automated systems analyze either the acoustic signal or the transcribed text in isolation, missing the interplay between how something is said and what is actually said. By fusing both streams of information, the new model captures pronunciation, rhythm, prosody, semantic coherence, and even emotional expression within a single end-to-end pipeline.

Technically, the system processes each speech recording in two parallel branches. The audio is normalized, filtered for background noise, and segmented using voice activity detection before being converted into a Log-Mel spectrogram, a time-frequency representation that maps frequencies onto a scale approximating human auditory perception. This representation preserves rich spectral detail, including pitch contours and pause durations, that coarser features like MFCCs tend to discard. Meanwhile, the corresponding transcript is tokenized and passed through BERT, the transformer-based language model, whose bidirectional self-attention produces contextual embeddings that encode the meaning of each word in relation to the whole sentence.

The heart of the framework is its fusion strategy. Rather than simply concatenating acoustic and semantic features and hoping for the best, the model applies a self-attention mechanism over the combined feature vector. This mechanism computes query, key, and value projections, scores the relevance of every feature to every other feature, and produces a weighted representation that emphasizes whichever pieces of information matter most for judging fluency. The authors chose self-attention over cross-attention because it captures dependencies between the already-merged modalities without adding computational complexity.

The fused representation is then fed into a multilayer perceptron with a softmax output layer that classifies each response into low, medium, or high fluency. Training relied on the Adam optimizer, ReLU activations in the hidden layers, and dropout regularization to prevent overfitting. Notably, the pipeline was built across two deep learning ecosystems: TensorFlow handled model training and optimization, while PyTorch implemented the transformer-based BERT embeddings, exploiting the strengths of both frameworks in one system.

The experiments drew on a Kaggle dataset of 8,520 English speech samples from 426 speakers, roughly balanced between male and female, and spanning American, British, Indian, Chinese, and other accents. The recordings total 28.1 hours, with utterances averaging 11.8 seconds, and are split into 5,964 training, 1,278 validation, and 1,278 testing samples across the three fluency classes. To harden the model against real-world messiness, the researchers applied audio augmentation techniques such as time stretching and white noise injection, and used ANOVA-based feature selection to retain only the most discriminative acoustic and semantic features.

The results are striking. The framework achieved 97.53% accuracy, 97.68% precision, 97.53% recall, and a 97.52% F1-score, outperforming single-modal baselines and well-known architectures including CNNs, LSTMs, BiLSTMs, SpeechTransformer, wav2vec 2.0, HuBERT, and Whisper. An ablation study underscores the value of the multimodal design: speech-only and text-only versions reached just 90.52% and 91.81% accuracy respectively, and removing the self-attention module, BERT, or the Log-Mel features each dragged performance down by several percentage points. The model also correlated strongly with human expert ratings, with Pearson and Spearman correlations of 0.941 and 0.932, and quadratic weighted kappa of 0.921.

The system is not lightweight. It carries 112.4 million parameters, occupies 427 MB, and demands 5.6 GB of GPU memory, though inference is fast at 21 milliseconds per sample. The authors are candid about the limitations: performance can degrade with poor-quality audio, heavy background noise, or erroneous speech recognition, and the framework was validated on a single dataset. Fairness across accents remains an open question, and no formal usability study with students and educators has yet been conducted, though the authors flag user-centered testing as a priority for future work.

Even so, the implications for education are considerable. English is the world’s lingua franca, and demand for scalable, objective speaking assessment far exceeds the supply of trained human raters. A framework that scores fluency, pronunciation, rhythm, and expressiveness consistently and instantly could be embedded in online courses, virtual classrooms, and language-learning apps, giving learners detailed feedback that no exam hall could provide at scale. The authors envision future deployment on cloud and edge platforms, validation on larger multilingual datasets, and adaptation to languages beyond English. If those steps succeed, the era of waiting weeks for a speaking test score may finally be drawing to a close.

Subject of Research: Automated oral English fluency assessment using multimodal deep learning with acoustic and semantic feature fusion

Article Title: A multimodal deep learning framework for automated oral English fluency assessment

Article References: Zhang, P. (2026). A multimodal deep learning framework for automated oral English fluency assessment. Discover Artificial Intelligence, 6(1), Article 1387. https://doi.org/10.1007/s44163-026-02325-6

Image Credits: AI Generated

DOI: 10.1007/s44163-026-02325-6

Keywords: multimodal deep learning, oral English fluency, automated assessment, Log-Mel spectrograms, BERT embeddings, self-attention, speech processing, natural language processing, educational technology, speech recognition, language testing, artificial intelligence

Cite Scienmag News

Blake Davidson. (October 11, 2026). AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy. Scienmag. https://scienmag.com/ai-that-listens-and-reads-multimodal-deep-learning-scores-spoken-english-with-97-5-accuracy/

Blake Davidson. "AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy." Scienmag, 11 October 2026, https://scienmag.com/ai-that-listens-and-reads-multimodal-deep-learning-scores-spoken-english-with-97-5-accuracy/. Accessed 11 October 2026.

Blake Davidson. "AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy." Scienmag. October 11, 2026. https://scienmag.com/ai-that-listens-and-reads-multimodal-deep-learning-scores-spoken-english-with-97-5-accuracy/

Tags: accuracy of AI in oral language proficiency testingadvancements in language proficiency grading toolsAI-based pronunciation and fluency scoringArtificial Intelligenceautomated assessmentautomated spoken English assessmentBERT embeddingschallenges in automated language assessmenteducational technologyend-to-end speech evaluation modelsintegration of acoustic and textual data in AIlanguage testingLog-Mel spectrogramsmultimodal deep learningmultimodal deep learning in language educationmultimodal speech analysis for language learningnatural language processingoral English fluencyprosody and emotional expression recognition in speechself-attentionspeech and transcript fusion for language proficiencyspeech processingspeech recognitionspeech signal processing with neural networks
Share26Tweet16
Previous Post

Body Weight, Anatomy and Access Route Drive Radiation Doses in Liver Drainage Procedures

Next Post

Bacterial Biofilms May Trigger the Autoantibodies Behind Lupus Kidney Damage

Related Posts

Overlapping Diesel and Gas Injections Push Natural Gas Truck Engines to Record Efficiency
Technology and Engineering

Overlapping Diesel and Gas Injections Push Natural Gas Truck Engines to Record Efficiency

October 11, 2026
Liquid Metal Lubricant Tames Molybdenum Friction From Room Temperature to 300 Degrees
Technology and Engineering

Liquid Metal Lubricant Tames Molybdenum Friction From Room Temperature to 300 Degrees

October 11, 2026
AI Learns to Say It Doesn’t Know: Neutrosophic Geometry Tackles Overconfident Deep Networks
Technology and Engineering

AI Learns to Say It Doesn’t Know: Neutrosophic Geometry Tackles Overconfident Deep Networks

October 11, 2026
Molecular Shape Holds the Key to Sharper Optical Pressure Standards
Technology and Engineering

Molecular Shape Holds the Key to Sharper Optical Pressure Standards

October 11, 2026
Digital Mindfulness and Compassion Therapies Show Modest but Real Mental Health Gains
Medicine

Digital Mindfulness and Compassion Therapies Show Modest but Real Mental Health Gains

October 11, 2026
New Open-Source Tool Automates Simulation of Pipeline Pitting Corrosion Monitoring
Technology and Engineering

New Open-Source Tool Automates Simulation of Pipeline Pitting Corrosion Monitoring

October 11, 2026
Next Post
Bacterial Biofilms May Trigger the Autoantibodies Behind Lupus Kidney Damage

Bacterial Biofilms May Trigger the Autoantibodies Behind Lupus Kidney Damage

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Bacterial Biofilms May Trigger the Autoantibodies Behind Lupus Kidney Damage
  • AI That Listens and Reads: Multimodal Deep Learning Scores Spoken English With 97.5% Accuracy
  • Body Weight, Anatomy and Access Route Drive Radiation Doses in Liver Drainage Procedures
  • AI Can Speed Vaccine Design, But Chemistry Still Decides Which Epitopes Survive

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading