Friday, September 25, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Earth Science

New AI Model Catches Deepfakes by Listening and Watching Simultaneously

September 25, 2026
in Earth Science
Violet Maxwell
By Violet Maxwell Scienmag Editorial Profile - Natural Hazards
Reading Time: 5 mins read
0
New AI Model Catches Deepfakes by Listening and Watching Simultaneously

New AI Model Catches Deepfakes by Listening and Watching Simultaneously

New AI Model Catches Deepfakes by Listening and Watching Simultaneously

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Deepfakes have long been treated as a visual problem: a swapped face, a warped expression, a flicker of unnatural texture around the mouth. But the most dangerous forgeries circulating today are multimodal, blending synthetic video with cloned voices, generated speech, and mismatched audio tracks that would fool a casual viewer completely. A team of researchers led by Leyan Wang, Jian Zhao, and Zhaofeng He, working at the Institute of Artificial Intelligence (TeleAI) at China Telecom along with collaborators at Beijing University of Posts and Telecommunications and several Chinese universities, has now unveiled a detection system built specifically for this messier reality. Their model, called ERF-BA-TFD+, was described in the open-access journal Vicinagearth in November 2025, and it recently took first place in the Audio-Visual Detection and Localization track of the Workshop on Deepfake Detection, Localization, and Interpretability, a competition centered on the demanding DDL-AV dataset.

The core insight behind the new system is that real audio and real video are locked together in ways that are extremely difficult to forge consistently. When a person speaks, lip movements, facial expressions, prosody, and timing all co-vary. Generators can produce convincing audio and convincing video separately, but keeping the two streams synchronized and semantically consistent across an entire clip is far harder. ERF-BA-TFD+ exploits exactly this gap. Rather than analyzing each modality in isolation, as most earlier detectors did, it processes audio and video features simultaneously and hunts for the subtle discrepancies that emerge when one stream has been manipulated and the other has not, or when both have been forged but their temporal alignment betrays the fabrication.

Technically, the framework rests on two powerful feature extractors. For the visual stream, the researchers chose MViTv2, a hierarchical multiscale vision transformer. Unlike standard vision transformers that process every token at a fixed scale, MViTv2 progressively pools features, letting it capture anomalies at multiple granularities, from pixel-level texture irregularities to coarse, motion-based artifacts such as inconsistent facial kinematics. For audio, the team employed BYOL-A, a self-supervised model pre-trained on a vast and diverse corpus of sound. The choice of a self-supervised encoder is strategic: supervised detectors trained to spot specific artifact types tend to overfit and fail when confronted with novel forgery methods, whereas BYOL-A learns a rich general-purpose representation that can flag unnatural prosody, lip-sync mismatches, and distortions that are often imperceptible to human ears.

The heart of the temporal forgery detection pipeline is a module called CRATrans, the Cross-Reconstruction Attention Transformer. Instead of simply concatenating audio and video features, which can dilute the distinctive signals in each stream, CRATrans adopts a cross-reconstruction paradigm. During training, the model is forced to reconstruct the feature sequence of one modality using features from the other as context: visual features are rebuilt from audio cues and vice versa. This adversarial exercise teaches the network fine-grained inter-modal temporal dependencies. In genuine footage, where the streams are aligned and semantically consistent, reconstruction error stays low. In manipulated content, the attempt to reconstruct one stream from the other produces a markedly higher error, a direct and robust indicator of forgery. Multi-head self-attention models long-range dependencies within each modality, while cross-attention handles information exchange between them, and the resulting attention weights and reconstruction errors help pinpoint anomalous regions at inference time.

Localization proceeds in a coarse-to-fine hierarchy. Independent frame-level classification modules screen each modality separately, producing per-frame anomaly scores that capture modality-specific inconsistencies before any fusion can obscure them. A dedicated Boundary Localization Module, inspired by the proposal-relation mechanism of BSN++, then converts anomalous frames into precise temporal boundaries. It builds a confidence matrix over all possible start-end frame pairs, sharpened by two complementary attention mechanisms: position-aware attention for global temporal context and channel-aware attention for relationships among feature channels. Two separate boundary modules, one per modality, ensure that precision in localizing a forged audio segment is not compromised by imprecision in video, and the final outputs are refined through soft non-maximum suppression, with parameters empirically tuned to merge overlapping proposals and suppress low-confidence detections.

Perhaps the most inventive component is the Evidence-based Reasoning Framework, or ERF, which addresses a blind spot that plagues nearly all temporal detectors: the full-fake video. A model with a finite temporal receptive field learns to spot forgeries by contrasting suspicious segments with authentic ones. But when an entire video is fabricated, there is no clean baseline to contrast against, and such clips can slip through undetected. ERF reasons statistically over the whole video’s detection output. Localized forgeries produce some segments with high confidence; fully real videos yield uniformly low scores. A full-fake video, by contrast, often produces moderately confident scores everywhere with no single peak. If the maximum confidence across predicted segments fails to exceed a threshold, ERF re-evaluates the video at a global level and can reclassify it as a potential full-fake. The module is lightweight, essentially a corrective rule layer, yet it closed a genuine gap in the model’s logic.

The experimental path to the final system was itself revealing, unfolding in three phases on the DDL-AV and LAV-DF datasets, both of which contain segmented clips and full-length videos and confront models with text-to-speech, voice cloning, voice swapping, face swapping, facial animation, and text-to-video generation. In the first phase, the baseline model scored impressively on LAV-DF, achieving an average precision of 0.9630 at an intersection-over-union threshold of 0.5, but training on DDL-AV actually degraded those lenient-threshold scores, a counter-intuitive result the authors attribute to hyperspecialization on subtle artifacts. Tellingly, the strictest metric, AP at 0.95, improved slightly, suggesting the fine-tuned model had grown more sensitive to forgeries requiring very precise localization even as it lost confidence on easy cases.

The second phase exposed and then fixed a deeper weakness. A bad-case analysis showed the fusion model was almost blind to audio-only forgeries, posting a near-zero AP of 0.0163 on that subset, apparently because the strong visual stream drowned out faint auditory manipulation signals. Integrating the Unified Multimodal Attention framework, which uses cross-attention to force a more judicious weighting of both streams, produced a dramatic rebound: AP at 0.5 on the audio-forgery subset jumped to 0.9243. The final phase added the ERF module and lifted the overall competition score to 0.78, with average recall in the top-100 detections on a long-video validation set rising from 0.6513 to 0.7886. The authors report that ERF-BA-TFD+ achieved state-of-the-art results on DDL-AV and outperformed most competing models on LAV-DF while also offering superior processing speed.

For a field racing to keep pace with generative models built on GANs, diffusion architectures, and variational autoencoders, the study offers both a practical tool and a design philosophy. Its lessons, that balanced cross-modal attention is essential for resisting single-modality attacks, and that global statistical reasoning can rescue detectors from the full-fake blind spot, form what the authors describe as a blueprint for future systems. Challenges remain, particularly severe or non-linear temporal desynchronization between audio and video and the approach of next-generation forgeries with fewer low-level artifacts. The team points toward zero-shot and few-shot learning as a route to rapid adaptation against unseen manipulation techniques, and both the LAV-DF dataset, available on Hugging Face, and the forthcoming full release of DDL-AV should give the wider community the means to build on this work as the arms race between fabrication and detection continues.

Subject of Research: Multimodal audio-visual deepfake detection using cross-modal reconstruction attention and evidence-based reasoning

Article Title: ERF-BA-TFD+: a multimodal model for audio-visual deepfake detection

Article References: Wang, L., Zhao, J., Zhang, X., Guo, X., Yuan, Y., Zhang, T., Chu, J., Jiang, Y., Yang, X., Jin, L., Zhang, C., & He, Z. (2025). ERF-BA-TFD+: a multimodal model for audio-visual deepfake detection. Vicinagearth, 2(1), Article 10. https://doi.org/10.1007/s44336-025-00021-0

Image Credits: AI Generated

DOI: 10.1007/s44336-025-00021-0

Keywords: deepfake detection, multimodal learning, audio-visual fusion, transformers, MViTv2, BYOL-A, temporal localization, DDL-AV dataset, LAV-DF dataset, digital forensics, generative adversarial networks, cross-attention

Cite Scienmag News

Violet Maxwell. (September 25, 2026). New AI Model Catches Deepfakes by Listening and Watching Simultaneously. Scienmag. https://scienmag.com/new-ai-model-catches-deepfakes-by-listening-and-watching-simultaneously/

Violet Maxwell. "New AI Model Catches Deepfakes by Listening and Watching Simultaneously." Scienmag, 25 September 2026, https://scienmag.com/new-ai-model-catches-deepfakes-by-listening-and-watching-simultaneously/. Accessed 25 September 2026.

Violet Maxwell. "New AI Model Catches Deepfakes by Listening and Watching Simultaneously." Scienmag. September 25, 2026. https://scienmag.com/new-ai-model-catches-deepfakes-by-listening-and-watching-simultaneously/

Tags: advancements in deepfake detection competitionsAI-based fake media identificationaudio-visual fusionaudio-visual synchronizationaudiovisual forensicsBYOL-Achallenges in detecting synthetic videos with cloned voicescross-attentioncross-modal forgery detection techniquesDDL-AV datasetdeepfake detectiondeepfake localization and interpretabilitydigital forensicsERF-BA-TFD+ modelgenerative adversarial networksLAV-DF datasetmultimodal AI modelsmultimodal deepfake detection systemsmultimodal learningMViTv2synthetic media forgeriestemporal localizationtransformers
Share26Tweet16
Previous Post

Automation Marches Into Data Warehouse Design, But Big Gaps Remain

Next Post

Hepatitis B Deaths Keep Rising in the Western Pacific Even as New Infections Fall, Decade-Long Analysis Finds

Related Posts

Machine Learning and Bayesian Statistics Rebuild Earthquake Fragility Curves from Italy’s 1980 Irpinia Disaster
Earth Science

Machine Learning and Bayesian Statistics Rebuild Earthquake Fragility Curves from Italy’s 1980 Irpinia Disaster

September 25, 2026
Fuzzy Logic and Image Analysis Team Up to Grade Sandstone Reservoir Quality
Earth Science

Fuzzy Logic and Image Analysis Team Up to Grade Sandstone Reservoir Quality

September 25, 2026
Sinking Rift: Satellites Reveal Kenya’s Nakuru County Is Slowly Collapsing
Earth Science

Sinking Rift: Satellites Reveal Kenya’s Nakuru County Is Slowly Collapsing

September 25, 2026
Ranking the World’s Climate Models for South America: Four Models Rise Above the Rest
Earth Science

Ranking the World’s Climate Models for South America: Four Models Rise Above the Rest

September 25, 2026
Spain Puts Shipwrecks on the Map, but Its Ocean Plans Leave Heritage on the Sidelines
Earth Science

Spain Puts Shipwrecks on the Map, but Its Ocean Plans Leave Heritage on the Sidelines

September 25, 2026
Ancient Fossils Reveal Microbe Partnerships Predate the Cambrian Explosion
Earth Science

Ancient Fossils Reveal Microbe Partnerships Predate the Cambrian Explosion

September 25, 2026
Next Post
Hepatitis B Deaths Keep Rising in the Western Pacific Even as New Infections Fall, Decade-Long Analysis Finds

Hepatitis B Deaths Keep Rising in the Western Pacific Even as New Infections Fall, Decade-Long Analysis Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Gut Microbes May Predict Survival in Dogs Receiving Cancer Immunotherapy
  • Radiation Oncology Advances Take Center Stage as MD Anderson Presents Nearly 70 Abstracts at ASTRO 2026
  • Brain Bleeds on Blood Thinners: New Guideline Reshapes Emergency Reversal Strategy
  • Hepatitis B Deaths Keep Rising in the Western Pacific Even as New Infections Fall, Decade-Long Analysis Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading