Sunday, October 4, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI Turns Marathi Text Into Realistic Indian Sign Language Videos

October 4, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
AI Turns Marathi Text Into Realistic Indian Sign Language Videos

AI Turns Marathi Text Into Realistic Indian Sign Language Videos

AI Turns Marathi Text Into Realistic Indian Sign Language Videos

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

For more than 70 million deaf and hard-of-hearing people in India, access to everyday information often depends on whether someone nearby can sign. Now a team of researchers in Pune has built an artificial intelligence pipeline that takes written Marathi—a language with almost no sign language technology supporting it—and converts it into fluent, visually convincing videos of Indian Sign Language (ISL). The work, published in Multimedia Tools and Applications, combines a custom translation network with a generative adversarial framework that animates a synthetic signer, and it runs fast enough to hint at near real-time use.

The challenge the researchers set out to solve is one of resources. Sign language translation systems for English, Chinese, and several European languages benefit from large parallel corpora—thousands of hours of signed video aligned with spoken-language text. Marathi, spoken by more than 80 million people, has no comparable dataset for ISL. Indian Sign Language itself is distinct from American or British sign languages, with its own grammar, its own spatial referencing system, and heavy reliance on facial expressions and handshapes that carry grammatical meaning. Building a text-to-video system for this low-resource pairing means solving two hard problems at once: translating Marathi sentences into a signed representation, and then rendering that representation as a moving human figure that looks natural rather than robotic.

The team, led by Prachi Pramod Waghmare with Ashwini Mangesh Deshpande and Supriya A. Mangale of Cummins College of Engineering for Women, affiliated with Savitribai Phule Pune University, addressed the first problem with what they call a Sparse Multi-Hierarchical Transformer. Transformers, the architecture behind modern large language models, normally attend to every token in a sequence, which becomes computationally expensive as sentences grow. The sparse, hierarchical variant prunes that attention: it organizes the Marathi input and the target gloss—a gloss being a written label for a sign, essentially the bridge vocabulary between a spoken and a signed language—into layers, attending selectively across words, phrases, and sentence-level structure. The result, according to the paper, is an efficient mapping from Marathi text to ISL gloss that preserves the reordering needed when a subject-verb-object spoken sentence becomes the spatially organized structure of a signed one.

The numbers reported for this translation stage are striking. The framework achieves BLEU-1, BLEU-2, and BLEU-3 scores of 99.93, with a BLEU-4 score of 56.23. BLEU scores measure how closely machine-generated output matches human reference translations, with the higher-order n-gram scores being progressively harder to satisfy because they demand matching longer consecutive sequences. Near-perfect unigram, bigram, and trigram scores indicate that virtually every sign and sign pair in the generated gloss sequences appears in the right place; the drop at BLEU-4 reflects the familiar difficulty of reproducing longer fluent stretches exactly, which is where sign order variability and synonymy among glosses bite hardest.

Once the gloss sequence exists, the system must turn it into video, and this is where the generative adversarial network comes in. GANs pit two neural networks against each other: a generator that produces candidate outputs and a discriminator that tries to tell them apart from real examples. Over training, the generator improves until its outputs fool the discriminator. The researchers’ variant, which they describe as a Pose-Enhanced Video-Transformer GAN, conditions the generation on explicit skeletal information rather than asking the network to invent human anatomy from scratch. OpenPose, a widely used landmark detection library, supplies the body, hand, and face keypoints that define a signer’s posture at every moment. These landmarks act as a structural scaffold: the generator learns to paint realistic video content onto a skeleton that already encodes where the arms, hands, and facial features must be for each sign.

The rendering stage, called a Pose-Adapted Progressive Video Synthesizer, builds the video in stages rather than all at once. Progressive synthesis—refining output from coarse to fine—has proven effective in image generation, and here it is adapted to the temporal dimension, so that the motion of the signer emerges coherently across frames instead of flickering between inconsistent poses. Temporal coherence is the notorious weakness of frame-by-frame generation: without it, fingers jitter, backgrounds shimmer, and the signer appears to morph. The pose conditioning plus progressive refinement is designed to suppress exactly those artifacts, keeping the signer’s body stable while the hands and face execute the signs.

Quality metrics for the generated videos back up the design. The framework reports a structural similarity index (SSIM) of 0.912, a measure of how closely the generated frames resemble real signer footage in structure and luminance, on a scale where 1.0 is identical. Peak signal-to-noise ratio (PSNR) comes in at 31.7 decibels, a reconstruction fidelity measure in which values above 30 dB are generally considered good for video synthesis. Perhaps most important for practical deployment, the system generates video at 15 frames per second with an average inference time of 120 milliseconds per sample. That is not broadcast speed, but it approaches the threshold where a translation service could produce signed video on demand rather than after a long offline render, which is what most earlier sign language production systems required.

The significance of the work extends beyond its benchmark numbers. Most sign language production research targets languages with abundant data, and several recent efforts—gloss-free translation models, sign language large language models, and diffusion-based generators—assume resources that simply do not exist for Marathi-ISL. By demonstrating a pipeline that works with a limited dataset and a single signer, the Pune team shows a viable path for other under-resourced language-sign pairs, of which there are hundreds worldwide. The hierarchical sparse attention mechanism also speaks to efficiency: sign language generation is computationally heavy because it couples language modeling with video synthesis, and trimming the attention cost at the translation stage buys headroom for the rendering stage.

The authors are candid about the limitations. Their Marathi-ISL dataset is small and features only one signer, which means the generated videos inherit a single person’s signing style, body proportions, and appearance. Generalizing to multiple signers, regional variation in ISL, and spontaneous rather than scripted sentences remains future work. The researchers frame the current system as a foundation for an adaptable sign language generation platform rather than a finished product. Ethical safeguards were observed: informed consent was obtained from the human signer whose recordings underpin the dataset, and the authors report no competing interests.

Still, the trajectory is clear. If systems like this one can be scaled to larger, multi-signer datasets and integrated with speech recognition, a deaf person could eventually receive signed video versions of news articles, government notices, classroom material, or medical instructions in their own sign language, generated on the fly. For Marathi speakers in the deaf community, who have until now been largely invisible to sign language technology, this research is an early but concrete step toward that future—and a template for bringing the benefits of generative AI to the languages and communities that need it most.

Subject of Research: Machine translation of Marathi text to Indian Sign Language gloss and pose-guided GAN-based sign language video generation

Article Title: Efficient marathi text-to-gloss translation and realistic isl video generation using pose-enhanced GAN

Article References: Waghmare, P. P., Deshpande, A. M., & Mangale, S. A. (2026). Efficient marathi text-to-gloss translation and realistic isl video generation using pose-enhanced GAN. Multimedia Tools and Applications, 85(9), Article 747. https://doi.org/10.1007/s11042-026-21908-0

Image Credits: AI Generated

DOI: 10.1007/s11042-026-21908-0

Keywords: Indian Sign Language, Marathi, text-to-gloss translation, generative adversarial network, pose estimation, OpenPose, transformer, video synthesis, sign language production, low-resource languages, accessibility, deep learning

Cite Scienmag News

Denise Maddox. (October 4, 2026). AI Turns Marathi Text Into Realistic Indian Sign Language Videos. Scienmag. https://scienmag.com/ai-turns-marathi-text-into-realistic-indian-sign-language-videos/

Denise Maddox. "AI Turns Marathi Text Into Realistic Indian Sign Language Videos." Scienmag, 4 October 2026, https://scienmag.com/ai-turns-marathi-text-into-realistic-indian-sign-language-videos/. Accessed 4 October 2026.

Denise Maddox. "AI Turns Marathi Text Into Realistic Indian Sign Language Videos." Scienmag. October 4, 2026. https://scienmag.com/ai-turns-marathi-text-into-realistic-indian-sign-language-videos/

Tags: accessibilityaccessibility tools for deaf and hard-of-hearing in IndiaAI-based sign language generation for low-resource languagescustom translation networks for sign languagedeep learningfacial expressions and handshapes in sign language AIgenerative adversarial networkgenerative adversarial networks for sign language animationIndian Sign LanguageIndian Sign Language (ISL) video synthesislow-resource languageslow-resource sign language datasets and AI solutionsMarathiMarathi text to Indian Sign Language video translationOpenPosepose estimationreal-time sign language video creationsign language productionsign language translation systems for regional languagessign language translation technology for Marathitext-to-gloss translationTransformervideo synthesis
Share26Tweet16
Previous Post

New Duplex Real-Time PCR Test Tracks Two Fox Babesia Parasites and Their Tick Vectors

Next Post

Self-Efficacy Scale Measures Men and Women Equally, Ecuadorian Study Finds

Related Posts

Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy
Technology and Engineering

Hybrid CNN-ViT Model With Triple Loss Boosts Image Search Accuracy

October 4, 2026
New AI Transformer Learns to See Interior Design Styles the Way Humans Do
Technology and Engineering

New AI Transformer Learns to See Interior Design Styles the Way Humans Do

October 4, 2026
New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data
Technology and Engineering

New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data

October 4, 2026
BatPose Turns Two Cameras and a Laptop Into a Markerless 3D Motion Capture Lab
Technology and Engineering

BatPose Turns Two Cameras and a Laptop Into a Markerless 3D Motion Capture Lab

October 4, 2026
Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men
Technology and Engineering

Two Weeks of Overeating Weakens the Gut Barrier and Ignites Liver Immunity in Healthy Men

October 4, 2026
New Model Captures Choked Gas Blasts and Wall Heat in Pressurized Vessel Discharge
Technology and Engineering

New Model Captures Choked Gas Blasts and Wall Heat in Pressurized Vessel Discharge

October 4, 2026
Next Post
Self-Efficacy Scale Measures Men and Women Equally, Ecuadorian Study Finds

Self-Efficacy Scale Measures Men and Women Equally, Ecuadorian Study Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Self-Efficacy Scale Measures Men and Women Equally, Ecuadorian Study Finds
  • AI Turns Marathi Text Into Realistic Indian Sign Language Videos
  • New Duplex Real-Time PCR Test Tracks Two Fox Babesia Parasites and Their Tick Vectors
  • Knowing the Risks Isn’t Enough: Ghana Study Reveals Obesity Knowledge Gap in Urban Adults

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,149 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading