Sunday, September 27, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language

September 27, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 4 mins read
0
AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language

AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language

AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

For more than 70 million Deaf and Hard-of-Hearing people worldwide, everyday communication still depends on human interpreters or imperfect technology. A new survey published in the International Journal of Machine Learning and Cybernetics argues that a quiet revolution is underway: the same large language models and vision-language models that power chatbots and image captioning are now being harnessed to translate sign language video directly into fluent spoken-language text, without the crude intermediaries that have constrained the field for decades. The review, led by Sufyan Danish and colleagues at Sejong University, Khalifa University and Princess Nourah bint Abdulrahman University, maps this emerging landscape and offers a roadmap for what comes next.

At the heart of the transformation is a shift away from so-called gloss-based translation. Historically, most sign language translation systems did not translate signing directly into sentences. Instead, they first converted signs into glosses, textual labels that roughly transcribe each sign into a word-like form. Glosses, however, are a lossy shorthand. Sign languages are fully-fledged languages with their own grammar, expressed through handshape, motion, facial expression and body posture, and the gloss layer discards much of this rich structure. It also requires expensive annotation by experts, which has kept progress confined to a handful of well-resourced languages. Gloss-free approaches, by contrast, learn to map continuous sign video straight to natural language, cutting out the bottleneck entirely.

The survey organizes the current wave of gloss-free systems into three architectural paradigms. The first involves adapter-based language models, in which compact neural modules are inserted into frozen large language models so that visual features extracted from sign video can be injected into a powerful text generator. Techniques such as low-rank adaptation, popularized by the LoRA method, allow researchers to tune these systems on relatively small sign language datasets without retraining billions of parameters. Systems like Sign2GPT and related work demonstrate that the linguistic knowledge already baked into large language models can be repurposed to produce grammatically coherent translations from signing alone.

The second paradigm centers on hierarchical tokenization frameworks. Sign video is dense and continuous, and naive approaches tend to compress it into representations that are either too coarse to capture nuance or too diffuse to decode. Newer methods reduce what researchers call representation density, structuring visual features at multiple temporal scales so that fast, fine-grained movements and slower, sentence-level rhythms are both preserved. Combined with sign back-translation, a data augmentation trick that generates synthetic training pairs, these frameworks have pushed benchmark scores upward on standard evaluation sets.

The third and arguably most exciting paradigm is vision-language pretraining. Models first trained on enormous collections of paired images and text, such as CLIP-style encoders and DINOv2 visual backbones, already possess a shared visual-linguistic space. Gloss-free translation systems exploit this by treating sign video as just another visual input to a pretrained vision-language model, then fine-tuning on sign data. The survey highlights systems including SignLLM, DiffSLT, and LLaVA-SLT, which extend this idea with diffusion-based generation and visual instruction tuning. The result is a family of models that translate signing with increasing fluency while also opening the door to sign language generation, producing signing avatars from text.

Underpinning all of this is the question of data. The survey reviews the modern dataset landscape, from How2Sign and the BBC-Oxford BSL corpus to Arabic sign language resources such as ArabSign and KaSL, and multilingual efforts like AfriSign. The imbalance is stark: American and British Sign Language dominate, while most of the world’s hundreds of sign languages remain severely under-resourced. This is where the motivation of the survey’s authors becomes concrete. During Hajj and Umrah, millions of pilgrims, including Deaf and Hard-of-Hearing individuals, gather annually in Saudi Arabia, and accessible translation in Arabic Sign Language is a matter of safety and dignity. Saudi Arabia has reported some 20 million deaf visitors expected to benefit from sign language services at the two holy mosques, a scale that no human interpreter workforce can match.

Evaluation remains a thorny problem. The survey examines the standard metrics, including BLEU, ROUGE, CIDEr and translation edit rate, and notes ongoing debates about standardized BLEU computation and the clarity of reported scores across papers. Automatic metrics, originally designed for text-to-text machine translation, struggle to judge whether a signed utterance has been rendered into semantically faithful English, particularly when valid paraphrases exist. Emerging evaluation approaches and human judgment studies are therefore essential complements, and the authors call for more consistent reporting so that comparisons between gloss-free frameworks are meaningful.

The survey also takes stock of practical knowledge-based systems already deployed to assist Deaf users, including real-time multilingual translation applications built on frameworks like MediaPipe for hand landmark tracking. These tools show promise but also expose limitations: signer variability, differing signing rates that alter the timing of manual signs and non-manual markers, and the challenge of mouthings, which research shows do not always align neatly with hand movements in languages like British Sign Language. Skeleton-based approaches and motion-visual fusion architectures are among the strategies being explored to make recognition robust across diverse signers and recording conditions.

Looking forward, the authors identify low-resource adaptation, ethical AI development and global accessibility as the defining frontiers. Low-rank adaptation and prefix-tuning techniques make it feasible to bootstrap systems for sign languages with only small corpora, while multilingual pretraining methods like mT5 hint at translation models that span many languages at once. Ethics is not an afterthought: Deaf communities have historically been skeptical of technologies developed without their involvement, and the survey’s framing of translation as a bridge for equitable access implicitly demands that future systems be co-designed with Deaf users rather than imposed on them.

The significance of this survey lies less in any single breakthrough than in its synthesis of a fast-moving field at an inflection point. Gloss-free sign language translation, once a niche goal, is now the mainstream research program, powered by the same foundation-model infrastructure reshaping the rest of artificial intelligence. If the roadmap holds, the coming years could see Deaf and hearing people converse directly through cameras and phones, with neural models handling the linguistic heavy lifting, a development that would rank among the most socially consequential applications of the large language model era.

Subject of Research: Gloss-free sign language translation using large language models and vision-language models for Deaf communication

Article Title: Leveraging large and vision language models for gloss-free sign language translation in deaf communication: a survey

Article References: Danish, S., Khan, S. U., Alghamdi, R. A., & Alghamdi, N. S. (2026). Leveraging large and vision language models for gloss-free sign language translation in deaf communication: a survey. International Journal of Machine Learning and Cybernetics, 17(10), Article 479. https://doi.org/10.1007/s13042-026-03323-x

Image Credits: AI Generated

DOI: 10.1007/s13042-026-03323-x

Keywords: sign language translation, gloss-free, large language models, vision-language models, multimodal AI, Deaf communication, neural machine translation, low-resource languages, Arabic sign language, vision-language pretraining, accessibility, transformers

Cite Scienmag News

Denise Maddox. (September 27, 2026). AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language. Scienmag. https://scienmag.com/ai-without-glosses-how-large-language-models-are-learning-to-translate-sign-language/

Denise Maddox. "AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language." Scienmag, 27 September 2026, https://scienmag.com/ai-without-glosses-how-large-language-models-are-learning-to-translate-sign-language/. Accessed 27 September 2026.

Denise Maddox. "AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language." Scienmag. September 27, 2026. https://scienmag.com/ai-without-glosses-how-large-language-models-are-learning-to-translate-sign-language/

Tags: accessibilityadvancements in assistive communication technologyAI-driven sign language interpretationArabic sign languageDeaf communicationgloss-based vs direct translationgloss-freelarge language modelslow-resource languagesmachine learning in deaf communicationmultimodal AIneural machine translationsign language grammar and structuresign language recognition technologysign language translationsign language video captioningsign language video translationspeech-to-text conversion for deaf communitiestransformersvision-language modelsvision-language pretraining
Share26Tweet16
Previous Post

Tiny Genetic Switches May Help Explain Why Embryos Fail to Implant

Next Post

Outsider Eyes: How a Non-Native Ethnographer Uncovered the Social Fault Lines of Water Scarcity in Rural India

Related Posts

AI Joins the Slide and the Chart to Predict Colorectal Cancer Risk
Technology and Engineering

AI Joins the Slide and the Chart to Predict Colorectal Cancer Risk

September 26, 2026
Iron-Carbon Nanocomposite Strips Lead and Cadmium from Wastewater with High Efficiency
Technology and Engineering

Iron-Carbon Nanocomposite Strips Lead and Cadmium from Wastewater with High Efficiency

September 26, 2026
New Open-Access Journal Puts Sustainable Polymers at the Center of Materials Science
Technology and Engineering

New Open-Access Journal Puts Sustainable Polymers at the Center of Materials Science

September 26, 2026
Metallurgy Enters a New Era as AI-Designed Alloys Redefine the Science of Metals
Technology and Engineering

Metallurgy Enters a New Era as AI-Designed Alloys Redefine the Science of Metals

September 26, 2026
AI Learns to Juggle the Internet of Things: Survey Maps Deep Reinforcement Learning’s Rise in Edge and Fog Computing
Technology and Engineering

AI Learns to Juggle the Internet of Things: Survey Maps Deep Reinforcement Learning’s Rise in Edge and Fog Computing

September 26, 2026
When AI Cannot See the Answer: The Mathematical Limit That Optimisation Cannot Cross
Technology and Engineering

When AI Cannot See the Answer: The Mathematical Limit That Optimisation Cannot Cross

September 26, 2026
Next Post
Outsider Eyes: How a Non-Native Ethnographer Uncovered the Social Fault Lines of Water Scarcity in Rural India

Outsider Eyes: How a Non-Native Ethnographer Uncovered the Social Fault Lines of Water Scarcity in Rural India

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Scientists Reverse-Engineer Insect Smell Receptors to Identify Moth Sex Pheromones
  • Machine Learning and Molecular Scaffolds Supercharge Production of a Licorice-Derived Drug Molecule
  • Outsider Eyes: How a Non-Native Ethnographer Uncovered the Social Fault Lines of Water Scarcity in Rural India
  • AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading