For more than 70 million Deaf and Hard-of-Hearing people worldwide, everyday communication still depends on human interpreters or imperfect technology. A new survey published in the International Journal of Machine Learning and Cybernetics argues that a quiet revolution is underway: the same large language models and vision-language models that power chatbots and image captioning are now being harnessed to translate sign language video directly into fluent spoken-language text, without the crude intermediaries that have constrained the field for decades. The review, led by Sufyan Danish and colleagues at Sejong University, Khalifa University and Princess Nourah bint Abdulrahman University, maps this emerging landscape and offers a roadmap for what comes next.
At the heart of the transformation is a shift away from so-called gloss-based translation. Historically, most sign language translation systems did not translate signing directly into sentences. Instead, they first converted signs into glosses, textual labels that roughly transcribe each sign into a word-like form. Glosses, however, are a lossy shorthand. Sign languages are fully-fledged languages with their own grammar, expressed through handshape, motion, facial expression and body posture, and the gloss layer discards much of this rich structure. It also requires expensive annotation by experts, which has kept progress confined to a handful of well-resourced languages. Gloss-free approaches, by contrast, learn to map continuous sign video straight to natural language, cutting out the bottleneck entirely.
The survey organizes the current wave of gloss-free systems into three architectural paradigms. The first involves adapter-based language models, in which compact neural modules are inserted into frozen large language models so that visual features extracted from sign video can be injected into a powerful text generator. Techniques such as low-rank adaptation, popularized by the LoRA method, allow researchers to tune these systems on relatively small sign language datasets without retraining billions of parameters. Systems like Sign2GPT and related work demonstrate that the linguistic knowledge already baked into large language models can be repurposed to produce grammatically coherent translations from signing alone.
The second paradigm centers on hierarchical tokenization frameworks. Sign video is dense and continuous, and naive approaches tend to compress it into representations that are either too coarse to capture nuance or too diffuse to decode. Newer methods reduce what researchers call representation density, structuring visual features at multiple temporal scales so that fast, fine-grained movements and slower, sentence-level rhythms are both preserved. Combined with sign back-translation, a data augmentation trick that generates synthetic training pairs, these frameworks have pushed benchmark scores upward on standard evaluation sets.
The third and arguably most exciting paradigm is vision-language pretraining. Models first trained on enormous collections of paired images and text, such as CLIP-style encoders and DINOv2 visual backbones, already possess a shared visual-linguistic space. Gloss-free translation systems exploit this by treating sign video as just another visual input to a pretrained vision-language model, then fine-tuning on sign data. The survey highlights systems including SignLLM, DiffSLT, and LLaVA-SLT, which extend this idea with diffusion-based generation and visual instruction tuning. The result is a family of models that translate signing with increasing fluency while also opening the door to sign language generation, producing signing avatars from text.
Underpinning all of this is the question of data. The survey reviews the modern dataset landscape, from How2Sign and the BBC-Oxford BSL corpus to Arabic sign language resources such as ArabSign and KaSL, and multilingual efforts like AfriSign. The imbalance is stark: American and British Sign Language dominate, while most of the world’s hundreds of sign languages remain severely under-resourced. This is where the motivation of the survey’s authors becomes concrete. During Hajj and Umrah, millions of pilgrims, including Deaf and Hard-of-Hearing individuals, gather annually in Saudi Arabia, and accessible translation in Arabic Sign Language is a matter of safety and dignity. Saudi Arabia has reported some 20 million deaf visitors expected to benefit from sign language services at the two holy mosques, a scale that no human interpreter workforce can match.
Evaluation remains a thorny problem. The survey examines the standard metrics, including BLEU, ROUGE, CIDEr and translation edit rate, and notes ongoing debates about standardized BLEU computation and the clarity of reported scores across papers. Automatic metrics, originally designed for text-to-text machine translation, struggle to judge whether a signed utterance has been rendered into semantically faithful English, particularly when valid paraphrases exist. Emerging evaluation approaches and human judgment studies are therefore essential complements, and the authors call for more consistent reporting so that comparisons between gloss-free frameworks are meaningful.
The survey also takes stock of practical knowledge-based systems already deployed to assist Deaf users, including real-time multilingual translation applications built on frameworks like MediaPipe for hand landmark tracking. These tools show promise but also expose limitations: signer variability, differing signing rates that alter the timing of manual signs and non-manual markers, and the challenge of mouthings, which research shows do not always align neatly with hand movements in languages like British Sign Language. Skeleton-based approaches and motion-visual fusion architectures are among the strategies being explored to make recognition robust across diverse signers and recording conditions.
Looking forward, the authors identify low-resource adaptation, ethical AI development and global accessibility as the defining frontiers. Low-rank adaptation and prefix-tuning techniques make it feasible to bootstrap systems for sign languages with only small corpora, while multilingual pretraining methods like mT5 hint at translation models that span many languages at once. Ethics is not an afterthought: Deaf communities have historically been skeptical of technologies developed without their involvement, and the survey’s framing of translation as a bridge for equitable access implicitly demands that future systems be co-designed with Deaf users rather than imposed on them.
The significance of this survey lies less in any single breakthrough than in its synthesis of a fast-moving field at an inflection point. Gloss-free sign language translation, once a niche goal, is now the mainstream research program, powered by the same foundation-model infrastructure reshaping the rest of artificial intelligence. If the roadmap holds, the coming years could see Deaf and hearing people converse directly through cameras and phones, with neural models handling the linguistic heavy lifting, a development that would rank among the most socially consequential applications of the large language model era.
Subject of Research: Gloss-free sign language translation using large language models and vision-language models for Deaf communication
Article Title: Leveraging large and vision language models for gloss-free sign language translation in deaf communication: a survey
Article References: Danish, S., Khan, S. U., Alghamdi, R. A., & Alghamdi, N. S. (2026). Leveraging large and vision language models for gloss-free sign language translation in deaf communication: a survey. International Journal of Machine Learning and Cybernetics, 17(10), Article 479. https://doi.org/10.1007/s13042-026-03323-x
Image Credits: AI Generated
DOI: 10.1007/s13042-026-03323-x
Keywords: sign language translation, gloss-free, large language models, vision-language models, multimodal AI, Deaf communication, neural machine translation, low-resource languages, Arabic sign language, vision-language pretraining, accessibility, transformers
Cite Scienmag News
Denise Maddox. (September 27, 2026). AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language. Scienmag. https://scienmag.com/ai-without-glosses-how-large-language-models-are-learning-to-translate-sign-language/
Denise Maddox. "AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language." Scienmag, 27 September 2026, https://scienmag.com/ai-without-glosses-how-large-language-models-are-learning-to-translate-sign-language/. Accessed 27 September 2026.
Denise Maddox. "AI Without Glosses: How Large Language Models Are Learning to Translate Sign Language." Scienmag. September 27, 2026. https://scienmag.com/ai-without-glosses-how-large-language-models-are-learning-to-translate-sign-language/

