<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>real-time sign language video creation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/real-time-sign-language-video-creation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 09:32:08 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>real-time sign language video creation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Turns Marathi Text Into Realistic Indian Sign Language Videos</title>
		<link>https://scienmag.com/ai-turns-marathi-text-into-realistic-indian-sign-language-videos/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 09:32:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[accessibility]]></category>
		<category><![CDATA[accessibility tools for deaf and hard-of-hearing in India]]></category>
		<category><![CDATA[AI-based sign language generation for low-resource languages]]></category>
		<category><![CDATA[custom translation networks for sign language]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[facial expressions and handshapes in sign language AI]]></category>
		<category><![CDATA[generative adversarial network]]></category>
		<category><![CDATA[generative adversarial networks for sign language animation]]></category>
		<category><![CDATA[Indian Sign Language]]></category>
		<category><![CDATA[Indian Sign Language (ISL) video synthesis]]></category>
		<category><![CDATA[low-resource languages]]></category>
		<category><![CDATA[low-resource sign language datasets and AI solutions]]></category>
		<category><![CDATA[Marathi]]></category>
		<category><![CDATA[Marathi text to Indian Sign Language video translation]]></category>
		<category><![CDATA[OpenPose]]></category>
		<category><![CDATA[pose estimation]]></category>
		<category><![CDATA[real-time sign language video creation]]></category>
		<category><![CDATA[sign language production]]></category>
		<category><![CDATA[sign language translation systems for regional languages]]></category>
		<category><![CDATA[sign language translation technology for Marathi]]></category>
		<category><![CDATA[text-to-gloss translation]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[video synthesis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234502</guid>

					<description><![CDATA[Researchers in Pune have developed a pose-enhanced generative adversarial network that translates Marathi text into fluent Indian Sign Language gloss and renders it as realistic, near real-time signer video.]]></description>
										<content:encoded><![CDATA[<p>For more than 70 million deaf and hard-of-hearing people in India, access to everyday information often depends on whether someone nearby can sign. Now a team of researchers in Pune has built an artificial intelligence pipeline that takes written Marathi—a language with almost no sign language technology supporting it—and converts it into fluent, visually convincing videos of Indian Sign Language (ISL). The work, published in Multimedia Tools and Applications, combines a custom translation network with a generative adversarial framework that animates a synthetic signer, and it runs fast enough to hint at near real-time use.</p>
<p>The challenge the researchers set out to solve is one of resources. Sign language translation systems for English, Chinese, and several European languages benefit from large parallel corpora—thousands of hours of signed video aligned with spoken-language text. Marathi, spoken by more than 80 million people, has no comparable dataset for ISL. Indian Sign Language itself is distinct from American or British sign languages, with its own grammar, its own spatial referencing system, and heavy reliance on facial expressions and handshapes that carry grammatical meaning. Building a text-to-video system for this low-resource pairing means solving two hard problems at once: translating Marathi sentences into a signed representation, and then rendering that representation as a moving human figure that looks natural rather than robotic.</p>
<p>The team, led by Prachi Pramod Waghmare with Ashwini Mangesh Deshpande and Supriya A. Mangale of Cummins College of Engineering for Women, affiliated with Savitribai Phule Pune University, addressed the first problem with what they call a Sparse Multi-Hierarchical Transformer. Transformers, the architecture behind modern large language models, normally attend to every token in a sequence, which becomes computationally expensive as sentences grow. The sparse, hierarchical variant prunes that attention: it organizes the Marathi input and the target gloss—a gloss being a written label for a sign, essentially the bridge vocabulary between a spoken and a signed language—into layers, attending selectively across words, phrases, and sentence-level structure. The result, according to the paper, is an efficient mapping from Marathi text to ISL gloss that preserves the reordering needed when a subject-verb-object spoken sentence becomes the spatially organized structure of a signed one.</p>
<p>The numbers reported for this translation stage are striking. The framework achieves BLEU-1, BLEU-2, and BLEU-3 scores of 99.93, with a BLEU-4 score of 56.23. BLEU scores measure how closely machine-generated output matches human reference translations, with the higher-order n-gram scores being progressively harder to satisfy because they demand matching longer consecutive sequences. Near-perfect unigram, bigram, and trigram scores indicate that virtually every sign and sign pair in the generated gloss sequences appears in the right place; the drop at BLEU-4 reflects the familiar difficulty of reproducing longer fluent stretches exactly, which is where sign order variability and synonymy among glosses bite hardest.</p>
<p>Once the gloss sequence exists, the system must turn it into video, and this is where the generative adversarial network comes in. GANs pit two neural networks against each other: a generator that produces candidate outputs and a discriminator that tries to tell them apart from real examples. Over training, the generator improves until its outputs fool the discriminator. The researchers&#8217; variant, which they describe as a Pose-Enhanced Video-Transformer GAN, conditions the generation on explicit skeletal information rather than asking the network to invent human anatomy from scratch. OpenPose, a widely used landmark detection library, supplies the body, hand, and face keypoints that define a signer&#8217;s posture at every moment. These landmarks act as a structural scaffold: the generator learns to paint realistic video content onto a skeleton that already encodes where the arms, hands, and facial features must be for each sign.</p>
<p>The rendering stage, called a Pose-Adapted Progressive Video Synthesizer, builds the video in stages rather than all at once. Progressive synthesis—refining output from coarse to fine—has proven effective in image generation, and here it is adapted to the temporal dimension, so that the motion of the signer emerges coherently across frames instead of flickering between inconsistent poses. Temporal coherence is the notorious weakness of frame-by-frame generation: without it, fingers jitter, backgrounds shimmer, and the signer appears to morph. The pose conditioning plus progressive refinement is designed to suppress exactly those artifacts, keeping the signer&#8217;s body stable while the hands and face execute the signs.</p>
<p>Quality metrics for the generated videos back up the design. The framework reports a structural similarity index (SSIM) of 0.912, a measure of how closely the generated frames resemble real signer footage in structure and luminance, on a scale where 1.0 is identical. Peak signal-to-noise ratio (PSNR) comes in at 31.7 decibels, a reconstruction fidelity measure in which values above 30 dB are generally considered good for video synthesis. Perhaps most important for practical deployment, the system generates video at 15 frames per second with an average inference time of 120 milliseconds per sample. That is not broadcast speed, but it approaches the threshold where a translation service could produce signed video on demand rather than after a long offline render, which is what most earlier sign language production systems required.</p>
<p>The significance of the work extends beyond its benchmark numbers. Most sign language production research targets languages with abundant data, and several recent efforts—gloss-free translation models, sign language large language models, and diffusion-based generators—assume resources that simply do not exist for Marathi-ISL. By demonstrating a pipeline that works with a limited dataset and a single signer, the Pune team shows a viable path for other under-resourced language-sign pairs, of which there are hundreds worldwide. The hierarchical sparse attention mechanism also speaks to efficiency: sign language generation is computationally heavy because it couples language modeling with video synthesis, and trimming the attention cost at the translation stage buys headroom for the rendering stage.</p>
<p>The authors are candid about the limitations. Their Marathi-ISL dataset is small and features only one signer, which means the generated videos inherit a single person&#8217;s signing style, body proportions, and appearance. Generalizing to multiple signers, regional variation in ISL, and spontaneous rather than scripted sentences remains future work. The researchers frame the current system as a foundation for an adaptable sign language generation platform rather than a finished product. Ethical safeguards were observed: informed consent was obtained from the human signer whose recordings underpin the dataset, and the authors report no competing interests.</p>
<p>Still, the trajectory is clear. If systems like this one can be scaled to larger, multi-signer datasets and integrated with speech recognition, a deaf person could eventually receive signed video versions of news articles, government notices, classroom material, or medical instructions in their own sign language, generated on the fly. For Marathi speakers in the deaf community, who have until now been largely invisible to sign language technology, this research is an early but concrete step toward that future—and a template for bringing the benefits of generative AI to the languages and communities that need it most.</p>
<p><strong>Subject of Research:</strong> Machine translation of Marathi text to Indian Sign Language gloss and pose-guided GAN-based sign language video generation</p>
<p><strong>Article Title:</strong> Efficient marathi text-to-gloss translation and realistic isl video generation using pose-enhanced GAN</p>
<p><strong>Article References:</strong> Waghmare, P. P., Deshpande, A. M., &amp; Mangale, S. A. (2026). Efficient marathi text-to-gloss translation and realistic isl video generation using pose-enhanced GAN. <em>Multimedia Tools and Applications, 85</em>(9), Article 747. <a href="https://doi.org/10.1007/s11042-026-21908-0" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21908-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21908-0" rel="noopener noreferrer">10.1007/s11042-026-21908-0</a></p>
<p><strong>Keywords:</strong> Indian Sign Language, Marathi, text-to-gloss translation, generative adversarial network, pose estimation, OpenPose, transformer, video synthesis, sign language production, low-resource languages, accessibility, deep learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234502</post-id>	</item>
	</channel>
</rss>
