<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>deep learning for visual speech recognition &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/deep-learning-for-visual-speech-recognition/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 09:57:02 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>deep learning for visual speech recognition &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Ensemble AI Reads Amharic Lips With Record Accuracy</title>
		<link>https://scienmag.com/ensemble-ai-reads-amharic-lips-with-record-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 09:57:02 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in visual speech recognition for underrepresented languages]]></category>
		<category><![CDATA[Amharic]]></category>
		<category><![CDATA[Assistive Technology]]></category>
		<category><![CDATA[Capsule Networks]]></category>
		<category><![CDATA[challenges in low-resource language processing]]></category>
		<category><![CDATA[classical image processing in speech recognition]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[datasets for Amharic lip reading]]></category>
		<category><![CDATA[deep learning for visual speech recognition]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[ensemble models for lip reading accuracy]]></category>
		<category><![CDATA[Ethiopia]]></category>
		<category><![CDATA[Ethiopian language technology development]]></category>
		<category><![CDATA[Gabor filters]]></category>
		<category><![CDATA[gradient boosting]]></category>
		<category><![CDATA[lip reading]]></category>
		<category><![CDATA[Lip reading technology for Amharic]]></category>
		<category><![CDATA[low-resource languages]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[multimedia tools for speech recognition]]></category>
		<category><![CDATA[neural network architectures for lip reading]]></category>
		<category><![CDATA[speech recognition in Semitic languages]]></category>
		<category><![CDATA[Transformer models in lip reading]]></category>
		<category><![CDATA[visual speech recognition]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221838</guid>

					<description><![CDATA[Researchers at Bahir Dar University have built an ensemble lip-reading model for Amharic that combines Gabor filters, Capsule Networks, and gradient boosting to achieve 95.63 percent classification accuracy.]]></description>
										<content:encoded><![CDATA[<p>Computers that can read lips have long been a staple of science fiction, but turning that fantasy into working technology has proven stubbornly difficult, especially for languages that lack the vast datasets that power English systems. Now, researchers at Bahir Dar University in Ethiopia report a significant step forward for Amharic, the second most widely spoken Semitic language in the world, with a lip-reading model that classifies spoken content from mouth movements alone with an accuracy of 95.63 percent. The work, published in the journal Multimedia Tools and Applications, combines classical image-processing techniques with modern deep learning architectures in an ensemble design that outperforms simpler approaches by a wide margin.</p>
<p>The research team, consisting of Zelalem Fetene from the Department of Statistics and Data Science and Seffi Gebeyehu from the Department of Software Engineering, set out to address a glaring gap in the field. While recent years have seen remarkable advances in visual speech recognition driven by three-dimensional convolutional neural networks, Transformer-based models, and attention mechanisms, these breakthroughs have overwhelmingly benefited well-resourced languages such as English and Mandarin. Amharic, spoken by tens of millions of people primarily in Ethiopia, has remained largely on the sidelines, with no adequately resourced or trained system capable of supporting a consistent speaker-independent lip-reading model.</p>
<p>The new system is built on a three-stage pipeline that begins with careful preprocessing of video frames. Raw images of speakers&#8217; faces are converted to grayscale, enhanced using contrast-limited adaptive histogram equalization, and resized to a standard dimension suitable for downstream analysis. This preprocessing stage is far from trivial: lip-reading systems are notoriously sensitive to lighting conditions, camera quality, and background clutter, and the authors draw on a body of literature showing that benchmarking and standardizing preprocessing methods can materially improve facial image analysis. By normalizing these variables early in the pipeline, the model receives cleaner, more consistent inputs than it would from raw video alone.</p>
<p>The heart of the feature-extraction stage lies in an unusual pairing of Gabor filters and Capsule Networks. Gabor filters, which have been used for decades in texture analysis, fingerprint classification, and face recognition, are mathematical operators tuned to detect oriented edges and fine-grained spatial patterns at multiple frequencies and orientations. Applied to lip images, they capture the subtle textural and geometric signatures of different mouth shapes. Capsule Networks, introduced by Geoffrey Hinton and colleagues in 2017, take a different approach: instead of simply detecting features, they preserve spatial hierarchies and relationships between parts, using a process called dynamic routing to determine how lower-level features combine into higher-level representations. For lip reading, where the relative position and deformation of the lips, teeth, and surrounding tissue carry meaning, this structural sensitivity is valuable.</p>
<p>Combining the two techniques allows the system to exploit complementary strengths. The Gabor filters provide robust, orientation-sensitive texture descriptors that are relatively invariant to illumination changes, while the Capsule Network learns to encode the part-whole relationships that distinguish one Amharic viseme, or visual speech unit, from another. The authors report that this hybrid feature-extraction strategy substantially improved discrimination compared with a Capsule Network operating alone. On their test set of 15,228 image components, the full ensemble achieved 95.63 percent classification accuracy, whereas a baseline Capsule-only model reached 91.4 percent. That four-percentage-point difference may sound modest, but in classification tasks with many categories, it represents a substantial reduction in error.</p>
<p>The final stage of the pipeline is a stacking Gradient Boosting Machine classifier. Gradient boosting, popularized by frameworks such as XGBoost, builds a strong predictor by sequentially adding decision trees that correct the errors of their predecessors. In a stacking configuration, the outputs of the feature-extraction stage are fed to the boosting algorithm, which learns how best to weigh and combine the evidence for each class. The authors cite prior work demonstrating that stacked ensemble learning can improve classification across domains ranging from multispectral image analysis to medical diagnostics, and their results suggest the same principle holds for visual speech. The ensemble approach also helps guard against the overfitting that can plague deep networks trained on limited data, a persistent concern for low-resource languages.</p>
<p>The term speaker-independent deserves careful attention, and the authors are refreshingly candid about its scope. In this study, speaker independence refers to robustness across different speakers within fixed Amharic content categories; the model was not evaluated on entirely unseen phrases. This distinction matters because true speaker-independent, cross-content lip reading remains one of the hardest open problems in the field, even for English. By acknowledging that cross-content evaluation on unseen phrases was not conducted, and by explicitly calling for such evaluation in future work, the researchers set a clear benchmark for what comes next. The honesty is notable in a field where inflated claims are not uncommon, and it does not diminish the practical significance of the result.</p>
<p>The potential applications are considerable. For the estimated hundreds of millions of people worldwide with hearing impairments, lip reading is already a vital communication strategy, but human lip reading is exhausting, error-prone, and highly dependent on context and the skill of the reader. An automated system could power assistive communication devices that transcribe speech in real time from video alone, giving deaf and hard-of-hearing users a new tool in conversations where sign language interpreters are unavailable. Equally important are noisy environments, where audio-based speech recognition fails outright: factories, construction sites, cockpits, and crowded public spaces all degrade microphone signals, but visual speech remains available to a camera. A reliable lip-reading system could serve as a transcription and command interface in precisely the settings where conventional voice assistants break down.</p>
<p>There is also a broader significance to building such systems for Amharic specifically. Language technology has a well-documented tendency to concentrate benefits in a handful of wealthy, English-speaking markets, leaving speakers of lower-resource languages with inferior tools or none at all. Recent research on lip reading for low-resource languages has explored transferring general speech knowledge learned from high-resource languages to language-specific models, and the Ethiopian work adds to this growing movement. The authors note that all videos used in the study were publicly available, sourced in accordance with YouTube&#8217;s terms of service, and that no personally identifiable information was included, an ethical posture appropriate for research involving facial imagery.</p>
<p>Looking ahead, the researchers outline a clear roadmap. They plan to increase the sample size to include a wider diversity of Ethiopian speaker profiles, which should strengthen the model&#8217;s generalization across faces, skin tones, and articulation styles. They also aim to optimize the model for real-time deployment on mobile devices, a step that would transform the system from a laboratory demonstration into a practical tool that could run on the smartphones that are increasingly the primary computing platform across Africa. Real-time mobile lip reading demands aggressive compression and inference optimization, since the Gabor filtering and capsule routing operations are computationally demanding, but the payoff would be assistive technology accessible to ordinary users rather than institutions. And with cross-content evaluation on unseen phrases flagged as the next scientific milestone, the Bahir Dar team has sketched a path from a strong fixed-category classifier toward a genuinely open-vocabulary Amharic lip-reading system. If that trajectory holds, one of the world&#8217;s major Semitic languages may soon have visual speech technology to rival anything available for English, and millions of speakers could gain access to communication tools that have long been taken for granted elsewhere.</p>
<p><strong>Subject of Research:</strong> Speaker-independent Amharic lip reading using an ensemble of Gabor filters, Capsule Networks, and a stacking gradient boosting classifier</p>
<p><strong>Article Title:</strong> A speaker-independent Amharic lip reading model using an ensemble approach</p>
<p><strong>Article References:</strong> Fetene, Z., &amp; Gebeyehu, S. (2026). A speaker-independent Amharic lip reading model using an ensemble approach. <em>Multimedia Tools and Applications, 85</em>(10), Article 782. <a href="https://doi.org/10.1007/s11042-026-21920-4" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21920-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21920-4" rel="noopener noreferrer">10.1007/s11042-026-21920-4</a></p>
<p><strong>Keywords:</strong> lip reading, Amharic, Gabor filters, Capsule Networks, gradient boosting, ensemble learning, visual speech recognition, low-resource languages, Ethiopia, machine learning, assistive technology, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221838</post-id>	</item>
	</channel>
</rss>
