<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>deep learning models for emotion recognition in virtual classrooms &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/deep-learning-models-for-emotion-recognition-in-virtual-classrooms/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 00:06:12 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>deep learning models for emotion recognition in virtual classrooms &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Reads Students&#8217; Faces to Measure Attention in Online Classes</title>
		<link>https://scienmag.com/ai-reads-students-faces-to-measure-attention-in-online-classes/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 00:06:12 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[affective computing]]></category>
		<category><![CDATA[AI-based facial recognition for online student engagement]]></category>
		<category><![CDATA[attentiveness index]]></category>
		<category><![CDATA[challenges of monitoring attention in online education]]></category>
		<category><![CDATA[cognitive science insights into emotional states during e-learning]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[DAiSEE dataset]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for emotion recognition in virtual classrooms]]></category>
		<category><![CDATA[detecting boredom and frustration during online lessons]]></category>
		<category><![CDATA[e-learning]]></category>
		<category><![CDATA[engagement detection]]></category>
		<category><![CDATA[facial expression recognition]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[impact of affective states on online learning outcomes]]></category>
		<category><![CDATA[importance of student emotional states in digital learning]]></category>
		<category><![CDATA[online education]]></category>
		<category><![CDATA[real-time affective state detection in e-learning]]></category>
		<category><![CDATA[real-time analytics]]></category>
		<category><![CDATA[remote classroom feedback systems using facial analysis]]></category>
		<category><![CDATA[role of]]></category>
		<category><![CDATA[technological advancements in virtual classroom engagement assessment]]></category>
		<category><![CDATA[webcam-based attention measurement in remote education]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=236206</guid>

					<description><![CDATA[Researchers have developed a real-time computer vision system that reads students' facial expressions during online lessons and converts four learning-centered emotional states into a single attentiveness score for instructors.]]></description>
										<content:encoded><![CDATA[<p>For millions of students logging into virtual classrooms every day, the question of whether anyone is actually paying attention has long been unanswerable. In a physical lecture hall, an instructor can scan the room for slumped shoulders, wandering eyes, and furrowed brows, then adjust the pace of a lesson on the fly. Online, that feedback loop collapses. A grid of silent video tiles offers little insight into who is absorbing the material and who has mentally checked out. Now, a team of researchers at Tel Aviv University and Washington State University has built a system that tries to restore that lost channel of communication, using nothing more than a standard webcam and a set of deep learning models that read a learner&#8217;s face in real time.</p>
<p>The new framework, described in the journal Multimedia Tools and Applications, tackles a problem that researchers call affective state recognition in e-learning. Rather than trying to detect the seven universal emotions that dominate much of facial expression research, the system focuses on four learning-centered states that education scientists have identified as especially meaningful during study: boredom, engagement, confusion, and frustration. These states matter because decades of cognitive science research show they shape how people learn. Positive emotional states are associated with gains in creative thinking, while negative states push learners toward prioritizing performance over mastery. Boredom in particular has been repeatedly linked to suboptimal learning outcomes, whereas mild confusion, somewhat counterintuitively, can signal that a student is grappling productively with difficult material.</p>
<p>At the heart of the system is a multioutput classification model built on a parallel-branch architecture. The researchers trained two versions: one based on a lightweight convolutional neural network with roughly 470,000 parameters, and another built on EfficientNetB2, a more capable backbone with about 31.7 million parameters that remains efficient enough for cloud deployment. Each branch of the network is dedicated to a single affective state and outputs a probability vector across four intensity levels, from very low to very high. The final output is a combined vector capturing the predicted intensity of boredom, confusion, engagement, and frustration simultaneously. This design distinguishes the work from most prior studies, which typically classified only the single dimension of engagement and ignored the richer emotional context surrounding it.</p>
<p>Training relied on DAiSEE, the first multilabel video dataset created for engagement recognition in the wild. It contains 9,068 ten-second video clips of 112 individuals, recorded at full high-definition resolution with ordinary webcams, and annotated with crowd-sourced labels that were correlated against a gold standard produced by expert psychologists. The dataset is notoriously imbalanced: high-intensity engagement labels dominate, while lower-intensity labels are comparatively rare, and the opposite pattern holds for boredom, confusion, and frustration. To counteract the bias this imbalance would otherwise introduce, the team employed a categorical focal loss, a function that applies a modulating term to standard cross-entropy so that learning concentrates on hard, easily misclassified examples rather than being swamped by the majority class.</p>
<p>Preprocessing was deliberately aggressive. Each ten-second clip originally contains 300 frames, but because facial expressions in instructional settings evolve slowly, on the order of one to two seconds, the pipeline subsamples to just ten frames per video, a 96.7 percent reduction in data volume. Each retained frame passes through a Viola-Jones Haar Cascade detector to isolate the face, and frames where facial landmarks or pupils cannot be reliably detected are discarded. Valid crops are resized to 64-by-64 pixels and converted to normalized grayscale. The result is a compact, standardized input that minimizes computational cost while preserving the affective signals that matter, an essential property for a system intended to run continuously during live lectures.</p>
<p>The most novel contribution, however, is not the classifier itself but what the researchers call the attentiveness index. A single classification of four emotional states does not directly tell an instructor whether a student is attentive, so the team needed a way to compress those predictions into one interpretable number. They had multiple instructors score a subset of dataset videos for attentiveness on a scale of one to ten, then applied multiple linear regression to learn weights for each affective state. The resulting formula weights engagement most heavily and positively, gives confusion a smaller positive weight, and assigns negative weights to boredom and frustration, with boredom penalized far more severely. Crucially, these learned coefficients align with the existing cognitive science literature: boredom is strongly associated with poor learning, frustration more weakly so, and moderate confusion often accompanies productive learning.</p>
<p>The index held up under statistical scrutiny. On an unseen test set, the computed index correlated strongly and significantly with the instructor-annotated ground truth, yielding a Pearson correlation coefficient of 0.700 and a Spearman rank coefficient of 0.585, both statistically significant. Three-fold cross-validation produced a mean R-squared of 0.495, and a t-test confirmed the model&#8217;s predictive power at p equals 0.0246. Meanwhile, the underlying classifiers achieved state-of-the-art performance on DAiSEE. The CNN-based model reached accuracies of 73.04 percent on boredom, 79.98 percent on engagement, 80.17 percent on confusion, and 85.37 percent on frustration, while the EfficientNet variant performed comparably, peaking at 80.32 percent on engagement. The researchers argue these results outperform prior end-to-end approaches on the same dataset, including methods built on C3D, I3D, and hybrid architectures combining EfficientNet with temporal convolutional and recurrent networks.</p>
<p>What elevates the work from a modeling exercise to a practical tool is the end-to-end pipeline built around it. The system is deployed on a cloud server and supports simultaneous multi-user logins through a web client built with Flask, JavaScript, and HTML5 Canvas. During a live session, the learner&#8217;s webcam feed is analyzed in real time, with detected affective states displayed alongside the learning content. Instructors receive a separate dashboard offering both graphical and tabular analytics: class-level trends in average attentiveness over time, post-session distributions of affective states, and timestamped logs of individual learners&#8217; dominant states and index scores. When cumulative class engagement drops below a threshold, the instructor receives an alert, and the analytics can pinpoint exactly when attention sagged during a lecture, for instance, or flag moments when confusion spiked while a difficult concept was being explained. The system can even aggregate data across multiple lectures to suggest how to structure future sessions for maximum engagement.</p>
<p>The researchers are explicit that the tool is meant as a pedagogical aid, not a surveillance instrument. They caution that affective computing models can be sensitive to lighting, camera quality, and demographic variation, and they strongly advise against punitive or comparative uses of the system. Real-world deployments, they stress, must be strictly opt-in, with informed consent from all participants, and individual analytics should be disabled by default in classroom settings in favor of aggregated, class-level feedback. They also acknowledge a limitation long noted in the field: basic universal emotions are rare during genuine learning sessions, which is precisely why the focus on learning-centered states matters.</p>
<p>Future directions outlined by the team include multimodal fusion, combining facial affect with head movement, gaze tracking, and noninvasive EEG signals, along with explainable AI techniques that would visually highlight which facial features drive each prediction. They also plan consent-based periodic self-reports from learners to calibrate the model against individual differences, and the construction of a new, demographically balanced dataset to validate generalizability across racial and cultural backgrounds. If those efforts succeed, the humble webcam could become one of the most informative instruments in the virtual classroom, giving online instructors something they have lacked since lectures moved to screens: a live read on the minds in the room.</p>
<p><strong>Subject of Research:</strong> Real-time computer vision analysis of learner affective states and attentiveness in online education</p>
<p><strong>Article Title:</strong> Learner attentiveness and engagement analysis in online education using computer vision</p>
<p><strong>Article References:</strong> Gogawale, S., Deshpande, M., Kumar, P., &amp; Ben-Gal, I. (2026). Learner attentiveness and engagement analysis in online education using computer vision. <em>Multimedia Tools and Applications, 85</em>(9), Article 743. <a href="https://doi.org/10.1007/s11042-026-21884-5" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21884-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21884-5" rel="noopener noreferrer">10.1007/s11042-026-21884-5</a></p>
<p><strong>Keywords:</strong> computer vision, online education, affective computing, engagement detection, deep learning, convolutional neural networks, DAiSEE dataset, attentiveness index, e-learning, facial expression recognition, focal loss, real-time analytics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">236206</post-id>	</item>
	</channel>
</rss>
