<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>VGG 16 &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/vgg-16/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 23 Sep 2026 21:44:27 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>VGG 16 &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Hybrid AI With Attention Achieves Record Face Expression Recognition Accuracy</title>
		<link>https://scienmag.com/hybrid-ai-with-attention-achieves-record-face-expression-recognition-accuracy/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 21:44:27 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[attention-driven modules in computer vision]]></category>
		<category><![CDATA[benchmark datasets for emotion classification]]></category>
		<category><![CDATA[challenges of real-world facial expression recognition]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[convolutional neural networks for emotion detection]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for emotion analysis]]></category>
		<category><![CDATA[facial expression recognition]]></category>
		<category><![CDATA[FER2013]]></category>
		<category><![CDATA[head pose variation in facial recognition]]></category>
		<category><![CDATA[high-accuracy emotion recognition models]]></category>
		<category><![CDATA[hybrid deep learning architecture]]></category>
		<category><![CDATA[multi-dataset evaluation in facial expression analysis]]></category>
		<category><![CDATA[multi-scale convolution]]></category>
		<category><![CDATA[occlusion]]></category>
		<category><![CDATA[pose variation]]></category>
		<category><![CDATA[RAF-DB]]></category>
		<category><![CDATA[resilience of AI models to visual variability]]></category>
		<category><![CDATA[ResNet-50]]></category>
		<category><![CDATA[robustness to occlusions in facial analysis]]></category>
		<category><![CDATA[VGG 16]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=210545</guid>

					<description><![CDATA[A Moroccan research team has developed a hybrid convolutional neural network with local and global attention modules that reaches 95.87 percent accuracy in facial expression recognition across multiple benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Facial expression recognition has long been one of the most deceptively difficult problems in computer vision. While humans read emotions from faces almost instantaneously, machines struggle with the sheer variability of real-world conditions: heads tilt, light shifts, glasses glint, hands partially cover mouths, and every face is unique. A new study published in Multimedia Tools and Applications by Abderrahim Ouza and Ali Choukri of Ibn Tofail University in Kenitra, Morocco, together with Mohamed El Ghmary of Mohammed V University in Rabat, tackles precisely these weaknesses. The team has built a hybrid deep learning architecture that combines two classic convolutional neural network backbones with attention-driven modules, and the results are striking. Across multiple benchmark datasets, including RAF-DB, FER2013 and CK+, the model reaches an accuracy of 95.87 percent, while showing markedly stronger resilience to occlusions and head pose changes than conventional approaches, even when tested on data drawn from a different dataset than the one used in training.</p>
<p>The central insight behind the work is that robust expression recognition requires two complementary kinds of visual understanding. The first is local: emotions are encoded in subtle movements of specific facial regions, such as the slight raising of the inner eyebrows that signals worry, or the tightening of the muscles around the eyes that distinguishes a genuine smile from a polite one. The second is global: the spatial arrangement of the entire face, the relationships between regions, and the overall configuration of features all contribute to how an expression is perceived. Most existing architectures emphasize one of these dimensions at the expense of the other. Purely convolutional pipelines excel at extracting local texture patterns but can miss long-range dependencies, while transformer-style models capture global context yet sometimes overlook the fine-grained local cues that carry the emotional signal.</p>
<p>To resolve this tension, the researchers designed a hybrid model that mixes several ResNet-50 branches with VGG-16 backbones. ResNet-50, with its residual connections, allows very deep networks to be trained without the vanishing gradient problems that once stalled deep learning, and it produces rich hierarchical feature maps. VGG-16, an older but still powerful architecture, is prized for its simple stack of small convolution filters that systematically capture texture at fine scale. By fusing representations from both families of networks, the model inherits the strengths of each: the depth-adapted abstractions of ResNet-50 and the fine-textured detail of VGG-16. The authors then attach two purpose-built modules on top of this fused backbone to refine what the networks see and how they interpret it.</p>
<p>The first of these is the Local Feature Enhancement, or LFE, module. Its role is to amplify the subtle variations that occur in the principal facial regions most relevant to emotional expression. In practical terms, the module learns to weight the feature maps so that informative local patterns, such as the curvature of the lips or the degree of eyelid opening, are emphasized while less relevant background or redundant information is suppressed. This kind of learned emphasis matters enormously in expression recognition, where the difference between two emotions can hinge on a few pixels. Without such focusing, a network may spread its capacity evenly across the image and fail to register the discriminative details that separate sadness from neutrality in a low-resolution photograph.</p>
<p>The second component, the Global Information Association, or GIA, module, works at the opposite end of the spatial scale. It captures the global structure of the face, modeling how distant regions of the image relate to one another. An expression is not merely a sum of parts; a raised mouth corner reads as joy only in combination with relaxed eyes, whereas the same mouth movement alongside furrowed brows may indicate a different state entirely. By associating information across the whole face, the GIA module gives the network a holistic understanding of the expression, allowing it to integrate the locally enhanced features into a coherent global interpretation.</p>
<p>Between these two modules, the architecture deploys a multi-scale convolutional design together with a self-attention-based approach known as ACMix. Multi-scale convolution means that the network examines the face simultaneously at several receptive field sizes, from small windows that catch fine wrinkles to larger windows that span whole facial regions. This mirrors the way human vision processes faces at multiple resolutions, and it helps the model remain accurate when faces appear at different sizes or distances. The self-attention mechanism, meanwhile, allows every part of the feature representation to dynamically attend to every other part, computing relevance weights on the fly. ACMix blends convolutional inductive biases with attention-style global reasoning, which improves the model&#8217;s stability under the challenging conditions, such as unusual illumination or partial occlusion, that routinely break less flexible architectures.</p>
<p>Evaluation was carried out on three of the most widely used benchmarks in the field. RAF-DB contains thousands of real-world, in-the-wild faces annotated with expressions, capturing the messy variety of unconstrained photography. FER2013, the dataset used for training and evaluation, is freely available through Kaggle and comprises more than 35,000 grayscale images cataloged across seven emotion categories: anger, disgust, fear, happiness, sadness, surprise and neutral expressions, including multiple views of the same subjects, making it the largest such resource of its kind. CK+ offers carefully controlled laboratory sequences with well-defined expressions, providing a complementary test of discriminative power. Achieving 95.87 percent accuracy across these tests places the hybrid model among the strongest performers reported in the literature, and the authors emphasize that its advantage grows when conditions become difficult, such as when glasses, scarves or hands obscure parts of the face, or when the head is rotated away from the camera.</p>
<p>Particularly notable is the model&#8217;s behavior in cross-dataset testing, one of the most demanding evaluations in expression recognition. A system trained on one dataset and tested on another must cope with differences in lighting, demographics, image quality and labeling conventions, and accuracy typically drops sharply in such settings. The Moroccan team reports that their model shows superior robustness under exactly these conditions, suggesting that the combination of local enhancement, global association and attention produces features that generalize rather than merely memorize dataset-specific quirks. This generalization is precisely the property that separates laboratory demonstrations from technology that can function reliably outside controlled environments, and it is a key reason the authors argue their approach is suited to real-world applications.</p>
<p>The potential applications span several domains where reading a face accurately is not a luxury but a safety requirement. In healthcare, precise identification of emotions can support monitoring of patients who cannot communicate verbally, alerting clinicians to pain or distress. In driver monitoring systems, a camera that reliably detects fatigue, distraction or confusion even when the driver&#8217;s face is partly occluded or turned could help prevent accidents. The authors point to these use cases, where missing a critical emotional state carries real consequences, as the natural fit for their architecture. The work was conducted with institutional resources at Ibn Tofail University and Sidi Mohamed Ben Abdellah University in Morocco, without external funding, and the datasets underpinning the study are openly available under appropriate licenses, which should make it straightforward for other research groups to reproduce and extend the results.</p>
<p>As facial expression recognition moves from the research bench into cars, clinics and everyday devices, studies like this one illustrate the recipe that is emerging as the field&#8217;s consensus: hybrid backbones that unite complementary convolutional designs, explicit modules for local detail and global context, and attention mechanisms that let the network decide what matters in each image. By combining ResNet-50 and VGG-16 with the LFE, GIA and ACMix components, the Moroccan team has demonstrated that this layered strategy can push accuracy toward 96 percent while keeping the model dependable when the real world, with its shadows, angles and obstructions, refuses to cooperate. The full details of the architecture and experiments are available in the article published in Multimedia Tools and Applications, volume 85, article number 775.</p>
<p><strong>Subject of Research:</strong> Hybrid deep learning with attention mechanisms for robust facial expression recognition</p>
<p><strong>Article Title:</strong> Hybrid convolutional neural network with attention mechanism to provide robust face expression recognition</p>
<p><strong>Article References:</strong> Ouza, A., Ghmary, M. E., &amp; Choukri, A. (2026). Hybrid convolutional neural network with attention mechanism to provide robust face expression recognition. <em>Multimedia Tools and Applications, 85</em>(10), Article 775. <a href="https://doi.org/10.1007/s11042-026-21933-z" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21933-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21933-z" rel="noopener noreferrer">10.1007/s11042-026-21933-z</a></p>
<p><strong>Keywords:</strong> facial expression recognition, deep learning, convolutional neural networks, attention mechanism, ResNet-50, VGG-16, multi-scale convolution, occlusion, pose variation, RAF-DB, FER2013, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">210545</post-id>	</item>
		<item>
		<title>Ensemble of Five Neural Networks Creates Richer Deep Dream Art</title>
		<link>https://scienmag.com/ensemble-of-five-neural-networks-creates-richer-deep-dream-art/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 05:11:44 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI art]]></category>
		<category><![CDATA[AI-generated dreamlike imagery]]></category>
		<category><![CDATA[bagging]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[convolutional neural networks for art]]></category>
		<category><![CDATA[Deep Dream]]></category>
		<category><![CDATA[Deep Dream ensemble]]></category>
		<category><![CDATA[deep learning for artistic visualization]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[ensemble learning in deep dreaming]]></category>
		<category><![CDATA[ensemble methods in AI art creation]]></category>
		<category><![CDATA[hierarchical feature detection in deep dreaming]]></category>
		<category><![CDATA[image feature aggregation in neural networks]]></category>
		<category><![CDATA[Inception v3]]></category>
		<category><![CDATA[Inception-ResNet-V2]]></category>
		<category><![CDATA[multi-model deep dream generation]]></category>
		<category><![CDATA[neural network hallucination techniques]]></category>
		<category><![CDATA[neural network image synthesis]]></category>
		<category><![CDATA[normalized cross-correlation]]></category>
		<category><![CDATA[SSIM]]></category>
		<category><![CDATA[VGG 16]]></category>
		<category><![CDATA[VGG 19]]></category>
		<category><![CDATA[VGG and Inception architectures]]></category>
		<category><![CDATA[Xception]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193818</guid>

					<description><![CDATA[Researchers combined five pretrained CNN architectures in a bagging ensemble to generate richer, more faithful Deep Dream images, achieving a loss of 8.5589 with SSIM of 0.2404 and NCC of 0.7622.]]></description>
										<content:encoded><![CDATA[<p>More than a decade after Google engineers first revealed the hypnotic, dreamlike images produced by a technique known as Deep Dream, a team of researchers from institutions across Iraq and the United Arab Emirates has found a way to make the algorithm&#8217;s hallucinations richer, more coherent and more faithful to the original picture. Their secret is not a brand-new network but a lesson machine learning has taught for years: sometimes a committee beats a single expert.</p>
<p>The study, published in Neural Computing and Applications, presents a bagging ensemble framework that combines the outputs of five pretrained convolutional neural network architectures: VGG 16, VGG 19, Xception, Inception v3 and Inception-ResNet-V2. Rather than asking one network to imagine the hidden patterns inside an image, the researchers let all five dream simultaneously and then aggregate their visions through iterative processes that enhance the features of the final image. The result, the team reports, is a Deep Dream output that maintains a high degree of similarity to the input while producing complex and aesthetically striking imagery.</p>
<p>Deep Dream works by exploiting the way convolutional neural networks perceive the world. These networks learn hierarchical features, with early layers detecting simple edges and textures and deeper layers responding to abstract concepts such as faces, eyes or animal shapes. When the technique amplifies the activations of chosen layers and feeds the result back through the network, the image spirals into the surreal patterns that made the method an internet sensation. The problem, however, is that a single architecture captures only one interpretation of the image&#8217;s features, and that narrow viewpoint can limit the quality and diversity of the generated dream.</p>
<p>The research team addressed this limitation by treating each pretrained network as a distinct dreamer. In each of the five CNN architectures, the model selects specific layers to activate, including some frozen layers, a choice that enhances the process of feature extraction. Because each network was trained on large image datasets and developed its own internal vocabulary of visual features, the ensemble effectively harvests a wider spectrum of patterns than any single model could reveal. The outputs are aggregated so that shared structures reinforce one another while idiosyncratic artifacts tend to fade, and a final scaling step refines the output image.</p>
<p>Bagging, short for bootstrap aggregating, is a classical ensemble strategy usually applied to prediction tasks, where multiple models trained on slightly different samples of data vote on an outcome. Applying it to image generation is less common, and it is here that the study makes its mark. Instead of sampling the data, the method samples the models themselves, drawing on architectures with different depths, block structures and activation behaviors. VGG 16 and VGG 19, with their stacks of small convolutional filters, emphasize texture-like patterns; the Inception family, with its multi-scale modules, blends information across several receptive field sizes; Xception&#8217;s depthwise separable convolutions decompose spatial and channel-wise filtering in yet another way. Each contributes a different stylistic signature to the composite dream.</p>
<p>To evaluate the approach objectively, the researchers tested their generated images using three metrics: loss, the Structural Similarity Index Measure (SSIM), and Normalized Cross-Correlation (NCC). SSIM quantifies how closely the structure of the generated image matches the original, capturing perceived changes in luminance, contrast and structure, while NCC measures the linear correlation between the two images. The generated images exhibited a loss value of 8.5589, with SSIM and NCC values of 0.2404 and 0.7622 respectively. The authors interpret these values as evidence that the ensemble method keeps the generated dreams anchored to the source image even as it introduces elaborate new patterns.</p>
<p>The numbers also hint at the fundamental tension at the heart of Deep Dream. An SSIM of 0.2404 indicates that the generated image departs substantially from its source in structural terms, which is precisely the point of the technique: the goal is transformation, not replication. The comparatively high NCC of 0.7622 shows that the two images remain strongly correlated in their overall signal, suggesting the algorithm amplifies what is already latent in the picture rather than inventing unrelated content. Striking that balance, the researchers argue, is where the bagging ensemble proves its robustness.</p>
<p>The work builds on the authors&#8217; earlier explorations of hybrid artistic models that combined Deep Dream with multiple CNN architectures, extending the idea into the ensemble learning framework that has proven so effective elsewhere in machine learning. It also joins a growing body of research that treats generative visual systems not merely as technical curiosities but as instruments of creative practice, from simulated visual hallucination studies in virtual reality to biometric and diagnostic applications that borrow Deep Dream&#8217;s feature-amplification machinery.</p>
<p>For artists and designers, the findings suggest a practical path forward: rather than settling for the idiosyncrasies of one pretrained model, creators can blend the perceptual biases of several networks to steer the mood and texture of their generated work. Because all five architectures are publicly available and pretrained, the approach does not require expensive training runs, making ensemble dreaming accessible to studios and hobbyists alike. The integration of CNN variants in a bagging ensemble framework, the authors conclude, paves the way for further innovations at the intersection of artificial intelligence and artistic expression, confirming the role of ensemble learning in supporting and enhancing visual art.</p>
<p>Beyond the gallery, the technique&#8217;s underlying principle, aggregating diverse feature extractors to stabilize and enrich a generative process, could inform other domains where neural networks are asked to reinterpret images, including data augmentation, style transfer and synthetic data generation for training vision systems. As generative AI continues to reshape how images are made and consumed, studies like this one remind us that some of the most compelling machine creativity comes not from bigger models but from teaching machines, like human artists, to compare notes.</p>
<p>The choice of architectures in the study is not arbitrary. All five networks were originally designed for large-scale image classification challenges, and each has since become a standard backbone in computer vision research. VGG-style networks, introduced in the mid-2010s, remain popular despite their age because their uniform structure of small convolutional filters produces feature maps that are easy to interpret and manipulate. The Inception lineage introduced the idea of running convolutions at several filter sizes in parallel within a single module, while Inception-ResNet-V2 added residual connections that ease the training of very deep stacks. Xception pushed this logic further by replacing standard filters with depthwise separable convolutions, a design that later influenced efficient mobile architectures. Drawing dream content from networks with such different inductive biases means the ensemble samples a genuinely diverse space of learned visual features.</p>
<p>A further practical advantage lies in the use of frozen layers. Because the networks are kept in their pretrained state, no gradient-based retraining is required, and the computational burden is limited to the forward and backward passes needed to amplify activations. This makes the method reproducible on modest hardware, an important consideration for creative practitioners who may not have access to large computing clusters. It also means the stylistic character of each network reflects the statistics of the dataset it originally learned from, so the ensemble implicitly blends the visual priors of multiple training runs without any additional data collection.</p>
<p>The evaluation strategy also reflects broader debates in generative image assessment. Metrics such as the Inception Score and the Fréchet Inception Distance are widely used for generative adversarial networks, but they measure distribution-level quality and diversity rather than fidelity to a specific source image. For a technique like Deep Dream, where the output is meant to be a transformation of a particular input, pairwise measures such as SSIM and NCC are more informative. SSIM, originally developed as a perceptual similarity measure grounded in the human visual system, has itself been the subject of ongoing refinement, with data-driven variants proposed to correct its known biases. Reporting both a structural similarity score and a correlation coefficient alongside a loss value therefore gives a more rounded picture of how the ensemble balances novelty against coherence than any single number could.</p>
<p>The work also sits within a lineage of research that repurposes feature-amplification machinery for tasks far removed from art. Deep Dream-style inversion has been used for data-free knowledge transfer between networks, and related mechanisms have appeared in cancellable biometric schemes, authentication systems, and diagnostic imaging pipelines, where amplifying latent features can highlight patterns that are otherwise subtle. In agriculture, image-to-image translation built on deep dreaming has been explored for crop disease datasets, suggesting that the same core operation can serve both aesthetic and analytical ends. The present study&#8217;s contribution to this lineage is architectural: it demonstrates that the aggregation principle, so successful in classification and prediction, transfers cleanly to a generative setting where there is no single correct output to converge upon.</p>
<p>There remain open questions that future work could address. The study relies on a fixed selection of layers within each network, and earlier investigations by overlapping author groups have shown that changing the targeted layers in a single model substantially alters the character of the resulting dream. A systematic exploration of how layer choice interacts with ensemble aggregation, or of how the number of contributing architectures affects the trade-off between richness and fidelity, would deepen the understanding of why the committee approach works. Extending the evaluation beyond pairwise similarity metrics to include human aesthetic judgments or distributional measures could also clarify how the perceived artistic quality of the composite images relates to their measured statistical properties, a question that ultimately lies at the boundary between machine learning and the psychology of visual perception.</p>
<p><strong>Subject of Research:</strong> Generating Deep Dream images using bagging ensemble learning across multiple pretrained convolutional neural network architectures</p>
<p><strong>Article Title:</strong> Generating deep dream images via bagging ensemble learning across multiple pretrained architectures</p>
<p><strong>Article References:</strong> Ali, L. R., Alkhazraji, W., Kadhim, Z. S., Jebur, S. A., Abbas, A. R., Jamil, A. S., Hussein, Z. K., Jaber, T. A., Hussein, H. A., Shaker, B. N., &amp; Hussain, A. J. (2026). Generating deep dream images via bagging ensemble learning across multiple pretrained architectures. <em>Neural Computing and Applications, 38</em>(17), Article 730. <a href="https://doi.org/10.1007/s00521-026-12432-1" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12432-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12432-1" rel="noopener noreferrer">10.1007/s00521-026-12432-1</a></p>
<p><strong>Keywords:</strong> Deep Dream, ensemble learning, bagging, convolutional neural networks, VGG 16, VGG 19, Xception, Inception v3, Inception-ResNet-V2, SSIM, normalized cross-correlation, AI art</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193818</post-id>	</item>
	</channel>
</rss>
