<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>comparison of convolutional neural networks and Vision Transformers &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/comparison-of-convolutional-neural-networks-and-vision-transformers/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 17:07:04 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>comparison of convolutional neural networks and Vision Transformers &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>CNN or Transformer? AI Duel Reveals Best Way to Read X-ray Patterns</title>
		<link>https://scienmag.com/cnn-or-transformer-ai-duel-reveals-best-way-to-read-x-ray-patterns/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 17:07:04 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced AI techniques for X-ray pattern interpretation]]></category>
		<category><![CDATA[AI applications in non-destructive testing]]></category>
		<category><![CDATA[AI-driven mineral phase detection]]></category>
		<category><![CDATA[CNN vs. Transformer in material science]]></category>
		<category><![CDATA[comparison of convolutional neural networks and Vision Transformers]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning architectures for X-ray pattern recognition]]></category>
		<category><![CDATA[deep learning for mineralogical phase quantification]]></category>
		<category><![CDATA[machine learning in waste material recycling]]></category>
		<category><![CDATA[mineral identification using AI]]></category>
		<category><![CDATA[mineral phase estimation in industrial residues]]></category>
		<category><![CDATA[mineral residues]]></category>
		<category><![CDATA[multi-task learning]]></category>
		<category><![CDATA[quantitative mineral analysis with neural networks]]></category>
		<category><![CDATA[quantitative phase analysis]]></category>
		<category><![CDATA[recycling]]></category>
		<category><![CDATA[Rietveld refinement]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[simulation-to-real]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[Vision Transformers]]></category>
		<category><![CDATA[X-ray diffraction]]></category>
		<category><![CDATA[X-ray diffraction analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248649</guid>

					<description><![CDATA[A controlled head-to-head study shows Vision Transformers beat convolutional networks on real X-ray diffraction data once fine-tuned, while CNNs remain more robust for zero-shot transfer.]]></description>
										<content:encoded><![CDATA[<p>Every year, the world&#8217;s factories and incinerators bury mountains of value in plain sight. Ashes from municipal waste plants, slags from metallurgical furnaces, and shredded residues from dead batteries and old electronics all contain minerals that could be recovered, reused, or safely locked away—but only if engineers know exactly what crystalline phases lurk inside them. A new study published in Results in Engineering by Yi Zhang and colleagues now shows that the choice of artificial intelligence architecture can make or break that knowledge, in a head-to-head contest between two of deep learning&#8217;s most celebrated designs.</p>
<p>The team, based at the AI Production Network Augsburg in Germany, set out to answer a deceptively simple question: when a machine reads an X-ray diffraction pattern, is a convolutional neural network or a Vision Transformer better at working out both which minerals are present and how much of each one the sample contains? Until now, no one had compared the two architectures under truly identical conditions on this task, because earlier studies had split the problem into separate stages—first identifying phases, then estimating their fractions—which made fair comparison impossible.</p>
<p>X-ray diffraction is the gold standard for quantitative phase analysis. When X-rays strike a finely ground powder, the crystal planes inside each mineral scatter the beam into a fingerprint of peaks at characteristic angles. The trouble is that real industrial residues produce nightmarish spectra: quartz and cristobalite, two forms of silicon dioxide, scatter at overlapping angles; cuprite and tenorite, both copper oxides, share diffraction features; and zirconia contaminating samples from grinding media produces a dense thicket of peaks that obscures everything beneath it. The traditional answer, Rietveld refinement, works well but demands expert users, careful initial parameters, and hours of computation for every sample—a bottleneck that simply cannot scale to the torrent of waste streams requiring characterization.</p>
<p>Deep learning promises to collapse that bottleneck, but the field has been divided over how to feed diffraction data to a neural network. Some researchers convert the one-dimensional intensity curve into a two-dimensional image and apply image-recognition networks, borrowing mature architectures from computer vision. That approach, however, discards spectral resolution through downsampling and can smear away the fine peak structure that carries the quantitative information. Zhang&#8217;s team instead treated the diffraction pattern as what it truly is: a one-dimensional sequence of intensities over diffraction angle, preserving its native structure.</p>
<p>Both contenders in the study shared the same multi-task skeleton. A feature extractor digests the raw spectrum into a compact embedding, which then feeds two parallel heads: a classification head that predicts whether each of seven target phases is present, and a regression head that estimates the weight fraction of every phase, constrained to sum to one. Crucially, the regression output is masked by the classification predictions, so phases judged absent automatically receive zero fraction. The models trained end-to-end on a composite loss combining binary cross-entropy for classification with mean squared error for regression. The convolutional network stacked three one-dimensional convolutional layers with shrinking kernel sizes to capture local peak shapes, while the Vision Transformer chopped the spectrum into patches and applied self-attention, letting every part of the spectrum communicate with every other part regardless of distance.</p>
<p>Because real labeled measurements are scarce and expensive, the researchers built a hybrid data strategy. They simulated tens of thousands of X-ray patterns from crystallographic data for seven phases—cristobalite, cuprite, graphite, copper, quartz, tenorite, and zirconia—using pseudo-Voigt peak profiles, exponentially decaying backgrounds, amorphous humps, and phase-specific noise calibrated so that each phase became measurably harder to detect, mimicking real laboratory degradation. The noise parameters were not guessed; they were optimized so that a template-based classifier&#8217;s detectability threshold degraded by roughly five percentage points per phase, a rigorous way of ensuring the simulated data posed genuine difficulty. On top of this synthetic mountain of 50,000 patterns sat a deliberately small set of just 105 real laboratory measurements of four reference phases, reflecting the data-poor reality of industrial practice.</p>
<p>The results, obtained under identical training data, loss functions, hyperparameter tuning, and evaluation metrics, tell a nuanced story. On purely simulated data, both architectures were nearly flawless at classification, with accuracy above 0.999, but the CNN reached its best regression performance early and plateaued, while the Transformer kept improving as the dataset grew, only converging with the CNN at the full 50,000 samples. The real drama unfolded on laboratory measurements. There, the Vision Transformer delivered perfect phase classification across all four reference phases and cut the regression error dramatically—reducing root mean square error by 31 percent and mean absolute error by 35 percent compared with the CNN. The lone exception was quartz, where the convolutional network held its ground, likely because quartz&#8217;s sharp, well-resolved peaks reward the local, receptive-field approach of convolutions.</p>
<p>The transfer experiments added another twist. When models trained only on simulated data were applied directly to real measurements without any adaptation—a zero-shot test—the CNN actually transferred more robustly than the Transformer. But a modest amount of fine-tuning on real samples largely closed the simulation-to-real gap for both architectures, and the Transformer benefited most from every additional real sample, achieving its best results when 72 real measurements were used for adaptation. Visualizations of the models&#8217; internal feature spaces using UMAP projections revealed why: the Transformer&#8217;s embeddings formed compact, well-separated clusters in which simulated and real spectra of the same phase already sat close together before any fine-tuning, whereas the CNN&#8217;s feature space remained diffuse and only partially realigned after adaptation. The convolutional network&#8217;s predictions, however, proved more stable across repeated random splits of the small real dataset, suggesting its strong built-in assumptions about local peak patterns act as a safeguard when data is scarce.</p>
<p>For practitioners, the message is refreshingly practical rather than doctrinaire. If labeled real measurements are available for fine-tuning, the Vision Transformer is the stronger choice, especially for materials dominated by broad, overlapping, or poorly crystalline features such as graphite&#8217;s smeared reflections. If the goal is zero-shot deployment on new instruments or new sample types with no adaptation data, the humble CNN remains a formidable competitor. Architecture selection, in other words, should follow the data budget and the spectral character of the target phases—not fashion.</p>
<p>The broader implications reach well beyond waste recycling. Quantitative phase analysis underpins everything from cement quality control to pharmaceutical polymorph screening to planetary geology, and any technique that removes the expert bottleneck from Rietveld refinement multiplies the throughput of entire laboratories. The Augsburg team is candid about the limits of the current work: only four of the seven simulated phases received experimental validation, the real dataset contains just 105 patterns, and the framework assumes a closed set of phases that may not hold in open-ended industrial screening. Future work will expand the real measurements, incorporate explicit domain adaptation techniques such as adversarial training, and add uncertainty quantification to flag low-confidence predictions. But the core finding stands as a milestone in applied machine learning: in the contest between convolution and attention, the winner depends on where you stand—and now, for the first time, scientists know exactly where that line is drawn.</p>
<p><strong>Subject of Research:</strong> Deep learning architectures for quantitative phase analysis of X-ray diffraction patterns from industrial mineral residues</p>
<p><strong>Article Title:</strong> A comparative study of convolutional neural networks and vision transformers for quantitative phase analysis from X-ray diffraction</p>
<p><strong>Article References:</strong> Zhang, Y., Eulitz, S., Michalak, A., Vollprecht, D., &amp; Mikelsons, L. (2026). A comparative study of convolutional neural networks and vision transformers for quantitative phase analysis from X-ray diffraction. <em>Results in Engineering, 32</em>, Article 113313. <a href="https://doi.org/10.1016/j.rineng.2026.113313" rel="noopener noreferrer">https://doi.org/10.1016/j.rineng.2026.113313</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.rineng.2026.113313" rel="noopener noreferrer">10.1016/j.rineng.2026.113313</a></p>
<p><strong>Keywords:</strong> X-ray diffraction, quantitative phase analysis, deep learning, convolutional neural networks, vision transformers, mineral residues, recycling, Rietveld refinement, transfer learning, simulation-to-real, multi-task learning, self-attention</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248649</post-id>	</item>
	</channel>
</rss>
