<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>t-SNE &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/t-sne/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 19:23:56 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>t-SNE &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Framework Maps the Shifting Landscape of Illegal Fundraising Schemes</title>
		<link>https://scienmag.com/ai-framework-maps-the-shifting-landscape-of-illegal-fundraising-schemes/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 19:23:56 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI and human judgment in financial monitoring]]></category>
		<category><![CDATA[AI-driven financial risk analysis]]></category>
		<category><![CDATA[Analytic Hierarchy Process]]></category>
		<category><![CDATA[big data in financial regulation]]></category>
		<category><![CDATA[cross-border financial fraud]]></category>
		<category><![CDATA[digital financial regulation]]></category>
		<category><![CDATA[expert-in-the-loop learning]]></category>
		<category><![CDATA[FATF]]></category>
		<category><![CDATA[financial crime]]></category>
		<category><![CDATA[financial crime taxonomy development]]></category>
		<category><![CDATA[FinBERT]]></category>
		<category><![CDATA[illegal fundraising]]></category>
		<category><![CDATA[Illegal fundraising schemes]]></category>
		<category><![CDATA[interdisciplinary collaboration in financial AI]]></category>
		<category><![CDATA[interdisciplinary financial regulation tools]]></category>
		<category><![CDATA[interpretability of AI frameworks in finance]]></category>
		<category><![CDATA[LDA]]></category>
		<category><![CDATA[online investment scams detection]]></category>
		<category><![CDATA[real case data in financial risk management]]></category>
		<category><![CDATA[risk taxonomy]]></category>
		<category><![CDATA[t-SNE]]></category>
		<category><![CDATA[text analytics]]></category>
		<category><![CDATA[TF-IDF]]></category>
		<category><![CDATA[topic modeling]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=231606</guid>

					<description><![CDATA[A hybrid topic-modeling and expert-calibration framework called TED-FinRisk turns more than 1,000 illegal fundraising case texts into an interpretable six-category risk taxonomy for regulators.]]></description>
										<content:encoded><![CDATA[<p>Illegal fundraising has become one of the most slippery targets in financial regulation. Schemes no longer announce themselves through a single channel or a familiar sales pitch; they spread across dispersed digital platforms, cross regional borders, and mutate their promotional narratives faster than official rule lists can be revised. Elder care investment traps, virtual currency offerings, supply-chain finance vehicles, agricultural ventures, film investment products, social e-commerce schemes, and overseas investment pitches all compete for the attention of regulators who must somehow sort them into coherent categories before they can be monitored, warned about, and prosecuted. A new study published in the Journal of Big Data argues that the answer lies not in replacing human judgment with machines, but in fusing the two into a single, interpretable pipeline.</p>
<p>The framework, named TED-FinRisk, was developed by Wenchuan Kuang and Yusiyuan Chen of Fudan University&#8217;s College of Computer Science and Artificial Intelligence, together with colleagues from Gientech Technology&#8217;s Financial Risk Lab, Mashang Consumer Finance&#8217;s AI Research Lab, and Pennsylvania State University. Its central claim is methodological: a risk taxonomy built from real case texts can carry genuine regulatory semantics if data-driven topic discovery is deliberately calibrated by domain experts rather than left to operate unsupervised. The researchers tested this idea on a corpus of more than 1,000 illegal fundraising cases and nearly 6,000 associated risk-related records, a body of material rich enough to capture the operational texture of contemporary schemes.</p>
<p>Technically, TED-FinRisk is a hybrid in the strict sense, chaining several well-established text analytics techniques into a sequence in which each stage constrains the next. The pipeline begins with term frequency–inverse document frequency, or TF-IDF, feature construction. TF-IDF assigns weights to words based on how often they appear in a given case document and how rare they are across the whole corpus, so that boilerplate legal language fades into the background while distinctive vocabulary, the telltale phrases of a Ponzi pitch or a crypto yield promise, rises to the surface. These weighted vectors become the raw material for everything that follows.</p>
<p>Before any topics are formally extracted, the researchers employ t-distributed stochastic neighbor embedding, or t-SNE, to assist topic exploration. t-SNE is a nonlinear dimensionality reduction technique that projects high-dimensional document vectors into a low-dimensional map while preserving local neighborhoods, allowing analysts to see clusters of similar cases at a glance. In TED-FinRisk this step serves as a visual diagnostic: it helps the team gauge how many natural groupings exist in the corpus and whether candidate topic counts are plausible before committing to a formal model. It is a pragmatic use of visualization as a sanity check rather than an end in itself.</p>
<p>The formal topic extraction is performed with latent Dirichlet allocation, or LDA, a generative probabilistic model that treats each document as a mixture of hidden topics and each topic as a probability distribution over words. Running LDA over the TF-IDF-weighted case texts yields candidate topics, each characterized by its most probable terms, which the researchers can then read as proto-categories of illegal fundraising behavior. But LDA output alone is notoriously noisy and does not automatically align with the categories regulators actually use. This is where the framework&#8217;s distinguishing move comes in: expert calibration through the Analytic Hierarchy Process, or AHP.</p>
<p>AHP is a structured decision-making method in which experts compare criteria pairwise and derive consistent priority weights. In TED-FinRisk, the AHP stage sits inside a threat–vulnerability–consequence perspective borrowed from the Financial Action Task Force, the intergovernmental body that sets global anti-money-laundering standards. Experts evaluate the data-driven topics against FATF-style considerations of threat, vulnerability, and consequence, assigning priorities that determine how the candidate topics are merged, split, renamed, and organized into a final taxonomy. The result is a classification scheme with six primary categories and a set of secondary labels, each linked to operational indicators that can feed downstream monitoring systems.</p>
<p>The scope of the corpus gives the taxonomy its breadth. The cases span scenarios ranging from elder care schemes that exploit aging populations&#8217; retirement savings, to virtual currency offerings riding waves of speculative enthusiasm, to supply-chain finance arrangements and agricultural ventures that wrap illicit fundraising in the language of legitimate business. Film investment products, social e-commerce schemes, and overseas investment vehicles round out the collection. This diversity matters because the study&#8217;s core argument is that a taxonomy derived from a narrow slice of cases will fail the moment schemes evolve beyond it, whereas a taxonomy grounded in thousands of records across many domains has a better chance of anticipating the next mutation.</p>
<p>Crucially, the authors do not stop at building the taxonomy; they test whether it corresponds to something real in the text. Once the expert-finalized labels are in place, the team runs a reclassification consistency check using FinBERT, a transformer language model pre-trained on financial text. FinBERT is tasked with assigning case documents to the taxonomy&#8217;s categories, and the researchers examine whether the model&#8217;s assignments align with the expert labels. If a taxonomy category cannot be recovered from the raw text by a semantic model, that category may be an artifact of expert convention rather than a pattern with genuine textual signature. The consistency check thus functions as an empirical bridge between human-defined regulatory semantics and machine-detectable language patterns, a form of expert-in-the-loop validation that the authors present as a reusable design principle.</p>
<p>The paper is explicit about its own character: it is a methodology-oriented case study, not a deployed production system. Its contributions are framed as a reusable taxonomy design process, a structured basis for future benchmarking, and a foundation for early-warning applications. That framing is honest about the state of the field. Static rule lists and purely expert-driven typologies, the authors note in their abstract, have become difficult to maintain as illegal fundraising migrates to dispersed digital channels and rapidly changing narratives. What they propose instead is a repeatable pipeline in which topic discovery, expert prioritization, and semantic validation reinforce one another, so that the taxonomy can be refreshed as new case texts arrive rather than rewritten from scratch.</p>
<p>For regulators and financial institutions, the significance of TED-FinRisk lies less in any single algorithm than in the architecture of collaboration it demonstrates. TF-IDF, t-SNE, LDA, AHP, and FinBERT are all established tools; the innovation is the disciplined order in which they are combined and the insistence that neither the data nor the experts have the final word alone. The six-category taxonomy with its operational indicators offers a template for how machine learning and regulatory practice can meet in the middle, producing classifications that are simultaneously discoverable in the data, meaningful to compliance officers, and testable by independent models. As financial crime continues to evolve at the speed of digital marketing, frameworks of this hybrid kind may become essential infrastructure for keeping watchlists, warning systems, and enforcement priorities a step ahead of the schemes they are designed to catch. The study, published open access on 3 October 2026 and supported by China&#8217;s National Key Research and Development Program, the National Natural Science Foundation of China, and regional science projects, invites other jurisdictions to replicate the process on their own case corpora and compare the resulting taxonomies, turning what has long been an artisanal exercise in expert judgment into a benchmarkable, evolving science.</p>
<p><strong>Subject of Research:</strong> Machine learning-based risk taxonomy construction for typologizing illegal fundraising from multi-source case texts</p>
<p><strong>Article Title:</strong> TED-FinRisk: a hybrid topic-enhanced framework for typologizing illegal fundraising risk from multi-source case texts</p>
<p><strong>Article References:</strong> TED-FinRisk: a hybrid topic-enhanced framework for typologizing illegal fundraising risk from multi-source case texts. (n.d.). <a href="https://doi.org/10.1186/s40537-026-01574-7" rel="noopener noreferrer">https://doi.org/10.1186/s40537-026-01574-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40537-026-01574-7" rel="noopener noreferrer">10.1186/s40537-026-01574-7</a></p>
<p><strong>Keywords:</strong> illegal fundraising, risk taxonomy, topic modeling, LDA, FinBERT, TF-IDF, t-SNE, Analytic Hierarchy Process, financial crime, FATF, text analytics, expert-in-the-loop learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">231606</post-id>	</item>
		<item>
		<title>Quantum Kernels Tackle Overlapping Anomalies That Defeat Classical Machine Learning</title>
		<link>https://scienmag.com/quantum-kernels-tackle-overlapping-anomalies-that-defeat-classical-machine-learning/</link>
		
		<dc:creator><![CDATA[Katie Riggs]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 11:06:07 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[acoustic anomaly detection]]></category>
		<category><![CDATA[acoustic monitoring]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[anomaly detection in factory settings]]></category>
		<category><![CDATA[classical machine learning limitations]]></category>
		<category><![CDATA[entanglement]]></category>
		<category><![CDATA[false positive rate]]></category>
		<category><![CDATA[fault diagnosis in manufacturing]]></category>
		<category><![CDATA[feature space]]></category>
		<category><![CDATA[Hilbert space]]></category>
		<category><![CDATA[industrial anomaly detection]]></category>
		<category><![CDATA[industrial condition monitoring]]></category>
		<category><![CDATA[machine learning for predictive maintenance]]></category>
		<category><![CDATA[non-intrusive vibration monitoring]]></category>
		<category><![CDATA[One-Class SVM]]></category>
		<category><![CDATA[overlapping data distributions]]></category>
		<category><![CDATA[Quantum kernel methods]]></category>
		<category><![CDATA[quantum kernels]]></category>
		<category><![CDATA[Quantum machine learning]]></category>
		<category><![CDATA[quantum machine learning advantages]]></category>
		<category><![CDATA[quantum vs classical kernel comparison]]></category>
		<category><![CDATA[RBF kernel]]></category>
		<category><![CDATA[support vector machines]]></category>
		<category><![CDATA[t-SNE]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227367</guid>

					<description><![CDATA[A new study shows quantum kernel methods dramatically outperform classical RBF kernels in industrial acoustic anomaly detection when anomaly distributions overlap, cutting false positive rates from 60 percent to 3 percent.]]></description>
										<content:encoded><![CDATA[<p>Quantum machine learning has long been accused of solving problems nobody actually has. A new study in the journal Quantum Machine Intelligence pushes back against that criticism with something refreshingly concrete: factory-floor sounds. Takao Tomono of Keio University and the Japan Aerospace Exploration Agency, together with Kazuya Tsujimura of Toppan Holdings, systematically compared quantum kernel methods against the classical radial basis function (RBF) kernel on real acoustic anomaly detection tasks, and found a dramatic advantage precisely where classical methods are expected to struggle: when the distributions of normal and anomalous data overlap so heavily that no amount of tuning can pry them apart.</p>
<p>The stakes are far from academic. Industrial anomaly detection is a cornerstone of modern manufacturing, where catching the faint acoustic signature of a worn bearing or a misaligned belt before it fails can save enormous maintenance costs and prevent catastrophic downtime. Acoustic and vibration monitoring is especially attractive because it is non-intrusive and sensitive to subtle mechanical degradation. Yet classical pipelines built on support vector machines hit a wall when multiple anomaly types coexist and their feature distributions intermingle in Euclidean space. The result is an unacceptably high false positive rate, which in practice means operators learn to ignore the alarms, defeating the entire purpose of the monitoring system.</p>
<p>The researchers built their study around two physical test rigs with deliberately contrasting difficulty. The first, an Open Belt Drive (OBD) system, consists of rubber and metal chain belt drives in which wooden chopsticks are inserted through pre-drilled holes, producing sudden, loud breaking sounds that define the anomaly class. The second, a Mini 4WD (M4W) course, sends a miniature four-wheel drive vehicle around a three-lane track where it first strikes a one-millimeter wooden stick, generating a small impact sound, and then passes over Magic Tape, producing a faint scratching noise. While chopstick breaks are obvious even to human ears, the stick and tape sounds of the M4W rig are extremely difficult to discern, and the two anomaly types occur sequentially within each roughly five-second lap.</p>
<p>From five-minute recordings of each system, the team extracted 30 normal samples per dataset and converted the audio into compact feature vectors using autoregressive (AR) modeling. Each 10-second segment was resampled from 48 kHz to 1100 Hz, cut into non-overlapping 0.4-second windows, and characterized by AR coefficients estimated with the Yule-Walker method, with median aggregation across windows providing robustness against transient noise spikes. The AR model order was varied from 2 to 10, and information-theoretic criteria, the Akaike Information Criterion and Bayesian Information Criterion, independently identified orders 9 to 10 as optimal, validating the experimental range. Because real industrial equipment rarely fails often enough to supply labeled anomaly data, the team adopted an unsupervised One-Class SVM that learns a decision boundary from normal operation data alone, flagging significant deviations as anomalies.</p>
<p>Against this pipeline, three kernels competed: the classical RBF baseline, and two quantum kernel architectures. Quantum Kernel 1 (QK1) uses linear entanglement, with nearest-neighbor CNOT gates connecting adjacent qubits in a chain, creating pairwise correlations between consecutive features at a circuit depth that scales linearly with qubit count. Quantum Kernel 2 (QK2) employs all-to-all entanglement, connecting every qubit pair to capture global, higher-order dependencies at the cost of quadratic circuit depth. Both use a data re-uploading strategy in which AR coefficients are repeatedly encoded into rotation gates across two layers, producing quantum amplitudes that depend non-linearly on the input through interference. Importantly, the quantum kernels were implemented via classical simulation, using adjusted RBF gamma parameters that approximate the correlation structures of the two entanglement topologies, a choice the authors candidly flag as a limitation that leaves confirmation on actual quantum hardware to future work.</p>
<p>On the OBD dataset, where anomalies are cleanly separable, all three methods eventually achieved perfect classification, but the quantum kernels got there faster. QK1 reached a perfect F1 score of 1.0 with just four AR features, QK2 matched it at seven, while the RBF kernel needed eight. Wilcoxon signed-rank tests across the full feature range confirmed these differences were statistically significant, with p-values of 0.0420 for QK1 versus RBF and 0.0203 for QK2 versus RBF. In practical terms, the quantum kernels offered a computational efficiency advantage, roughly halving the feature extraction overhead, but no fundamental capability gap, since classical methods converged to the same asymptotic solution.</p>
<p>The M4W dataset told an entirely different story. There, QK2 achieved an F1 score of 0.893 with a false positive rate of just 0.033, while the RBF kernel collapsed to an F1 of 0.428 with a staggering false positive rate of 0.600, meaning 60 percent of normal operation was flagged as anomalous. The statistical significance was emphatic, with a p-value of 0.0023. QK1 landed in between at F1 = 0.661 but became unstable at higher feature orders. Most tellingly, the RBF kernel&#8217;s F1 trajectory remained flat, never exceeding 0.47 despite a fivefold increase in feature count, a signature of structural failure rather than insufficient data. ROC trajectory analysis made the mechanism visible: RBF wandered horizontally in ROC space, oscillating between false positive rates of 0.6 and 1.0 without ever approaching the ideal operating point, while QK2 systematically drove its false positive rate down from 0.93 to 0.03 as features were added.</p>
<p>To explain why, the researchers turned to t-SNE visualization of the AR feature space, backed by rigorous statistical testing. For OBD, Kruskal-Wallis tests confirmed complete class separation in both projected dimensions, with all pairwise Mann-Whitney comparisons passing Bonferroni-corrected significance thresholds. For M4W, the picture was one of fundamental inseparability: the second dimension showed no significant structure at all (H = 0.4, p = 0.942), and Magic Tape anomalies proved statistically indistinguishable from every other class. A striking paradox emerged in the decision score distributions: QK2 achieved its superior performance despite a centroid separation between normal and anomalous classes roughly three times smaller than RBF&#8217;s. The quantum kernel&#8217;s advantage lies not in pushing class means apart but in reshaping the distributions to minimize their overlap, exploiting non-linear correlations through high-dimensional Hilbert space projections that Euclidean geometry simply cannot represent.</p>
<p>The study also connects its unsupervised findings to earlier supervised work by the same group, including apple defect detection in shipping inspection that achieved F1 scores of 0.8 to 0.9 with only 24 training samples per class, far outperforming classical RBF. Across both paradigms, a consistent pattern emerges: quantum kernel benefits appear in small-data regimes where classical feature distributions overlap, while showing diminishing returns beyond moderate feature dimensionality, a phenomenon attributed to exponential concentration that may impose fundamental scalability limits. The authors propose a practical screening framework: run a preliminary t-SNE analysis with classical features, and if Kruskal-Wallis tests reveal poor separation or multiple inseparable class pairs, quantum kernels may be worth the investment. With false alarms triggering unnecessary inspections and production stoppages, a drop from a 60 percent to a 3 percent false positive rate is not a marginal improvement but the difference between a monitoring system operators trust and one they switch off. As quantum hardware matures, this work suggests that messy, overlapping, real-world sensor data, not pristine synthetic benchmarks, may be where quantum machine learning first earns its keep.</p>
<p><strong>Subject of Research:</strong> Quantum kernel methods for unsupervised industrial acoustic anomaly detection with overlapping data distributions</p>
<p><strong>Article Title:</strong> Potential of multi-anomalies detection using quantum machine learning</p>
<p><strong>Article References:</strong> Tomono, T., &amp; Tsujimura, K. (2026). Potential of multi-anomalies detection using quantum machine learning. <em>Quantum Machine Intelligence, 8</em>(2), Article 106. <a href="https://doi.org/10.1007/s42484-026-00448-8" rel="noopener noreferrer">https://doi.org/10.1007/s42484-026-00448-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s42484-026-00448-8" rel="noopener noreferrer">10.1007/s42484-026-00448-8</a></p>
<p><strong>Keywords:</strong> quantum machine learning, quantum kernels, anomaly detection, One-Class SVM, acoustic monitoring, false positive rate, t-SNE, feature space, industrial condition monitoring, RBF kernel, entanglement, Hilbert space</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227367</post-id>	</item>
		<item>
		<title>AI Faces a Hard Test: Detecting Rare Martensite in Steel Micrographs</title>
		<link>https://scienmag.com/ai-faces-a-hard-test-detecting-rare-martensite-in-steel-micrographs/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:03:07 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[application of deep learning in steel microstructure analysis]]></category>
		<category><![CDATA[challenges of AI in detecting rare phases]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[data imbalance impact on microstructure detection]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for micrograph analysis]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[GLCM]]></category>
		<category><![CDATA[Grad-CAM]]></category>
		<category><![CDATA[limitations of AI in materials quality assurance]]></category>
		<category><![CDATA[machine learning class imbalance in metallography]]></category>
		<category><![CDATA[Martensite]]></category>
		<category><![CDATA[Martensite identification in steel micrographs]]></category>
		<category><![CDATA[microstructural phase classification in high carbon steel]]></category>
		<category><![CDATA[microstructure classification]]></category>
		<category><![CDATA[rare microstructure detection in steel]]></category>
		<category><![CDATA[rare-phase detection]]></category>
		<category><![CDATA[ResNet50]]></category>
		<category><![CDATA[ResNet50 application in metallurgical microstructure classification]]></category>
		<category><![CDATA[t-SNE]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[transfer learning in materials science]]></category>
		<category><![CDATA[ultra-high carbon steel]]></category>
		<category><![CDATA[use of SEM micrographs for AI training]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202416</guid>

					<description><![CDATA[A new baseline study shows that deep learning can learn meaningful microstructural information from severely imbalanced ultra-high carbon steel images, but reliable detection of the rare Martensite phase remains constrained by data scarcity and class overlap.]]></description>
										<content:encoded><![CDATA[<p>Deep learning has transformed how scientists classify materials, but a new study shows just how difficult the task becomes when the most important feature is also the rarest. In research published in the Journal of Materials Science: Metallurgy, Ehsun Saeed evaluated whether a transfer-learning framework could reliably detect Martensite, a hard and brittle phase, in ultra-high carbon steel micrographs when that phase accounted for a mere 3.75 percent of the dataset. The work was conceived as a baseline diagnostic rather than a claim of industrial readiness, and its findings carry a candid message about the limits of artificial intelligence in metallurgical quality assurance.</p>
<p>The dataset consisted of 961 scanning electron microscopy micrographs spanning seven microstructural classes, drawn from the Ultra-High Carbon Steel micrograph collection maintained by the National Institute of Standards and Technology. Spheroidite dominated with 374 images, while Martensite was represented by only 36. This tenfold disparity in representation creates exactly the kind of class imbalance that undermines conventional machine learning, where optimisation naturally favours majority classes and can produce seemingly respectable accuracy while missing the minority class entirely.</p>
<p>Saeed adopted a frozen ResNet50 transfer-learning architecture, using ImageNet pre-trained weights as a fixed feature extractor and retraining only the newly added classification layers. Images were resized to 224 by 224 pixels, normalised, and augmented online with random rotations, flips, and shifts during training. A stratified train-test split of 80 to 20 preserved class proportions, leaving just seven Martensite images in the test partition. Baseline training used the Adam optimiser with categorical cross-entropy loss for 30 epochs, and two reference classifiers, a random predictor and a majority-class predictor, served as lower bounds.</p>
<p>The baseline model achieved a peak validation accuracy of 52.8 percent, well above the random baseline of 14.3 percent and the majority-class baseline of 38.9 percent. Martensite detection, however, remained modest, with precision of 0.600, recall of 0.429, and an F1-score of 0.500. Because overall accuracy can mask poor minority-class recognition, the study also reported balanced accuracy, macro-averaged metrics, and the Matthews Correlation Coefficient, all of which provide a more honest picture under severe imbalance.</p>
<p>To probe robustness, stratified five-fold cross-validation was performed, yielding an average accuracy of 0.412, a macro F1-score of 0.133 plus or minus 0.009, and an MCC of 0.134 plus or minus 0.029. These figures reveal weak multi-class generalisation and suggest that the principal challenge is not instability across partitions but genuine class overlap and insufficient minority-class representation. The gap between the single split and the cross-validated results further indicates that performance estimates are sensitive to dataset composition.</p>
<p>Two imbalance-aware strategies were then tested. Class-weighted loss, with weights inversely proportional to class frequencies, boosted Martensite recall dramatically to 0.857 but crashed overall accuracy to 29.0 percent, illustrating a flood of false positives and an overcompensation for the rarity of the minority class. Focal loss, which down-weights easily classified majority samples, fared better, delivering the strongest overall performance with 56.0 percent accuracy, 36.5 percent balanced accuracy, an MCC of 0.388, and improved Martensite precision of 0.750 while holding recall steady. The results show that loss-function design can shift the balance between sensitivity and precision, but cannot conjure data that does not exist.</p>
<p>A binary Martensite-versus-non-Martensite experiment offered further insight, achieving an AUROC of 0.908 and an MCC of 0.527, substantially stronger than the multi-class formulation. Threshold optimisation raised recall from 28.6 percent to 57.1 percent while preserving precision at 80.0 percent. This suggests that ambiguity among seven overlapping classes compounds the difficulty, and that application-specific threshold tuning can extract meaningful rare-phase discrimination from the same underlying model.</p>
<p>Interpretability and confounding analyses rounded out the study. Grad-CAM visualisations showed that activation maps concentrated on microstructural regions rather than on scale bars or image borders, and border-crop robustness tests confirmed that removing up to 15 percent of the image perimeter left performance largely unchanged. t-SNE projection of the penultimate-layer features, however, revealed that Martensite samples were dispersed throughout the feature space with no distinct cluster, overlapping heavily with Pearlite, Spheroidite, and Network microstructures. A magnification audit exposed substantial variation in imaging scale across classes, from a mean of 89 times for Network to 13,441 times for Pearlite, and magnification-only classification with a Random Forest achieved an AUROC of 0.810, indicating that imaging scale acts as a partial confounder without fully explaining model behaviour.</p>
<p>Handcrafted texture descriptors, long the workhorses of metallographic image analysis, were evaluated for comparison. Features derived from Gray-Level Co-occurrence Matrices and Local Binary Patterns, classified with Support Vector Machine and Random Forest models under stratified cross-validation, achieved performance comparable to the CNN, suggesting that informative texture signals exist in the data and that physically interpretable baselines remain valuable benchmarks in data-scarce settings.</p>
<p>The study positions itself as a reproducible baseline rather than a deployable solution, and the implications are clear. Reliable rare-phase detection in industrial metallography will require larger minority-class datasets, scale-normalised imaging protocols, quantitative microstructural descriptors fused with learned features, and specialised strategies such as few-shot learning, anomaly detection, and metric learning. Until then, the author argues, such models are best regarded as decision-support tools subject to expert review, uncertainty-aware thresholds, and external validation across instruments, compositions, and processing conditions, rather than as autonomous replacements for the trained metallurgist&#8217;s eye.</p>
<p><strong>Subject of Research:</strong> Rare-phase detection of Martensite in ultra-high carbon steel micrographs using deep learning and handcrafted texture features under severe class imbalance</p>
<p><strong>Article Title:</strong> A baseline study of rare-phase detection in ultra-high carbon steel microstructures using deep learning and handcrafted texture features</p>
<p><strong>Article References:</strong> Saeed, E. (2026). A baseline study of rare-phase detection in ultra-high carbon steel microstructures using deep learning and handcrafted texture features. <em>Journal of Materials Science: Metallurgy, 1</em>(1), Article 12. <a href="https://doi.org/10.1007/s44492-026-00012-2" rel="noopener noreferrer">https://doi.org/10.1007/s44492-026-00012-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44492-026-00012-2" rel="noopener noreferrer">10.1007/s44492-026-00012-2</a></p>
<p><strong>Keywords:</strong> ultra-high carbon steel, Martensite, deep learning, ResNet50, transfer learning, class imbalance, microstructure classification, focal loss, Grad-CAM, t-SNE, GLCM, rare-phase detection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202416</post-id>	</item>
	</channel>
</rss>
