<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>focal loss &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/focal-loss/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 08:52:30 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>focal loss &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Post-Processing Sharpens Rainfall Forecasts Where Numerical Models Fall Short</title>
		<link>https://scienmag.com/ai-post-processing-sharpens-rainfall-forecasts-where-numerical-models-fall-short/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 08:52:30 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in meteorology]]></category>
		<category><![CDATA[channel attention]]></category>
		<category><![CDATA[climate and extreme weather forecasting]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for weather forecasting]]></category>
		<category><![CDATA[EfficientNet]]></category>
		<category><![CDATA[extreme rainfall]]></category>
		<category><![CDATA[flash flood prediction]]></category>
		<category><![CDATA[flood preparedness]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[heavy rainfall event prediction]]></category>
		<category><![CDATA[improving operational weather forecasts]]></category>
		<category><![CDATA[multi-task learning]]></category>
		<category><![CDATA[neural network weather models]]></category>
		<category><![CDATA[numerical weather prediction]]></category>
		<category><![CDATA[post-processing]]></category>
		<category><![CDATA[post-processing numerical weather models]]></category>
		<category><![CDATA[PostRainBench]]></category>
		<category><![CDATA[precipitation bias correction]]></category>
		<category><![CDATA[precipitation forecasting]]></category>
		<category><![CDATA[Rainfall forecast correction]]></category>
		<category><![CDATA[rainfall prediction accuracy]]></category>
		<category><![CDATA[Swin Transformer]]></category>
		<category><![CDATA[weather forecasting]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=243873</guid>

					<description><![CDATA[A new multi-scale deep learning framework called Rain-PostEFNet substantially improves numerical weather prediction rainfall forecasts, with the largest gains in detecting rare heavy precipitation events.]]></description>
										<content:encoded><![CDATA[<p>Every day, weather services around the world rely on numerical weather prediction models to answer one of society&#8217;s most consequential questions: where and how hard will it rain? These models, built on the equations of atmospheric physics, have transformed forecasting over the past half century, yet they remain stubbornly imperfect when it comes to precipitation. They carry systematic biases, resolve rainfall only at coarse spatial scales, and routinely miss the extreme downpours that trigger flash floods and landslides. A new study published in Applied Intelligence by Farah Naz, Lei She, Xincan Sui, Chenghong Zhang and Jie Shao tackles this gap from a fresh angle, introducing a deep learning framework called Rain-PostEFNet that sits on top of existing numerical forecasts and corrects them, with particularly striking gains for heavy rainfall events.</p>
<p>The core idea behind the work is post-processing. Rather than replacing numerical weather prediction altogether, the researchers treat the model output as a starting point that can be refined by a neural network trained on observations. This approach acknowledges a practical reality: numerical models encode decades of atmospheric science and remain the backbone of operational forecasting, but their raw precipitation fields often need correction before they reach emergency managers, hydrologists and the public. Previous deep learning post-processing efforts have shown promise, but the authors identify two persistent weaknesses. First, many networks fail to capture meteorological features that operate across multiple spatial scales, from broad synoptic rain bands down to localized convective cells. Second, and perhaps more fundamentally, extreme precipitation is vanishingly rare in training data, creating a severe class imbalance that causes standard networks to systematically underpredict the very events that matter most for disaster preparedness.</p>
<p>Rain-PostEFNet addresses these weaknesses through a carefully assembled architecture that borrows proven components from computer vision research. For feature extraction, the framework employs EfficientNet, a convolutional neural network family renowned for its efficient scaling of network depth, width and input resolution. On top of that sits ECA-Net, an efficient channel attention mechanism that adaptively reweights the importance of different feature channels, allowing the model to emphasize the meteorological signals most relevant to rainfall at each location and scale. The third pillar is FPN-Swin-Unet, which combines a feature pyramid network with a Swin transformer-based U-Net to refine features across multiple resolutions. The U-Net lineage traces back to biomedical image segmentation, where encoder-decoder structures proved adept at recovering fine spatial detail, and the Swin transformer adds hierarchical, window-based attention that captures long-range dependencies without prohibitive computational cost.</p>
<p>What distinguishes the framework from many predecessors, however, is its multi-task learning strategy. Instead of treating precipitation forecasting as a single regression problem, Rain-PostEFNet jointly optimizes two related tasks: classifying whether rain will fall in a given category, and regressing the actual precipitation amounts. The classification branch is trained with a weighted focal loss, a function specifically designed for imbalanced datasets. Focal loss down-weights the easy, abundant examples of light or no rain and focuses the network&#8217;s learning capacity on the rare, hard examples of intense precipitation. The regression branch uses the familiar mean squared error loss to sharpen the quantitative accuracy of predicted rainfall amounts. By training both objectives simultaneously, the network learns representations that serve detection and estimation at once, which the authors argue is key to improving predictive skill for rare heavy rain events.</p>
<p>Evaluating a precipitation forecasting model fairly requires standardized benchmarks, and the team turned to PostRainBench, a comprehensive benchmark dataset for post-processing numerical weather prediction that spans three climatologically distinct regions: China, Germany and Korea. This diversity matters, because rainfall regimes differ dramatically across these countries, from the monsoon-driven deluges of East Asia to the more moderate frontal systems of Central Europe. Performance was measured with three established verification metrics: the Critical Success Index, which captures the ratio of correct rain detections to all predicted and observed events; the Heidke Skill Score, which measures improvement over random chance; and overall accuracy. Together these metrics reward models that find the balance between detecting real rain and avoiding false alarms, a trade-off that has long challenged forecast verification.</p>
<p>The results are remarkable, especially for the Chinese dataset. For moderate rainfall, Rain-PostEFNet achieved relative improvements in Critical Success Index of 63.94 percent in China, 4.28 percent in Germany and 2.1 percent in Korea compared with state-of-the-art baseline methods. For heavy rainfall, the gains were even larger in China, reaching 78.34 percent, alongside 9.09 percent in Germany and 7.68 percent in Korea. The Heidke Skill Score told a similar story, rising by 52.18 percent in China and 2.95 percent in Germany for moderate rain, and by 52.4 percent in China and 6.78 percent in Germany for heavy rain. The outsized Chinese improvements likely reflect both the challenging nature of that region&#8217;s precipitation and the severity of the baseline errors there, but the consistent direction of improvement across all three countries suggests the framework captures something genuinely general about how to correct numerical forecasts.</p>
<p>To verify that the performance gains truly stem from the proposed design rather than sheer model size, the authors conducted an ablation study, systematically removing components and measuring the impact. The analysis confirmed that both the multi-task learning framework and the attention mechanisms contribute meaningfully to refining the numerical predictions. In other words, the joint classification-regression objective and the adaptive channel attention are not decorative additions; each plays a measurable role in the network&#8217;s ability to detect and quantify rainfall. This kind of component-level validation is increasingly expected in machine learning research, where headline numbers can sometimes mask architectures whose benefits are poorly understood.</p>
<p>The practical implications extend well beyond the leaderboard. Accurate precipitation forecasting underpins disaster preparedness, water resource management and climate modeling, and the events that current systems handle worst, extreme rainfall episodes, are precisely those with the highest human and economic stakes. A post-processing layer like Rain-PostEFNet could be deployed alongside existing operational numerical models, sharpening the rainfall fields that feed flood warning systems, reservoir operations and agricultural planning. Because the approach refines rather than replaces numerical forecasts, it complements the substantial investments meteorological agencies have already made in physical modeling, and it can be retrained as those models evolve. The authors also note the broader context of artificial intelligence entering earth system science, from neural global forecasting models to data-driven process understanding, positioning their work as part of a quiet revolution in how weather prediction is built.</p>
<p>Transparency was a priority for the team. The PostRainBench dataset is publicly available on GitHub, and the Rain-PostEFNet code has been released as well, allowing other researchers to reproduce the experiments, apply the framework to new regions and build on the architecture. The research was supported by the National Key R&amp;D Program of China, the Joint Laboratory for Artificial Intelligence and Digital Meteorological Applications Research, the Yibin Science and Technology Program and the Sichuan Science and Technology Program, reflecting institutional commitment to integrating artificial intelligence into meteorological applications. The work was conducted by researchers at the University of Electronic Science and Technology of China, the Sichuan Artificial Intelligence Research Institute, the Institute of Plateau Meteorology of the China Meteorological Administration and the joint laboratory in Beijing.</p>
<p>Challenges remain before such systems become routine in operational forecasting. The severe class imbalance that motivates the focal loss never fully disappears, and performance gains in Germany and Korea, while consistent, are far more modest than in China, hinting that regional climate characteristics, dataset size and baseline quality all shape how much post-processing can deliver. Extreme events, by definition, offer few training examples, and no amount of architectural ingenuity fully substitutes for observations of rare phenomena. Still, the study demonstrates that multi-scale feature refinement and imbalance-aware training can extract substantially more skill from existing numerical forecasts, particularly for the heavy rainfall that matters most. As climate change intensifies the hydrological cycle and pushes precipitation extremes beyond the historical distributions on which both physical models and neural networks were trained, tools that squeeze additional accuracy from every forecast will only grow in value. Rain-PostEFNet offers a template for how computer vision and meteorology can converge on that task, one corrected rain map at a time.</p>
<p><strong>Subject of Research:</strong> Deep learning post-processing of numerical weather prediction precipitation forecasts</p>
<p><strong>Article Title:</strong> Enhancing NWP precipitation forecasting with Rain-PostEFNet: A multi-scale deep learning approach</p>
<p><strong>Article References:</strong> Naz, F., She, L., Sui, X., Zhang, C., &amp; Shao, J. (2026). Enhancing NWP precipitation forecasting with Rain-PostEFNet: A multi-scale deep learning approach. <em>Applied Intelligence, 56</em>(14), Article 408. <a href="https://doi.org/10.1007/s10489-026-07457-x" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07457-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07457-x" rel="noopener noreferrer">10.1007/s10489-026-07457-x</a></p>
<p><strong>Keywords:</strong> precipitation forecasting, numerical weather prediction, deep learning, post-processing, extreme rainfall, multi-task learning, channel attention, EfficientNet, Swin transformer, PostRainBench, focal loss, flood preparedness</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">243873</post-id>	</item>
		<item>
		<title>New Lightweight Framework Puts Rigorous Deepfake Detection Within Reach of Modest Hardware</title>
		<link>https://scienmag.com/new-lightweight-framework-puts-rigorous-deepfake-detection-within-reach-of-modest-hardware/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 01:57:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[accessible deepfake detection methods]]></category>
		<category><![CDATA[affordable AI tools for small labs]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[benchmarking AI models for deepfake detection]]></category>
		<category><![CDATA[collaborative international AI research]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[ConvNeXt]]></category>
		<category><![CDATA[cross-dataset evaluation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deepfake detection]]></category>
		<category><![CDATA[deepfake detection on modest hardware]]></category>
		<category><![CDATA[deepfakes]]></category>
		<category><![CDATA[ethical concerns of deepfake technology]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[generative content manipulation]]></category>
		<category><![CDATA[Light-wDEF]]></category>
		<category><![CDATA[lightweight deepfake detection framework]]></category>
		<category><![CDATA[MTCNN]]></category>
		<category><![CDATA[multimodal detection]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[reproducibility in AI research]]></category>
		<category><![CDATA[resource-efficient computing]]></category>
		<category><![CDATA[societal hazards of digital forgeries]]></category>
		<category><![CDATA[societal impact of deepfakes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=242971</guid>

					<description><![CDATA[Researchers have introduced Light-wDEF, a resource-efficient and fully reproducible framework that enabled a fair comparison of ten visual deepfake detectors across five public datasets and identified ConvNeXt-Base as the top performer on modest hardware.]]></description>
										<content:encoded><![CDATA[<p>Deepfake technology has moved from laboratory curiosity to a genuine societal hazard in a remarkably short span of time. By synthesizing and manipulating digital content, modern generative pipelines can produce videos in which a person&#8217;s identity, facial appearance, or actions are altered so convincingly that even careful viewers struggle to spot the forgery. The consequences reach far beyond embarrassment: fabricated footage can damage personal reputations, erode dignity, and corrode the integrity of public discourse. Detection research has kept pace in accuracy, but often at a hidden cost. State-of-the-art detectors are typically benchmarked on clusters of high-end GPUs, with pipelines whose sampling choices, augmentations, and training details are documented unevenly. That combination makes results hard to reproduce and hard to compare, and it effectively locks out research groups, small labs, and students who lack access to expensive computing infrastructure. A new study published in Applied Intelligence argues that this inequity is not just a convenience problem but a scientific one, because unreproducible benchmarks undermine the very reliability that deepfake detection desperately needs.</p>
<p>The paper, authored by Muhammad Yousaf, Sarwar Shah Khan, Babar Shah, and colleagues working across institutions in Pakistan, the United Arab Emirates, and Portugal, introduces the Lightweight Deepfake Evaluation Framework, or Light-wDEF. The framework is designed from the ground up for large-scale, rigorous evaluation of visual deepfake detectors on modest hardware, including the class of free cloud accelerators typified by a Colab T4 GPU. Rather than chasing leaderboard-topping accuracy with ever larger models, the authors focused on a different bottleneck: experimental standardization. Light-wDEF packages a complete evaluation workflow in which every stage, from frame selection to final metric computation, is deterministic, logged, and resumable. The goal is that two researchers running the same configuration on different continents should obtain numerically identical results, a property the authors support by fixing random seeds, archiving checkpoints, and recording full environment metadata for exact replication.</p>
<p>Technically, the framework rests on several carefully chosen components. Deterministic frame sampling replaces ad hoc video frame selection with a reproducible rule, ensuring that the same frames are drawn from every video in every run. Face extraction, a notoriously expensive preprocessing step, is handled by batched MTCNN detection, which groups faces into batches to maximize throughput on limited GPU memory. Storage is organized as a streaming, append-only archive with lazy offset indexing, meaning extracted face crops are written once and later read by reference rather than duplicated, which keeps disk usage predictable even across datasets containing hundreds of thousands of frames. Data augmentation is performed online using the Albumentations library, so transformed samples are generated on the fly during training rather than precomputed and stored. Dataset splits are stratified with class-aware sampling to preserve the balance between real and fake classes, a critical detail in forensic benchmarks where manipulation types are unevenly represented.</p>
<p>Training under Light-wDEF is equally deliberate. The framework uses the AdamW optimizer, which decouples weight decay from gradient updates and has become a standard choice for stable deep learning training, and it optionally employs focal loss, a formulation that down-weights easy examples and concentrates learning on hard, ambiguous cases, a useful property when forgeries vary widely in visual quality. Per-epoch checkpointing with exact resume capability means that a training run interrupted by hardware limits, timeouts, or crashes can continue precisely where it stopped, without silently altering the experimental trajectory. Together these choices address a chronic weakness in the deepfake literature: pipelines that are so fragile and opaque that reported numbers cannot be trusted as a basis for comparison between competing methods.</p>
<p>To demonstrate the framework&#8217;s value, the authors trained and evaluated ten modern visual models under strictly identical conditions across five widely used public datasets: Celeb-DF, FaceForensics++, the Google DeepFakeDetection dataset, FakeAVCeleb, and the DFDC Preview set. This cross-dataset design matters because a detector that excels on the data it was trained on may collapse when confronted with a different generation method, a phenomenon known as poor generalization. By holding every hyperparameter, augmentation, and sampling decision constant, Light-wDEF isolates the effect of architecture itself, allowing a fair architectural bake-off rather than a comparison confounded by tuning differences. The evaluated architectures spanned the modern computer vision landscape, including convolutional families such as ResNet, EfficientNet, and ConvNeXt alongside transformer-based designs descended from the Vision Transformer lineage.</p>
<p>The headline result is that ConvNeXt-Base, a modernized convolutional network, emerged as the most effective visual model in the study. It achieved an area under the ROC curve of at least 0.96 and an F1 score of at least 0.94 on the evaluated sets, leading on AUC, balanced accuracy, F1, precision, and recall, all while remaining compatible with T4-class resource constraints. The authors&#8217; broader conclusion is arguably more interesting than the single winner: convolutional families provide the best practical balance of discrimination and efficiency. This finding pushes back against the assumption that vision transformers, which dominate many image recognition benchmarks, automatically translate into better forensic performance under realistic compute budgets. For practitioners deploying detectors in the real world, where latency, memory, and energy all carry costs, the efficiency side of that balance is not a luxury but a requirement.</p>
<p>The study also delivers a sobering caveat. On multimodal corpora such as FakeAVCeleb, which combines video manipulation with audio cues, visual-only detection showed clear limitations. A detector that watches faces alone cannot exploit inconsistencies in speech, lip synchronization, or audio artifacts, and the results quantify how much that blind spot costs. The authors explicitly frame this as motivation for future work on multimodal fusion, in which audio and visual evidence are combined, and on domain adaptation, which would help detectors trained on one set of generation methods transfer to unseen ones. In an era when generative tools increasingly manipulate voice alongside face, the finding suggests that purely visual detection, however well engineered, will remain an incomplete shield.</p>
<p>Beyond its specific results, the paper makes a case about how the field should conduct itself. Reproducibility in machine learning research is frequently discussed but rarely enforced, and deepfake detection is particularly vulnerable because datasets are large, preprocessing is complex, and subtle choices like frame sampling can shift reported metrics by meaningful margins. By releasing fixed seeds, archived checkpoints, and environment metadata alongside the framework, the authors provide a template that other groups can adopt even if they do not use Light-wDEF itself. The emphasis on resource-conscious design also carries a democratizing implication: a student with a free cloud notebook can now run a benchmarking protocol equivalent in rigor to one executed on a dedicated GPU cluster, which widens the pool of researchers able to contribute credible results to the field.</p>
<p>The stakes of getting this right are considerable. As deepfake generation tools become cheaper and more accessible, the arms race between creation and detection intensifies, and society&#8217;s ability to trust digital media hangs on the reliability of the detectors deployed at scale. Frameworks like Light-wDEF do not catch a single forgery; instead, they make the science of catching forgeries more trustworthy, comparable, and inclusive. If the field embraces standardized, reproducible, resource-efficient evaluation, progress can be measured honestly, weak claims can be exposed quickly, and the best detection ideas, whatever their architectural origin, can be identified and deployed where they matter most. In that sense, the quiet engineering discipline embodied in this study may prove as consequential as any individual detection model it helped to evaluate.</p>
<p><strong>Subject of Research:</strong> A lightweight, reproducible evaluation framework for benchmarking visual deepfake detection models on resource-constrained hardware</p>
<p><strong>Article Title:</strong> Light-wDEF: lightweight deepfake evaluation framework for efficient and reliable visual detection</p>
<p><strong>Article References:</strong> Yousaf, M., Khan, S. S., Shah, B., Khan, M. S., Bacha, M. A., Akbar, M. S., Ali, I., &amp; Moreira, F. (2026). Light-wDEF: lightweight deepfake evaluation framework for efficient and reliable visual detection. <em>Applied Intelligence, 56</em>(14), Article 411. <a href="https://doi.org/10.1007/s10489-026-07460-2" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07460-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07460-2" rel="noopener noreferrer">10.1007/s10489-026-07460-2</a></p>
<p><strong>Keywords:</strong> deepfakes, Light-wDEF, deep learning, ConvNeXt, cross-dataset evaluation, reproducibility, computer vision, MTCNN, focal loss, multimodal detection, resource-efficient computing, benchmarking</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">242971</post-id>	</item>
		<item>
		<title>AI Framework Combines Contrastive Learning and Graph Attention to Spot Financial Legal Risk</title>
		<link>https://scienmag.com/ai-framework-combines-contrastive-learning-and-graph-attention-to-spot-financial-legal-risk/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 21:36:19 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[addressing label sparsity in financial datasets]]></category>
		<category><![CDATA[AI frameworks for risk control]]></category>
		<category><![CDATA[analyzing corporate transaction networks]]></category>
		<category><![CDATA[anti-money laundering]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[contrastive learning in finance]]></category>
		<category><![CDATA[detecting fraud rings and hidden transactions]]></category>
		<category><![CDATA[financial credit risk]]></category>
		<category><![CDATA[Financial legal risk detection]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[fraud detection]]></category>
		<category><![CDATA[GATv2]]></category>
		<category><![CDATA[GATv2 for legal risk detection]]></category>
		<category><![CDATA[graph attention networks for risk analysis]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[graph neural networks in finance]]></category>
		<category><![CDATA[heterogeneous graphs]]></category>
		<category><![CDATA[label sparsity]]></category>
		<category><![CDATA[legal risk mitigation]]></category>
		<category><![CDATA[legal violation identification in financial data]]></category>
		<category><![CDATA[machine learning for financial compliance]]></category>
		<category><![CDATA[MoCo]]></category>
		<category><![CDATA[momentum contrastive learning (MoCo)]]></category>
		<category><![CDATA[risk management]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=239320</guid>

					<description><![CDATA[A new framework pairs momentum contrastive learning with the GATv2 graph attention network to detect rare, legally entangled financial credit risks that conventional scoring models miss.]]></description>
										<content:encoded><![CDATA[<p>Financial institutions face a paradox at the heart of modern risk control: the most dangerous borrowers and corporate networks are precisely the ones that appear least often in the data. Severe legal violations—fraud rings, hidden related-party transactions, litigation entanglements—make up a vanishingly small fraction of millions of transaction records, and conventional credit scoring models, built on the assumption that loan applicants are statistically independent of one another, simply cannot see the webs of connection that link a defaulting borrower to a shell company, a guarantor, and a court case. A new study published in Discover Artificial Intelligence proposes a way out of this trap, pairing two machine learning techniques—momentum contrastive learning (MoCo) and the graph attention network GATv2—into a single framework designed specifically for the legal dimension of financial credit risk.</p>
<p>The study, authored by Di Teng of Harbin Finance University, addresses two stubborn problems that have limited earlier attempts to apply graph neural networks to financial compliance. The first is label sparsity. In a typical compliance dataset, confirmed instances of serious legal violations are so rare that a supervised model trained directly on them tends to overfit, memorizing the few known bad actors rather than learning generalizable patterns of risk. The second is a structural flaw in the standard Graph Attention Network, or GAT, one of the most widely used architectures for learning from networked data. In GAT, the attention weight assigned to a connection is computed in a way that is effectively static: for a given edge, the score depends only weakly on which node is asking the question. That rigidity makes it hard for the model to trace how legal risk actually propagates through a financial network, where the significance of a relationship changes dramatically depending on the perspective of the entity involved.</p>
<p>Teng&#8217;s framework unfolds in two stages. In the first, MoCo performs unsupervised pre-training on the unlabeled graph of financial entities. MoCo, originally developed for computer vision, works by maintaining two encoders: a query encoder that is updated by normal backpropagation, and a momentum encoder whose parameters drift slowly toward the query encoder according to the update rule in which the key parameters are replaced by a weighted blend of their previous values and the query parameters. The momentum encoder&#8217;s outputs are pushed into a large queue of negative samples, giving the model a vast dictionary of examples against which each new node can be compared. The training objective, an InfoNCE loss, pushes each node&#8217;s representation toward its positive counterpart and away from thousands of negatives, forcing the network to learn high-order semantic features without ever seeing a label. Crucially, the study adapts the queue mechanism to serve a compliance-specific purpose: buffering rare, high-risk legal event nodes so that their features are preserved rather than drowned out by the overwhelming majority of benign entities.</p>
<p>In the second stage, the pre-trained embeddings flow into GATv2, a refined attention architecture that fixes the static attention problem of its predecessor. GATv2 computes attention scores by applying a learnable linear transformation and a LeakyReLU non-linearity to the concatenation of two nodes&#8217; features before taking the inner product with a learnable attention vector. This reordering of operations makes the attention mechanism strictly more expressive: the importance of a neighbor now genuinely depends on the query node, allowing the model to assign different weights to the same relationship depending on whose risk is being assessed. After normalization with a softmax function across each node&#8217;s neighborhood, the attention weights are used to aggregate neighbor features into an updated node representation, and multiple attention heads run in parallel to capture different facets of the risk landscape.</p>
<p>The full pipeline is trained with a composite loss function that reflects the realities of compliance work. The supervised component is Focal Loss, a formulation that down-weights easy, well-classified examples and concentrates the model&#8217;s gradient signal on the hard cases—precisely the meticulously disguised violations that fraud rings engineer to look ordinary. An auxiliary contrastive loss from the pre-training stage is retained during fine-tuning, weighted by a hyperparameter, to preserve the structure of the learned feature manifold, and an L2 regularization term guards against overfitting. When the model&#8217;s predicted probability for an entity exceeds a decision threshold calibrated by ROC analysis, the entity is flagged as high-risk, triggering a detailed legal compliance review such as an anti-money-laundering check, a contract compliance audit, or a litigation risk warning.</p>
<p>The framework was evaluated on two datasets. The first is the public Lending Club dataset of personal loans and defaults. The second is Fin-Law-CN, a proprietary heterogeneous graph built for the study containing roughly 20,000 entity nodes and 80,000 typed edges. Its nodes include borrowers, enterprises, guarantors, court-case records, and transaction events, while its edges capture borrower-enterprise links, ownership and control relations, guarantees, litigation involvement, and simulated supply-chain transactions. Legal-risk labels were mapped into three tiers—low risk, medium risk, and high risk—based on observable default, litigation, and compliance-warning outcomes, and the data were cleaned by merging duplicate entities and normalizing inconsistent identifiers before graph construction. The data were split 7:1:2 into training, validation, and test sets, and the model was benchmarked against GCN, GAT, GraphSAGE, RGCN, and the Heterogeneous Graph Transformer under a unified grid search over learning rates, dropout rates, and embedding dimensions.</p>
<p>The results were decisive. On Fin-Law-CN, where conventional models struggled with the intricate mix of relationship types, the proposed approach achieved an F1-Score roughly 4 to 6 percentage points higher than the baselines, with an AUC of 0.86. On the classification benchmarks, the MoCo-GATv2 model reached an AUC of 0.96, compared with 0.85 for standard GAT and 0.76 for GCN, and its precision-recall curve dominated across recall levels. Training curves showed the model converging to a loss of about 0.7 within roughly 100 epochs while validation accuracy climbed to approximately 0.89, against 0.85 for GAT and 0.79 for GCN—evidence that the combination is not only more accurate but also more stable during training.</p>
<p>The study&#8217;s ablation experiments dissected where the gains come from. Removing the MoCo pre-training module dropped accuracy, F1-Score, and AUC from roughly 98, 96, and 97 percent to about 90, 88, and 89 percent, confirming that self-supervised representation learning is critical when labels are scarce. Replacing GATv2 with standard GAT was even more damaging, sinking the metrics to around 85, 83, and 84 percent and underscoring the value of dynamic attention for tracing risk propagation. Dropping multi-head attention cost a few more points. Parameter sensitivity analyses showed performance peaking at embedding dimensions of 64 or 128, and the MoCo queue size experiments revealed a sweet spot: downstream F1 rose from about 0.80 at a queue of 256 to a peak of 0.932 near 16,384, then plateaued and slightly declined at larger sizes, suggesting diminishing returns and potential redundancy from excessive negative sampling. Attention-head analysis identified heads H1, H5, and H8 as the most influential in the shallow layers, while some heads contributed almost nothing, hinting at room for pruning.</p>
<p>Practical deployment considerations were addressed as well. On a single NVIDIA RTX 3090 GPU, the eight-head configuration required about 1.2 seconds per training epoch on Fin-Law-CN and sustained inference latency of roughly 15 milliseconds per subgraph—fast enough for real-time risk control. The architecture also incorporates a human-machine feedback loop: expert reviewers examine flagged high-risk events, their judgments re-label or update annotations in the graph database, and the model is retrained on the enriched data in a version-controlled cycle that continuously sharpens its accuracy.</p>
<p>The author is candid about the framework&#8217;s limits. Validation so far rests on one public and one proprietary dataset, both drawn from a single regulatory context, and cross-market testing in the United States, Europe, and beyond remains to be done. The experiments did not include dedicated adversarial attacks, deliberately disguised fraud-ring subsets, or tests of temporal propagation delays, and the model lacks a full explainable-AI module capable of producing legally auditable, case-level justifications for its flags—something regulators are likely to demand. Building and maintaining large heterogeneous financial-legal graphs also carries substantial deployment cost. Future work, the study notes, will incorporate temporal graph neural networks to capture how risk spreads across evolving enterprise networks, explainability methods such as GNNExplainer and SHAP-style attribution, federated learning for privacy-preserving collaboration across institutions, and robustness testing under structural adversarial attacks. If those extensions succeed, the combination of contrastive pre-training and dynamic graph attention could become a standard weapon in the fight against the hidden networks behind financial crime.</p>
<p><strong>Subject of Research:</strong> Machine learning methods for legal prevention and control of financial credit risk</p>
<p><strong>Article Title:</strong> Research on legal prevention and control of financial credit risk based on MoCo and GATv2 algorithm</p>
<p><strong>Article References:</strong> Teng, D. (2026). Research on legal prevention and control of financial credit risk based on MoCo and GATv2 algorithm. <em>Discover Artificial Intelligence, 6</em>(1), Article 1356. <a href="https://doi.org/10.1007/s44163-026-02297-7" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02297-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02297-7" rel="noopener noreferrer">10.1007/s44163-026-02297-7</a></p>
<p><strong>Keywords:</strong> MoCo, GATv2, graph neural networks, financial credit risk, contrastive learning, legal risk mitigation, anti-money laundering, fraud detection, label sparsity, heterogeneous graphs, focal loss, risk management</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">239320</post-id>	</item>
		<item>
		<title>Lightweight AI Sees Through Murky Water to Spot Defects in Underground Drainage Pipelines</title>
		<link>https://scienmag.com/lightweight-ai-sees-through-murky-water-to-spot-defects-in-underground-drainage-pipelines/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 09:53:31 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in submerged infrastructure monitoring]]></category>
		<category><![CDATA[AI-powered sonar for sewer defect detection]]></category>
		<category><![CDATA[artificial intelligence in urban infrastructure maintenance]]></category>
		<category><![CDATA[attention mechanism]]></category>
		<category><![CDATA[automated underground pipe flaw detection]]></category>
		<category><![CDATA[compact underwater robot inspection systems]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[defect detection]]></category>
		<category><![CDATA[drainage pipelines]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[lightweight AI hardware for sewer inspection]]></category>
		<category><![CDATA[lightweight neural networks]]></category>
		<category><![CDATA[low-visibility pipeline imaging solutions]]></category>
		<category><![CDATA[multi-scale feature extraction]]></category>
		<category><![CDATA[pipeline inspection]]></category>
		<category><![CDATA[real-time sewer defect detection]]></category>
		<category><![CDATA[sediment-penetrating sonar imaging]]></category>
		<category><![CDATA[ShuffleNet]]></category>
		<category><![CDATA[sonar]]></category>
		<category><![CDATA[submerged pipeline crack identification]]></category>
		<category><![CDATA[turbid water pipeline inspection technology]]></category>
		<category><![CDATA[Underground drainage pipeline inspection]]></category>
		<category><![CDATA[underwater acoustics]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=237516</guid>

					<description><![CDATA[Researchers at Hohai University have developed ShuffleNet-MSAA, a lightweight AI system that uses adaptive attention and sonar imaging to detect defects in flooded underground drainage pipelines where cameras fail.]]></description>
										<content:encoded><![CDATA[<p>Beneath the streets of nearly every city on Earth runs a hidden network of drainage pipelines, most of them flooded, dark, and increasingly past their design life. Inspecting them has always been a miserable compromise: cameras fail in turbid water where visibility drops to centimeters, and sending crews into confined, submerged spaces is slow and dangerous. Now a team at Hohai University in China reports a new artificial intelligence system that lets compact sonar devices do the job instead, detecting cracks and defects in drainage pipes even when the water is thick with sediment. The system, described in the journal Multimedia Tools and Applications, is light enough to run on the kind of modest hardware that can actually be mounted on a small underwater robot, which may be the detail that pushes it out of the laboratory and into real sewers.</p>
<p>The research, led by Kao Ge and Qing-Bang Han, tackles a problem that has long frustrated engineers: sonar is the only practical sensor in fully flooded, opaque pipes, but sonar images are notoriously ugly to look at and harder to interpret. Acoustic imaging produces grainy, low-contrast pictures in which a small fracture can occupy just a handful of pixels against a background of acoustic speckle, reverberation from pipe walls, and noise from the water itself. Conventional computer vision models, trained mostly on crisp optical photographs, tend to collapse under these conditions. Deep learning models that do work well on sonar are usually too large and power-hungry to embed in the battery-constrained vehicles that would actually crawl through a municipal drain.</p>
<p>The Hohai team&#8217;s answer is an architecture they call ShuffleNet-MSAA, and its design reflects a careful trade-off between intelligence and efficiency. The backbone of the network is ShuffleNet, a convolutional neural network originally engineered for smartphones and other mobile devices. ShuffleNet achieves its speed through a clever trick known as channel shuffling, in which information is mixed across feature channels in a way that preserves accuracy while drastically cutting the number of computations. That makes it an ideal starting point for a detector that must live on embedded hardware rather than in a data center. But a lightweight backbone alone cannot solve the fundamental signal problems of sonar imagery, so the researchers layered two specialized modules on top of it.</p>
<p>The first is a multi-scale feature extractor. Defects in a drainage pipe come in wildly different sizes: a hairline crack may span a few pixels while a collapsed section dominates the frame. Networks that look at an image through a single receptive field tend to miss one extreme or the other. By extracting and fusing features at multiple scales simultaneously, the new model can attend to both tiny anomalies and large structural failures within the same pass. The authors report that this multi-scale strategy was particularly important for small targets, which are precisely the defects most likely to be overlooked by human inspectors and by earlier automated systems alike.</p>
<p>The second innovation is the multi-scale adaptive attention module, the MSAA of the system&#8217;s name. Attention mechanisms in modern AI work roughly the way human attention does: they let a network decide which parts of an input deserve focus and which can be ignored. What distinguishes the new module is that it operates pixel by pixel and adapts dynamically to each image, weighting the contribution of every location according to how likely it is to contain genuine defect information rather than noise. Notably, the researchers implemented the module using median filtering, a classical signal-processing technique prized for suppressing impulsive noise without blurring edges. Combining this statistical robustness with learned attention allows the network to quiet the acoustic clutter of a turbid pipe while sharpening the faint signatures of real defects.</p>
<p>Even a well-designed network can be undermined by the data it learns from. Sonar defect datasets are inherently imbalanced: catastrophic failures are rare, minor defects are common, and some defect categories may be represented by only a handful of examples, producing what statisticians call a long-tailed distribution. Standard training procedures end up favoring the frequent categories and neglecting the rare ones, which is exactly backwards for infrastructure monitoring, where catching an uncommon but severe fault matters most. To counter this, the team introduced a dual-weighted focal loss, or DW-FL, a modified training objective that extends the focal loss concept originally developed for dense object detection. The dual weighting scheme adjusts the loss both for class frequency and for example difficulty, forcing the network to invest its learning capacity in the rare, hard-to-classify defects that conventional training would gloss over.</p>
<p>The experimental results are the strongest part of the story. The researchers evaluated ShuffleNet-MSAA on both a custom dataset of drainage pipe sonar images and a public sonar defect benchmark, comparing it against state-of-the-art baselines built on both convolutional neural networks and Transformer architectures. The proposed system outperformed these competitors in both detection and classification performance, while remaining lightweight enough for deployment on resource-constrained platforms. The authors emphasize that the model showed high accuracy, robustness, and generalization across the two datasets, a combination that matters because a detector tuned to one pipe network&#8217;s acoustics often fails when moved to another. The work was supported by the Natural Science Foundation of China, the Key Research and Development Project of Changzhou in Jiangsu Province, a Jiangsu Provincial graduate innovation program, and the company AutoSubsea Vehicles Inc., an affiliation that hints at commercial ambitions for the technology.</p>
<p>What makes this study worth wider attention is how it fits into a broader shift in how societies maintain their buried infrastructure. Cities worldwide are grappling with aging water systems, and the cost of undetected pipe failures, from sinkholes to sewage contamination, runs into billions annually. Earlier automation efforts relied on closed-circuit television, which works only in dry or near-dry pipes, or on hand-crafted acoustic features that required expert tuning. The new work builds on a decade of progress in sonar deep learning, including transfer learning approaches, hybrid CNN-Transformer frameworks, and attention modules adapted from general computer vision, but it is among the first to target the specific, punishing conditions of underground drainage: full submersion, heavy sediment, severe noise, and a hard ceiling on compute. The pixel-wise adaptive attention design also echoes a trend in medical imaging, where adaptive spatial weighting has recently improved segmentation under noisy conditions, suggesting a convergence of techniques across domains that share the same core problem, extracting weak signals from hostile data.</p>
<p>There are, of course, caveats. The published results come from curated datasets, and the true test will come when the system meets the chaotic reality of a working sewer, with its shifting debris, variable flow, and acoustic surprises no benchmark anticipated. The authors state that their data are available upon request, which should allow independent groups to probe the model&#8217;s limits. Still, the engineering philosophy on display, pairing an efficient mobile-grade backbone with task-specific attention and an imbalance-aware loss, offers a template that other underwater robotics applications, from subsea pipeline leak detection to seabed mapping, could readily adopt. As autonomous inspection vehicles grow cheaper and acoustic sensors grow better, the bottleneck is increasingly the intelligence onboard. Work like ShuffleNet-MSAA suggests that bottleneck is now being cleared, one carefully weighted pixel at a time, and that the hidden arteries beneath our cities may soon be examined routinely, safely, and without a single human entering the water.</p>
<p><strong>Subject of Research:</strong> Lightweight deep learning for sonar-based defect detection in underground drainage pipelines</p>
<p><strong>Article Title:</strong> ShuffleNet-MSAA: a lightweight multi-scale adaptive attention mechanism for enhanced sonar-based defect detection in underground drainage pipelines</p>
<p><strong>Article References:</strong> Ge, K., &amp; Han, Q.-B. (2026). ShuffleNet-MSAA: a lightweight multi-scale adaptive attention mechanism for enhanced sonar-based defect detection in underground drainage pipelines. <em>Multimedia Tools and Applications, 85</em>(9), Article 739. <a href="https://doi.org/10.1007/s11042-026-21816-3" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21816-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21816-3" rel="noopener noreferrer">10.1007/s11042-026-21816-3</a></p>
<p><strong>Keywords:</strong> sonar, defect detection, drainage pipelines, ShuffleNet, attention mechanism, deep learning, multi-scale feature extraction, focal loss, underwater acoustics, pipeline inspection, lightweight neural networks, computer vision</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">237516</post-id>	</item>
		<item>
		<title>AI Reads Students&#8217; Faces to Measure Attention in Online Classes</title>
		<link>https://scienmag.com/ai-reads-students-faces-to-measure-attention-in-online-classes/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 00:06:12 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[affective computing]]></category>
		<category><![CDATA[AI-based facial recognition for online student engagement]]></category>
		<category><![CDATA[attentiveness index]]></category>
		<category><![CDATA[challenges of monitoring attention in online education]]></category>
		<category><![CDATA[cognitive science insights into emotional states during e-learning]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional neural networks]]></category>
		<category><![CDATA[DAiSEE dataset]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for emotion recognition in virtual classrooms]]></category>
		<category><![CDATA[detecting boredom and frustration during online lessons]]></category>
		<category><![CDATA[e-learning]]></category>
		<category><![CDATA[engagement detection]]></category>
		<category><![CDATA[facial expression recognition]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[impact of affective states on online learning outcomes]]></category>
		<category><![CDATA[importance of student emotional states in digital learning]]></category>
		<category><![CDATA[online education]]></category>
		<category><![CDATA[real-time affective state detection in e-learning]]></category>
		<category><![CDATA[real-time analytics]]></category>
		<category><![CDATA[remote classroom feedback systems using facial analysis]]></category>
		<category><![CDATA[role of]]></category>
		<category><![CDATA[technological advancements in virtual classroom engagement assessment]]></category>
		<category><![CDATA[webcam-based attention measurement in remote education]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=236206</guid>

					<description><![CDATA[Researchers have developed a real-time computer vision system that reads students' facial expressions during online lessons and converts four learning-centered emotional states into a single attentiveness score for instructors.]]></description>
										<content:encoded><![CDATA[<p>For millions of students logging into virtual classrooms every day, the question of whether anyone is actually paying attention has long been unanswerable. In a physical lecture hall, an instructor can scan the room for slumped shoulders, wandering eyes, and furrowed brows, then adjust the pace of a lesson on the fly. Online, that feedback loop collapses. A grid of silent video tiles offers little insight into who is absorbing the material and who has mentally checked out. Now, a team of researchers at Tel Aviv University and Washington State University has built a system that tries to restore that lost channel of communication, using nothing more than a standard webcam and a set of deep learning models that read a learner&#8217;s face in real time.</p>
<p>The new framework, described in the journal Multimedia Tools and Applications, tackles a problem that researchers call affective state recognition in e-learning. Rather than trying to detect the seven universal emotions that dominate much of facial expression research, the system focuses on four learning-centered states that education scientists have identified as especially meaningful during study: boredom, engagement, confusion, and frustration. These states matter because decades of cognitive science research show they shape how people learn. Positive emotional states are associated with gains in creative thinking, while negative states push learners toward prioritizing performance over mastery. Boredom in particular has been repeatedly linked to suboptimal learning outcomes, whereas mild confusion, somewhat counterintuitively, can signal that a student is grappling productively with difficult material.</p>
<p>At the heart of the system is a multioutput classification model built on a parallel-branch architecture. The researchers trained two versions: one based on a lightweight convolutional neural network with roughly 470,000 parameters, and another built on EfficientNetB2, a more capable backbone with about 31.7 million parameters that remains efficient enough for cloud deployment. Each branch of the network is dedicated to a single affective state and outputs a probability vector across four intensity levels, from very low to very high. The final output is a combined vector capturing the predicted intensity of boredom, confusion, engagement, and frustration simultaneously. This design distinguishes the work from most prior studies, which typically classified only the single dimension of engagement and ignored the richer emotional context surrounding it.</p>
<p>Training relied on DAiSEE, the first multilabel video dataset created for engagement recognition in the wild. It contains 9,068 ten-second video clips of 112 individuals, recorded at full high-definition resolution with ordinary webcams, and annotated with crowd-sourced labels that were correlated against a gold standard produced by expert psychologists. The dataset is notoriously imbalanced: high-intensity engagement labels dominate, while lower-intensity labels are comparatively rare, and the opposite pattern holds for boredom, confusion, and frustration. To counteract the bias this imbalance would otherwise introduce, the team employed a categorical focal loss, a function that applies a modulating term to standard cross-entropy so that learning concentrates on hard, easily misclassified examples rather than being swamped by the majority class.</p>
<p>Preprocessing was deliberately aggressive. Each ten-second clip originally contains 300 frames, but because facial expressions in instructional settings evolve slowly, on the order of one to two seconds, the pipeline subsamples to just ten frames per video, a 96.7 percent reduction in data volume. Each retained frame passes through a Viola-Jones Haar Cascade detector to isolate the face, and frames where facial landmarks or pupils cannot be reliably detected are discarded. Valid crops are resized to 64-by-64 pixels and converted to normalized grayscale. The result is a compact, standardized input that minimizes computational cost while preserving the affective signals that matter, an essential property for a system intended to run continuously during live lectures.</p>
<p>The most novel contribution, however, is not the classifier itself but what the researchers call the attentiveness index. A single classification of four emotional states does not directly tell an instructor whether a student is attentive, so the team needed a way to compress those predictions into one interpretable number. They had multiple instructors score a subset of dataset videos for attentiveness on a scale of one to ten, then applied multiple linear regression to learn weights for each affective state. The resulting formula weights engagement most heavily and positively, gives confusion a smaller positive weight, and assigns negative weights to boredom and frustration, with boredom penalized far more severely. Crucially, these learned coefficients align with the existing cognitive science literature: boredom is strongly associated with poor learning, frustration more weakly so, and moderate confusion often accompanies productive learning.</p>
<p>The index held up under statistical scrutiny. On an unseen test set, the computed index correlated strongly and significantly with the instructor-annotated ground truth, yielding a Pearson correlation coefficient of 0.700 and a Spearman rank coefficient of 0.585, both statistically significant. Three-fold cross-validation produced a mean R-squared of 0.495, and a t-test confirmed the model&#8217;s predictive power at p equals 0.0246. Meanwhile, the underlying classifiers achieved state-of-the-art performance on DAiSEE. The CNN-based model reached accuracies of 73.04 percent on boredom, 79.98 percent on engagement, 80.17 percent on confusion, and 85.37 percent on frustration, while the EfficientNet variant performed comparably, peaking at 80.32 percent on engagement. The researchers argue these results outperform prior end-to-end approaches on the same dataset, including methods built on C3D, I3D, and hybrid architectures combining EfficientNet with temporal convolutional and recurrent networks.</p>
<p>What elevates the work from a modeling exercise to a practical tool is the end-to-end pipeline built around it. The system is deployed on a cloud server and supports simultaneous multi-user logins through a web client built with Flask, JavaScript, and HTML5 Canvas. During a live session, the learner&#8217;s webcam feed is analyzed in real time, with detected affective states displayed alongside the learning content. Instructors receive a separate dashboard offering both graphical and tabular analytics: class-level trends in average attentiveness over time, post-session distributions of affective states, and timestamped logs of individual learners&#8217; dominant states and index scores. When cumulative class engagement drops below a threshold, the instructor receives an alert, and the analytics can pinpoint exactly when attention sagged during a lecture, for instance, or flag moments when confusion spiked while a difficult concept was being explained. The system can even aggregate data across multiple lectures to suggest how to structure future sessions for maximum engagement.</p>
<p>The researchers are explicit that the tool is meant as a pedagogical aid, not a surveillance instrument. They caution that affective computing models can be sensitive to lighting, camera quality, and demographic variation, and they strongly advise against punitive or comparative uses of the system. Real-world deployments, they stress, must be strictly opt-in, with informed consent from all participants, and individual analytics should be disabled by default in classroom settings in favor of aggregated, class-level feedback. They also acknowledge a limitation long noted in the field: basic universal emotions are rare during genuine learning sessions, which is precisely why the focus on learning-centered states matters.</p>
<p>Future directions outlined by the team include multimodal fusion, combining facial affect with head movement, gaze tracking, and noninvasive EEG signals, along with explainable AI techniques that would visually highlight which facial features drive each prediction. They also plan consent-based periodic self-reports from learners to calibrate the model against individual differences, and the construction of a new, demographically balanced dataset to validate generalizability across racial and cultural backgrounds. If those efforts succeed, the humble webcam could become one of the most informative instruments in the virtual classroom, giving online instructors something they have lacked since lectures moved to screens: a live read on the minds in the room.</p>
<p><strong>Subject of Research:</strong> Real-time computer vision analysis of learner affective states and attentiveness in online education</p>
<p><strong>Article Title:</strong> Learner attentiveness and engagement analysis in online education using computer vision</p>
<p><strong>Article References:</strong> Gogawale, S., Deshpande, M., Kumar, P., &amp; Ben-Gal, I. (2026). Learner attentiveness and engagement analysis in online education using computer vision. <em>Multimedia Tools and Applications, 85</em>(9), Article 743. <a href="https://doi.org/10.1007/s11042-026-21884-5" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21884-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21884-5" rel="noopener noreferrer">10.1007/s11042-026-21884-5</a></p>
<p><strong>Keywords:</strong> computer vision, online education, affective computing, engagement detection, deep learning, convolutional neural networks, DAiSEE dataset, attentiveness index, e-learning, facial expression recognition, focal loss, real-time analytics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">236206</post-id>	</item>
		<item>
		<title>AI Learns to Read Ultrasound Echoes and Sort Weld Flaws with Expert Precision</title>
		<link>https://scienmag.com/ai-learns-to-read-ultrasound-echoes-and-sort-weld-flaws-with-expert-precision/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 06:34:02 +0000</pubDate>
				<category><![CDATA[Space]]></category>
		<category><![CDATA[advancements in automated weld inspection methods]]></category>
		<category><![CDATA[aerospace safety]]></category>
		<category><![CDATA[AI-powered non-destructive testing in aerospace]]></category>
		<category><![CDATA[challenges of large ultrasonic data volumes]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[complicating defect detection]]></category>
		<category><![CDATA[computer vision techniques applied to ultrasonic signals]]></category>
		<category><![CDATA[convolutional neural network]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for weld flaw classification]]></category>
		<category><![CDATA[defect classification]]></category>
		<category><![CDATA[defect region identification in weld inspection]]></category>
		<category><![CDATA[FiLM layer]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[integration of AI and ultrasonic testing technologies]]></category>
		<category><![CDATA[KAIST]]></category>
		<category><![CDATA[machine learning accuracy in ultrasonic flaw detection]]></category>
		<category><![CDATA[non-destructive testing]]></category>
		<category><![CDATA[phased-array ultrasonic testing]]></category>
		<category><![CDATA[phased-array ultrasonic testing automation]]></category>
		<category><![CDATA[physics-informed preprocessing for ultrasound signal analysis]]></category>
		<category><![CDATA[reliability of AI in critical aerospace component testing]]></category>
		<category><![CDATA[signal preprocessing]]></category>
		<category><![CDATA[ultrasound echo signals]]></category>
		<category><![CDATA[weld inspection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233914</guid>

					<description><![CDATA[KAIST researchers combined physics-based preprocessing with a FiLM-conditioned convolutional neural network to classify weld defects in phased-array ultrasonic testing, reaching 87.66 percent three-class accuracy and 99.26 percent binary defect detection.]]></description>
										<content:encoded><![CDATA[<p>Every welded joint in an aircraft fuselage, a rocket propellant tank, or a launch-vehicle structure carries a silent question: is it sound? Answering that question falls to non-destructive testing, and in recent years the workhorse of subsurface inspection has become phased-array ultrasonic testing, or PAUT. The technique steers beams of ultrasound electronically through metal, sweeping across welds and returning streams of echoes that trained inspectors must interpret one by one. Now a team at KAIST in South Korea has shown that a deep learning model, guided by physics-informed preprocessing and a clever conditioning layer borrowed from computer vision, can classify those echoes with an accuracy that rivals expert judgment, and can flag defective regions with 99.26 percent reliability.</p>
<p>The research, published in the International Journal of Aeronautical and Space Sciences by Jaeho Lee, Chen Ciang Chia, and Jung-Ryul Lee, tackles a problem that has long frustrated automation efforts. PAUT produces enormous volumes of A-scan signals, one-dimensional traces of echo amplitude over time, and those traces are notoriously messy. Echoes bounce off the weld root, the toe, and the cap, producing geometry signals whose amplitudes can rival those of genuine defects. The initial pulse generated near the probe contaminates the beginning of every beam. Noise drifts in from the environment. To a human inspector with years of experience, these impostors are recognizable; to a neural network trained on raw signals, they are indistinguishable from cracks and porosity.</p>
<p>The KAIST team built their dataset from 40 single V-groove butt weld specimens, half carbon steel and half stainless steel, with thicknesses of 15 and 25 millimeters. Each specimen contained three artificially introduced flaws drawn from six categories: lack of fusion, crack, porosity, slag, undercut, and incomplete penetration. Certified ASNT Level II inspectors scanned every specimen with an Olympus OmniScan MX2, positioning two probes on opposite sides of the weld so that each defect could be observed from both directions. Crucially, the researchers then treated each side as an independent observation, so the model would learn to classify defects from single-sided data alone, a realistic constraint for field inspections where dual-sided access is not always possible. ASNT Level III experts annotated defect type, position, and beam angle, establishing ground truth as regions of interest on the C-scan images.</p>
<p>The first pillar of the framework is a two-stage preprocessing scheme that encodes inspector expertise directly into the data. The first stage, time-of-flight masking, exploits a simple geometric fact: defects occur inside the weld and its heat-affected zone, at a calculable distance from the weld surface. Using the probe offset, wedge delay, beam angle, and acoustic velocity stored alongside the inspection data, the team computed the spatial position of every sample in every A-scan and zeroed out anything lying more than 6 millimeters beyond the weld surface. This single operation eliminated the initial pulse and much of the ambient noise in one stroke.</p>
<p>The second stage is more subtle and more ingenious. Geometry signals, unlike defect signals, appear consistently at similar positions along the scan axis throughout an entire inspection, because they originate from fixed weld features rather than from localized flaws. The researchers exploited this by building a statistical threshold for every beam angle and time index. For each index, they pooled amplitudes from a three-by-three neighborhood across all scan positions, sorted the resulting distribution, fitted a line to its lower 80 percent, and then shifted that line upward by a multiple of the standard deviation. Where the shifted line crossed the sorted distribution, they set the threshold. Defect signals, being sparse and isolated, produce a low-sloped distribution with a sharp rise only at the tail, so the threshold lands high and the defect survives. Geometry signals, being densely distributed along the scan axis, produce a steeper distribution that intersects the offset line early, so nearly all their samples fall below the threshold and are erased. The team tuned the offset multiplier empirically, using a value of 30 for the louder carbon steel specimens and 8 for the quieter stainless steel ones, and extended each surviving window by 20 samples on either side to preserve the full defect waveform.</p>
<p>The second pillar of the framework is the neural architecture itself. Rather than simply concatenating hand-crafted features with a convolutional neural network&#8217;s output, as earlier work had done, the team inserted a feature-wise linear modulation, or FiLM, layer into the network. FiLM, originally developed for visual reasoning, works by generating per-channel scale and shift parameters from auxiliary information and applying them to intermediate feature maps. Here, three physically meaningful features computed from each A-scan, a signed distance ratio locating the echo peak relative to the weld surfaces, the peak amplitude, and the full width at half maximum of the echo, were fed through a fully connected layer to produce ten scale and ten shift values. These then modulated the ten feature channels of the CNN just before its final convolutional stage. In effect, the domain knowledge no longer waited passively at the classification head; it actively reshaped how the network perceived the raw signal.</p>
<p>The feature set itself was carefully revised from prior work. The team dropped the skip count, the number of back-wall reflections in the scan plan, because it reflects inspection geometry rather than defect character. They merged the binary inside-or-outside indicator with absolute distance into a single signed distance ratio normalized by the local weld width, producing a dimensionless descriptor that remains meaningful across welds of different sizes. Peak amplitude was added to capture reflection strength, which differs between strongly reflecting planar flaws and weakly reflecting volumetric ones. The six defect types were grouped into two broad classes: volumetric defects, comprising slag and porosity, and linear defects, comprising lack of fusion, cracks, incomplete penetration, and undercut. Undercut, though geometrically a surface notch, produces a sharp, narrow specular echo that behaves like a linear defect in the ultrasound data, so it joined that class.</p>
<p>Evaluation was deliberately stringent. Instead of randomly shuffling individual A-scans, which would let near-identical neighboring signals from the same defect leak between training and test sets and inflate performance, the team partitioned data at the level of complete single-sided inspection files, 80 in total, in tenfold cross-validation. They also trained on 196 labeled regions of interest, including normal regions deliberately covering geometry signals and noise, and addressed class imbalance with a class-weighted focal loss. The results were clear-cut. Preprocessing alone lifted mean accuracy by 6.35 percentage points and macro F1 by 10.98 points, driven largely by fewer normal signals being mistaken for defects. The FiLM layer, nearly useless without preprocessing because its input features were corrupted by noise, added a further gain when combined with it, raising volumetric defect accuracy by 8.01 percentage points, from 58.34 to 78.71 percent, the weakest class in the entire problem.</p>
<p>The binary confusion matrices tell perhaps the most compelling story. Without preprocessing, between 30 and 33 percent of normal signals were falsely flagged as defects, a rate that would drown any inspector in false alarms. With preprocessing, that false positive rate collapsed to 2.30 percent, while the false negative rate fell below half a percent. Merging the two defect classes into a single defect-versus-normal decision, the full framework achieved 99.26 percent accuracy, meaning the system can reliably and rapidly localize defect-bearing regions across an entire inspection before any human looks at the data.</p>
<p>The researchers are careful about scope. Their specimens all used a common scan plan with near-normal incidence and a single skip, and they note that extending the framework to other thicknesses, materials, groove geometries, and probe configurations, or to small-radius cylindrical welds requiring different scan strategies, will demand additional representative data. They also anticipate that larger and more varied datasets will allow finer-grained classification into more than three classes, using additional features such as depth ratio and neighboring A-scans. Even so, the study demonstrates a principle with broad implications for safety-critical industries: when deep learning is fed physics-informed preprocessing and conditioned on the same physical features experts use, it does not merely mimic inspection, it accelerates it. For aerospace maintenance, where every hour of inspection time carries real cost and every missed flaw carries real risk, an algorithm that can triage ultrasound data with expert-level reliability is not a laboratory curiosity. It is the beginning of a new division of labor, in which machines screen the flood of echoes and humans spend their expertise where it matters most.</p>
<p><strong>Subject of Research:</strong> Deep learning classification of weld defects from phased-array ultrasonic testing signals</p>
<p><strong>Article Title:</strong> Weld Defect Classification in Phased-Array Ultrasonic Testing Using a CNN with FiLM Layer</p>
<p><strong>Article References:</strong> Lee, J., Chia, C. C., &amp; Lee, J.-R. (2026). Weld Defect Classification in Phased-Array Ultrasonic Testing Using a CNN with FiLM Layer. <em>International Journal of Aeronautical and Space Sciences</em>. <a href="https://doi.org/10.1007/s42405-026-01251-2" rel="noopener noreferrer">https://doi.org/10.1007/s42405-026-01251-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s42405-026-01251-2" rel="noopener noreferrer">10.1007/s42405-026-01251-2</a></p>
<p><strong>Keywords:</strong> phased-array ultrasonic testing, weld inspection, deep learning, convolutional neural network, FiLM layer, non-destructive testing, defect classification, aerospace safety, signal preprocessing, class imbalance, focal loss, KAIST</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233914</post-id>	</item>
		<item>
		<title>Tiny Transformer Offers Early Warning Against Stealthy Attacks on Industrial IoT</title>
		<link>https://scienmag.com/tiny-transformer-offers-early-warning-against-stealthy-attacks-on-industrial-iot/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 01:31:42 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced persistent threats]]></category>
		<category><![CDATA[AI in industrial network threat monitoring]]></category>
		<category><![CDATA[AI-driven intrusion detection for industrial networks]]></category>
		<category><![CDATA[CICAPT-IIoT dataset]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[Compact AI models for resource-constrained devices]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[Early detection of advanced persistent threats]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[Edge computing security in industrial IoT]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[industrial IoT]]></category>
		<category><![CDATA[Industrial IoT cybersecurity]]></category>
		<category><![CDATA[intrusion detection]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Machine learning for IoT attack prevention]]></category>
		<category><![CDATA[Multi-stage cyberattack detection in industrial systems]]></category>
		<category><![CDATA[provenance data]]></category>
		<category><![CDATA[self-attention]]></category>
		<category><![CDATA[Self-attention mechanisms in cyber threat detection]]></category>
		<category><![CDATA[Sequence modeling in industrial cybersecurity]]></category>
		<category><![CDATA[Small transformer models for edge device security]]></category>
		<category><![CDATA[Stealthy cyberattack identification using transformers]]></category>
		<category><![CDATA[Transformer]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224886</guid>

					<description><![CDATA[Researchers in Beijing have built a 0.27-megabyte transformer model that detects multi-stage APT attacks in industrial IoT telemetry with high precision and microsecond latency, positioning lightweight attention-based sequence modeling as a calibrated early-warning layer for edge defenses.]]></description>
										<content:encoded><![CDATA[<p>Industrial systems that once ran in isolation are now stitched into networks of connected sensors, controllers, and gateways, and that connectivity has opened the door to one of the most dangerous categories of cyberattack: the advanced persistent threat, or APT. These intrusions are patient, multi-stage campaigns in which an adversary quietly maps a network, escalates privileges, and moves laterally toward critical assets, all while hiding inside enormous streams of ordinary operational telemetry. A new study published in the International Journal of Machine Learning and Cybernetics by Ramadhani Zuberi Nyangusi and Hongsong Chen of the University of Science and Technology Beijing tackles this problem with an unusually small piece of artificial intelligence: a compact transformer model designed to run on the resource-starved edge devices that guard industrial Internet of Things deployments.</p>
<p>The appeal of transformer architectures in cybersecurity is easy to understand. Since the landmark 2017 paper &#8220;Attention Is All You Need,&#8221; self-attention mechanisms have transformed natural language processing and, more recently, sequence modeling in security applications, because they can weigh the relationships between events in a sequence regardless of how far apart those events occur. For APT detection, that matters enormously. An attacker&#8217;s footprint is rarely a single anomalous packet; it is a chain of individually unremarkable actions whose significance emerges only when they are read together. Larger transformer models, however, carry millions of parameters and demand memory and compute budgets that typical IIoT gateways simply cannot provide, which is why many high-performing research models never leave the laboratory.</p>
<p>Nyangusi and Chen&#8217;s answer is a deliberately stripped-down transformer. Their framework uses just two encoder layers and two attention heads, with a model dimension of 64 and a feed-forward dimension of 128. The result is a network with only 69,057 trainable parameters and an approximate model size of 0.27 megabytes, small enough to plausibly sit on edge hardware rather than requiring a cloud round-trip for every decision. The design philosophy is context-awareness at minimal cost: rather than analyzing entire provenance graphs or long event histories, the system organizes provenance events into short temporal windows, allowing the attention mechanism to capture local temporal behavior while keeping the computational footprint tiny.</p>
<p>Class imbalance is the second central challenge the researchers confront head-on. In real industrial telemetry, malicious events are vanishingly rare compared with benign ones, and models trained naively on such data tend to achieve high accuracy while missing most actual attacks. The framework therefore employs focal loss, an imbalance-aware optimization objective that down-weights easy, well-classified examples and concentrates learning effort on the difficult minority cases that matter most. Just as importantly, the authors are explicit about methodology hygiene: decision thresholds are selected on a validation set before final testing, avoiding the test-set-driven calibration that can silently inflate reported performance in detection research.</p>
<p>The evaluation rests on the CICAPT-IIoT dataset, a publicly available provenance-based APT attack dataset for IIoT environments released by the Canadian Institute for Cybersecurity at the University of New Brunswick. Provenance data records the causal history of system activity, which makes it a natural substrate for spotting multi-stage intrusions. Across five random seeds, the best sequence-level configuration used a four-event temporal window and achieved a malicious precision of 0.8547 plus or minus 0.0205, a recall of 0.5025 plus or minus 0.0353, an F1-score of 0.6321 plus or minus 0.0263, a ROC-AUC of 0.8840 plus or minus 0.0131, and a PR-AUC of 0.5909 plus or minus 0.0193. Reporting across multiple seeds and including variance, rather than a single best run, gives these numbers a credibility that single-shot benchmarks often lack.</p>
<p>Those figures tell an honest and nuanced story. Precision above 0.85 means that when the model raises an alarm, it is right the vast majority of the time, which is exactly what operators of critical infrastructure need, since false alarms in a factory or power grid carry real operational costs. Recall near 0.50, by contrast, means the model catches roughly half of malicious sequences, and the modest PR-AUC reflects the brutal arithmetic of extreme class imbalance. The authors do not paper over this trade-off. Instead, they position the framework explicitly as what it is: a compact, calibrated early-warning component for IIoT APT detection, not a universal replacement for all classical classifiers. In a layered defense, a lightweight sensor that reliably flags high-confidence threats at the edge has clear value even if deeper analysis systems handle the harder cases.</p>
<p>The resource profiling is where the work becomes genuinely striking for anyone thinking about deployment. CPU inference latency measured 0.0609 plus or minus 0.0002 milliseconds per four-event window, a figure so low that the model could, in principle, evaluate thousands of windows per second on modest hardware. Combined with the 0.27-megabyte footprint, this suggests the framework could be embedded directly into gateways, industrial PCs, or even constrained embedded devices, screening provenance streams continuously and escalating only suspicious sequences to heavier backend analysis. That division of labor, tiny models at the edge and heavyweight forensics in the core, is increasingly seen as the realistic architecture for securing sprawling industrial estates.</p>
<p>To test whether the approach generalizes beyond its home dataset, the researchers performed an external validation on Windows-APT 2025, a dataset of APT-inspired attack scenarios on Windows systems. The same temporal-window pipeline showed it could transfer to ATT&amp;CK-mapped Windows host-alert detection, suggesting the design is not merely tuned to the quirks of one provenance dataset. The authors are careful to note that this second setting uses a different telemetry source and a proxy-label structure, so the transfer result is suggestive rather than definitive. Even so, the ability of one lightweight pipeline to operate across both IIoT provenance data and Windows host alerts hints at a portable pattern for early-stage threat detection across heterogeneous environments.</p>
<p>The study also situates itself within a rapidly crowding field. Recent years have produced transformer-based intrusion detectors, hybrid CNN-BiLSTM and Swin-transformer hybrids, diffusion-transformer models for imbalanced IoT learning, provenance-graph frameworks with masked representation learning, and knowledge-distillation approaches aimed at explainable detection. Many of these achieve strong classification metrics, but the Beijing team argues that too few provide evidence of deployment feasibility under edge-oriented resource constraints, and many gloss over the precision-recall trade-off that severe imbalance imposes. By publishing parameter counts, model sizes, latency figures, and seed-level variance alongside detection metrics, this work offers a template for how lightweight security AI should be evaluated: not just how well it detects, but whether it can actually run where the threats arrive.</p>
<p>For the operators of factories, utilities, and critical infrastructure, the takeaway is pragmatic rather than sensational. Advanced persistent threats will not be defeated by a single algorithm, and a detector that catches half of malicious sequences is not a silver bullet. But a 69,000-parameter model that fits in a fraction of a megabyte, responds in microseconds, and delivers high-precision alerts from raw provenance windows represents a meaningful building block for defense in depth. As industrial networks grow and attackers grow more patient, the future of cybersecurity may depend less on ever-larger models in distant data centers and more on swarms of small, fast, honest sentinels watching quietly at the edge, and this research shows exactly what such sentinels can, and cannot, yet do.</p>
<p><strong>Subject of Research:</strong> Lightweight transformer-based detection of advanced persistent threats in industrial Internet of Things environments</p>
<p><strong>Article Title:</strong> A lightweight transformer-based framework for context-aware APT detection in industrial IoT</p>
<p><strong>Article References:</strong> Nyangusi, R. Z., &amp; Chen, H. (2026). A lightweight transformer-based framework for context-aware APT detection in industrial IoT. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 484. <a href="https://doi.org/10.1007/s13042-026-03324-w" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03324-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03324-w" rel="noopener noreferrer">10.1007/s13042-026-03324-w</a></p>
<p><strong>Keywords:</strong> advanced persistent threats, industrial IoT, transformer, intrusion detection, edge computing, provenance data, focal loss, class imbalance, CICAPT-IIoT dataset, self-attention, cybersecurity, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224886</post-id>	</item>
		<item>
		<title>AI Listens to the Womb: Deep Learning Model Predicts Preterm Birth from Uterine Electrical Signals</title>
		<link>https://scienmag.com/ai-listens-to-the-womb-deep-learning-model-predicts-preterm-birth-from-uterine-electrical-signals/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 21:03:37 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in maternal-fetal medicine]]></category>
		<category><![CDATA[AI in obstetrics]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[continuous wavelet transform]]></category>
		<category><![CDATA[convolutional neural network]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning in biomedical signal processing]]></category>
		<category><![CDATA[deep learning models for pregnancy]]></category>
		<category><![CDATA[early detection of preterm delivery]]></category>
		<category><![CDATA[electrohysterogram]]></category>
		<category><![CDATA[electrohysterogram (EHG) monitoring]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[neural network-based pregnancy monitoring]]></category>
		<category><![CDATA[noninvasive monitoring]]></category>
		<category><![CDATA[noninvasive preterm labor detection]]></category>
		<category><![CDATA[personalized pregnancy risk scoring]]></category>
		<category><![CDATA[predictive medicine]]></category>
		<category><![CDATA[pregnancy risk assessment tools]]></category>
		<category><![CDATA[Preterm birth]]></category>
		<category><![CDATA[Preterm birth prediction]]></category>
		<category><![CDATA[Signal Processing]]></category>
		<category><![CDATA[uterine electrical signal analysis]]></category>
		<category><![CDATA[uterine electromyography]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212491</guid>

					<description><![CDATA[A new deep learning model called CWT-AuxNet converts wavelet-based time-frequency representations of uterine electrical recordings into individualized preterm birth risk scores, achieving an AUC of 0.932 at the patient level.]]></description>
										<content:encoded><![CDATA[<p>Every year, an estimated 15 million babies around the world are born preterm, and complications of prematurity remain a leading cause of death among newborns. For clinicians, the central challenge has always been anticipation: identifying, weeks in advance, which pregnancies will end too soon. Now a team of researchers in China has unveiled a deep learning system that reads the electrical chatter of the pregnant uterus and converts it into a personalized risk score, achieving strikingly high accuracy in distinguishing women who will deliver early from those who will carry to term. The study, published in Medical &amp; Biological Engineering &amp; Computing, describes a model called CWT-AuxNet, which reached an area under the receiver operating characteristic curve, or AUC, of 0.932 at the level of individual patients, a figure that places it among the strongest noninvasive preterm birth predictors reported to date.</p>
<p>The signal at the heart of the work is the electrohysterogram, or EHG, a recording of the electrical activity that coordinates contractions of the uterine muscle. Electrodes placed on the abdominal wall pick up these faint potentials noninvasively, much as an electrocardiogram captures the heart&#8217;s rhythm. Researchers have known for decades that the EHG changes character as pregnancy progresses: as labor approaches, the electrical bursts that sweep across the uterus shift toward lower frequencies and become more synchronized, a physiological signature of the muscle preparing for coordinated contractions. What has been harder is turning that knowledge into a reliable clinical test, because EHG recordings are long, noisy, subtly different between patients, and, crucially, scarce. Datasets of labeled recordings from women who later delivered preterm are small, which has made it difficult to train the data-hungry deep neural networks that have transformed other areas of medicine.</p>
<p>That scarcity and complexity, the authors argue, is precisely why most previous efforts relied on conventional machine learning pipelines, in which experts hand-craft features from the signal, such as entropy measures, frequency-band power ratios, or nonlinear descriptors, and then feed them to a classifier. Such approaches work, but they inherit the biases and blind spots of the features humans choose to extract. Deep learning, by contrast, can in principle discover discriminative patterns directly from raw data. The new study tackles the obstacles that have kept end-to-end deep learning out of the EHG field with a three-part strategy: a wavelet-based representation of the signal, an auxiliary feature that injects established physiological knowledge, and a training objective engineered to cope with severe class imbalance.</p>
<p>The first ingredient is the continuous wavelet transform, or CWT, a mathematical technique that decomposes a signal into a family of wavelets stretched and shifted across different scales. Unlike the Fourier transform, which tells you which frequencies are present in a recording but discards when they occurred, the wavelet transform preserves both time and frequency information simultaneously. Applied to an EHG recording, the CWT produces a two-dimensional image, a scalogram, in which the horizontal axis is time, the vertical axis is frequency, and brightness encodes the local energy of the signal. This transformation is a natural fit for uterine electrical activity, whose relevant patterns, such as the gradual downward shift of contraction-related frequencies, unfold over time. It also converts a one-dimensional signal into an image-like input, allowing the researchers to exploit the full power of convolutional neural networks, the same architecture that revolutionized image recognition.</p>
<p>On top of this time-frequency representation, the team added what they call an auxiliary feature: the peak amplitude, or PA, of the normalized power spectrum in the low-frequency band. This measure, highlighted in recent work as an effective standalone predictor of premature birth, captures how strongly the EHG energy is concentrated at the frequencies associated with preterm labor. By feeding this engineered feature into the network alongside the learned wavelet representations, CWT-AuxNet blends data-driven pattern discovery with domain knowledge accumulated over years of EHG research. The architecture itself is multibranch and convolutional, meaning parallel streams of filters process different aspects of the input before their outputs are merged, enabling the model to extract fine-grained features at the level of short signal windows rather than forcing a single judgment on an entire recording.</p>
<p>The third innovation addresses a problem that has quietly inflated results across the EHG literature: imbalance. Preterm deliveries are, fortunately, the minority outcome, so datasets contain far more term recordings than preterm ones. Naively trained classifiers tend to default to predicting the majority class, and oversampling techniques such as SMOTE, which synthesize artificial minority examples, have been shown in critical reanalyses to produce overly optimistic performance estimates. Instead of resampling the data, CWT-AuxNet uses a cost-sensitive loss function built on focal loss, a technique originally developed for dense object detection in computer vision. Focal loss down-weights the contribution of easy, well-classified examples and concentrates the gradient signal on hard, ambiguous ones, while class-specific weighting compensates for the rarity of preterm cases. The result is a network that learns to care about the minority class without fabricating synthetic data.</p>
<p>Because the model produces predictions for individual windows of the EHG recording rather than a single verdict per patient, the researchers designed a two-tier decision strategy. At the window level, the network assigns a risk score to each short segment of signal, capturing fine-grained fluctuations in uterine electrical behavior. At inference time, these window-level outputs are aggregated into a user-level decision through a dedicated strategy, yielding one individualized assessment of preterm risk per patient. This hierarchical design mirrors how a clinician might reason: noticing suspicious moments in a long monitoring session and then weighing them together to form an overall judgment. It also makes the system more robust, since a single noisy segment cannot dominate the final decision.</p>
<p>The performance numbers tell a compelling story. CWT-AuxNet achieved an AUC of 0.741 at the window level, indicating strong discrimination even when judging brief, isolated segments of signal where information is inherently limited. When window-level predictions were aggregated to the user level, performance climbed to an AUC of 0.932, meaning the model ranked individual patients&#8217; preterm risk with high reliability. In head-to-head comparisons, the model consistently outperformed both traditional machine learning baselines built on hand-crafted features and earlier deep learning approaches. The authors also employed gradient-based visualization techniques, in the spirit of Grad-CAM, to probe which regions of the time-frequency representations drove the network&#8217;s decisions, offering a degree of interpretability that is essential for any technology hoping to enter prenatal care.</p>
<p>The implications reach beyond a single benchmark. Preterm birth prediction has long been dominated by clinical measures with limited predictive power when applied early: cervical length measured by ultrasound and fetal fibronectin testing, for example, show modest predictive value in threatened preterm labor. An EHG-based approach offers something different, a continuous, noninvasive window into the physiological maturation of the uterus itself, potentially usable in routine prenatal visits with standard surface electrodes. The study was supported by the National Natural Science Foundation of China, and the research team, led by co-first authors Xinliang Wen and Shengnan Zhuan with corresponding authors Lai Jiang and Xu Zhang, spans the University of Science and Technology of China, Bengbu Medical University, and the First Affiliated Hospital of USTC, combining expertise in microelectronics, life sciences, and obstetrics.</p>
<p>Challenges remain before CWT-AuxNet or any successor reaches the delivery ward. EHG datasets are still small and drawn largely from a limited number of recording centers, and the field has been burned before by methods that excelled on a single benchmark but failed to generalize. External validation on independent, multi-center cohorts, prospective clinical studies, and careful attention to calibration of risk scores will all be necessary. Yet the study marks a meaningful shift in how the problem is framed: rather than asking humans to define what distinguishes a preterm EHG recording, the wavelet-driven network learns those distinctions itself, guided by physiological priors and trained with an objective that respects the reality of imbalanced clinical data. If that approach holds up in the clinic, the faint electrical whispers of the uterus could become one of obstetrics&#8217; most valuable early warning systems, giving mothers and doctors the most precious resource of all, time.</p>
<p><strong>Subject of Research:</strong> Deep learning prediction of preterm birth from electrohysterogram signals</p>
<p><strong>Article Title:</strong> A deep learning method for preterm birth prediction using wavelet representations and auxiliary features from electrohysterogram</p>
<p><strong>Article References:</strong> Wen, X., Zhuan, S., Gao, X., Jiang, L., &amp; Zhang, X. (2026). A deep learning method for preterm birth prediction using wavelet representations and auxiliary features from electrohysterogram. <em>Medical &amp;amp; Biological Engineering &amp;amp; Computing</em>. <a href="https://doi.org/10.1007/s11517-026-03602-3" rel="noopener noreferrer">https://doi.org/10.1007/s11517-026-03602-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11517-026-03602-3" rel="noopener noreferrer">10.1007/s11517-026-03602-3</a></p>
<p><strong>Keywords:</strong> preterm birth, electrohysterogram, deep learning, continuous wavelet transform, convolutional neural network, focal loss, class imbalance, uterine electromyography, predictive medicine, signal processing, noninvasive monitoring, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212491</post-id>	</item>
		<item>
		<title>New AI Method Learns From Rare Positives Hidden in Unlabeled Data</title>
		<link>https://scienmag.com/new-ai-method-learns-from-rare-positives-hidden-in-unlabeled-data/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 01:46:39 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancing reliability of AI models]]></category>
		<category><![CDATA[assumptions in positive-unlabeled learning]]></category>
		<category><![CDATA[challenges of unlabeled data]]></category>
		<category><![CDATA[class imbalance in machine learning]]></category>
		<category><![CDATA[data mining]]></category>
		<category><![CDATA[data mining for hidden positive signals]]></category>
		<category><![CDATA[financial misstatement detection]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[fraud detection]]></category>
		<category><![CDATA[handling missing and mislabeled data]]></category>
		<category><![CDATA[imbalanced classification]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in medicine and finance]]></category>
		<category><![CDATA[new algorithms for unbalanced datasets]]></category>
		<category><![CDATA[novel methods for PU learning]]></category>
		<category><![CDATA[positive-unlabeled learning]]></category>
		<category><![CDATA[PU learning]]></category>
		<category><![CDATA[rare positive example detection]]></category>
		<category><![CDATA[risk estimation]]></category>
		<category><![CDATA[SAR assumption]]></category>
		<category><![CDATA[SCAR assumption]]></category>
		<category><![CDATA[semi-supervised learning in machine learning]]></category>
		<category><![CDATA[weakly supervised learning]]></category>
		<category><![CDATA[XGBoost]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=209621</guid>

					<description><![CDATA[Researchers have developed a focused positive-unlabeled learning method that uses focal loss to achieve state-of-the-art performance on severely imbalanced datasets under realistic labeling assumptions.]]></description>
										<content:encoded><![CDATA[<p>Machine learning has transformed fields from medicine to finance, but one stubborn problem continues to undermine its reliability in the real world: most of the data we care about is either missing or mislabeled. In many practical settings, practitioners have only a handful of confirmed positive examples, surrounded by a vast ocean of unlabeled data that contains a hidden mixture of positives and negatives. This scenario, known as positive-unlabeled (PU) learning, has been studied intensively for two decades, yet most existing methods quietly assume that the classes are roughly balanced and that labeled examples are drawn randomly from the positive population. A new study published in Data Mining and Knowledge Discovery by Elias Zavitsanos and Georgios Paliouras of the Institute of Informatics and Telecommunications at NCSR Demokritos in Greece confronts those assumptions head-on and delivers a method that thrives precisely where others falter.</p>
<p>The core difficulty is easy to state but hard to solve. In classical binary classification, an algorithm sees both labeled positives and labeled negatives and learns a decision boundary between them. In PU learning, negative labels simply do not exist. Instead, the algorithm receives a small set of confirmed positives and a large unlabeled pool. Because the unlabeled pool is dominated by negatives but contaminated with positives, treating every unlabeled example as negative injects label noise into training. Traditional approaches either try to identify reliable negatives from the unlabeled set before training, or they incorporate assumptions about the class prior, the underlying proportion of positive examples in the data. Both strategies become brittle when the dataset is severely imbalanced, meaning positives may account for only one to a few percent of all examples.</p>
<p>Zavitsanos and Paliouras observed that in imbalanced PU data, the mathematics of existing risk estimators works against the practitioner. The estimators used by state-of-the-art methods, such as the unbiased PU estimator (uPU) and its non-negative successor (nnPU), contain a term weighted by the positive class prior. When that prior is tiny, the contribution of the few labeled positives nearly vanishes, and the training objective is dominated by the unlabeled data. The model therefore learns to be excellent at recognizing negatives, and mediocre at catching the rare positives that matter most. Matters worsen under the Probabilistic Gap assumption, a realistic refinement of labeling theory in which positive examples that resemble negatives are the least likely to have been labeled. The very positives a model most needs to learn from are the ones most likely to be hidden in the unlabeled pool.</p>
<p>The researchers&#8217; answer is a new empirical risk estimator they call iFPU, for imbalanced focused PU learning. The key ingredient is focal loss, a function originally developed in computer vision for dense object detection, where positive targets are similarly swamped by an enormous number of easy negatives. Focal loss reshapes the standard cross-entropy objective by multiplying each example&#8217;s loss by a factor that shrinks as the model&#8217;s confidence in the correct answer grows. Easy examples, which the model already classifies correctly with high probability, contribute almost nothing to the gradient. Hard examples, sitting near the decision boundary, receive exponentially amplified weight, controlled by a focusing parameter gamma. The result is that training concentrates its capacity on exactly the boundary cases that define the Probabilistic Gap.</p>
<p>Incorporating focal loss into a PU risk estimator is technically delicate. The authors derive a non-negative risk estimator that combines three terms: the focal loss of the labeled positives treated as positive, an unbiased correction subtracting the focal loss the positives would incur if treated as negative, and the focal loss of the unlabeled pool treated as negative. A max operator clamps the correction at zero to prevent the notorious problem of negative empirical risk, which causes overfitting in neural networks trained with earlier unbiased estimators. When the correction term goes negative within a training mini-batch, the method performs a step of gradient ascent instead, deliberately nudging the model away from overfitting that batch. The whole procedure slots into standard stochastic optimization, requires no preprocessing, resampling, or manipulation of the data, and can be attached to essentially any classifier trained by cost minimization, from multilayer perceptrons to gradient-boosted trees.</p>
<p>The authors also supply a rigorous theoretical analysis. They prove that the population-level iFPU risk is identical to the focal classification risk under full supervision, for any labeling mechanism, by virtue of a mixture identity between the unlabeled distribution and the class-conditional densities. They further characterize the bias that appears when labeled positives are not selected completely at random, showing that this bias depends only on the labeling propensity and is orthogonal to the choice of loss function. Under the SCAR assumption, they establish an estimation error bound that vanishes at the standard statistical rate, proportional to the inverse square root of the number of labeled positives and unlabeled examples combined, using Rademacher complexity tools. Crucially, they show focal loss is Lipschitz continuous and bounded on a restricted score range, which makes these guarantees possible. Although focal loss itself introduces a deliberate bias by reweighting errors, it remains classification-calibrated, meaning the Bayes-optimal classifier under focal loss coincides with the Bayes-optimal classifier under zero-one loss.</p>
<p>Empirically, the method was tested on 14 publicly available benchmark datasets for imbalanced binary classification, spanning positive rates from roughly 1 to 14 percent. The experimental design was deliberately demanding: positive examples were progressively hidden in the unlabeled pool at rates of 25, 50, and 75 percent, under both the SCAR and the more realistic SAR labeling assumptions, generating 840 experimental runs in total. Notably, the authors avoided hyperparameter tuning altogether, using the recommended default gamma of 3, because tuning in PU settings is itself fraught with assumptions due to the absence of negatively labeled validation data. Under SCAR, iFPU outperformed the neural risk estimators uPU, nnPU, and i-NNPU, and when paired with an XGBoost classifier it matched or exceeded strong competitors including the two-step NNIF anomaly-detection method, PU Hellinger Decision Trees, the label-bias estimation method LBE, and the SAREM expectation-maximization framework, trailing only slightly behind the PU Hellinger Random Forest ensemble.</p>
<p>The picture shifts decisively in favor of iFPU under the more realistic SAR assumption, where positives resembling negatives are less likely to be labeled. Here the competing methods degrade noticeably, while iFPU maintains its performance, and in the hardest scenario, with only 25 percent of positives labeled, it ranks first overall, surpassing the PU Hellinger Random Forest by five percentage points in PR-AUC. Statistical tests confirmed significant differences among methods, and paired comparisons showed iFPU significantly outperforming all alternatives in the most challenging configuration. A sensitivity analysis demonstrated that the method remains robust even when the class prior is misspecified by factors of two or four in either direction, degrading meaningfully only when the prior is severely underestimated. On the three most imbalanced datasets, Cover, Poker, and Satellite, iFPU showed clearly higher mean and median PR-AUC than its strongest rival.</p>
<p>To demonstrate real-world value, the researchers applied their method to financial misstatement detection, a problem where PU data arise naturally. Auditors and regulators typically discover accounting misstatements years after reports are filed, and often only a fraction of misstatements have been identified when a model is trained. Using data on publicly traded US companies spanning 2000 to 2014, with 47,086 firm-year records described by 28 financial indices and derived accounting features, the team simulated realistic detection delays in which approximately 40 percent of positive training labels were missing at training time. Misstatements, whether deliberately concealed frauds or subtle errors resembling normal accounts, fit the Probabilistic Gap assumption perfectly. Models built on a TabTransformer architecture with gated MLP modules and equipped with the calibrated iFPU risk achieved R-precision scores roughly three times higher than prior baselines such as RUSBoost, and outperformed both earlier specialized models and the PU Hellinger Random Forest, setting a new state of the art in this application.</p>
<p>The significance of this work extends beyond any single benchmark. By combining a principled risk-estimation framework with a loss function engineered for imbalance, the authors show that PU learning can be made practical in exactly the conditions that dominate high-stakes applications, from disease gene identification to fraud detection, where positives are rare, partially labeled, and deceptively similar to negatives. The method&#8217;s plug-and-play compatibility with modern classifiers, its robustness to prior misspecification, and its theoretical grounding distinguish it from heuristic preprocessing pipelines. The authors point to extensions toward semi-supervised and multi-class settings, integration with pre-trained tabular foundation models such as TabPFN, and output calibration via temperature scaling as promising future directions. For now, iFPU offers practitioners a rare commodity in weakly supervised machine learning: a method whose assumptions match reality rather than convenience.</p>
<p><strong>Subject of Research:</strong> Positive-unlabeled machine learning from highly imbalanced datasets</p>
<p><strong>Article Title:</strong> Focused PU learning from imbalanced data</p>
<p><strong>Article References:</strong> Focused PU learning from imbalanced data. (n.d.). <a href="https://doi.org/10.1007/s10618-026-01264-1" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01264-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01264-1" rel="noopener noreferrer">10.1007/s10618-026-01264-1</a></p>
<p><strong>Keywords:</strong> PU learning, imbalanced classification, weakly supervised learning, focal loss, risk estimation, machine learning, fraud detection, financial misstatement detection, XGBoost, SCAR assumption, SAR assumption, data mining</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">209621</post-id>	</item>
		<item>
		<title>AI Faces a Hard Test: Detecting Rare Martensite in Steel Micrographs</title>
		<link>https://scienmag.com/ai-faces-a-hard-test-detecting-rare-martensite-in-steel-micrographs/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:03:07 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[application of deep learning in steel microstructure analysis]]></category>
		<category><![CDATA[challenges of AI in detecting rare phases]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[data imbalance impact on microstructure detection]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning for micrograph analysis]]></category>
		<category><![CDATA[focal loss]]></category>
		<category><![CDATA[GLCM]]></category>
		<category><![CDATA[Grad-CAM]]></category>
		<category><![CDATA[limitations of AI in materials quality assurance]]></category>
		<category><![CDATA[machine learning class imbalance in metallography]]></category>
		<category><![CDATA[Martensite]]></category>
		<category><![CDATA[Martensite identification in steel micrographs]]></category>
		<category><![CDATA[microstructural phase classification in high carbon steel]]></category>
		<category><![CDATA[microstructure classification]]></category>
		<category><![CDATA[rare microstructure detection in steel]]></category>
		<category><![CDATA[rare-phase detection]]></category>
		<category><![CDATA[ResNet50]]></category>
		<category><![CDATA[ResNet50 application in metallurgical microstructure classification]]></category>
		<category><![CDATA[t-SNE]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[transfer learning in materials science]]></category>
		<category><![CDATA[ultra-high carbon steel]]></category>
		<category><![CDATA[use of SEM micrographs for AI training]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202416</guid>

					<description><![CDATA[A new baseline study shows that deep learning can learn meaningful microstructural information from severely imbalanced ultra-high carbon steel images, but reliable detection of the rare Martensite phase remains constrained by data scarcity and class overlap.]]></description>
										<content:encoded><![CDATA[<p>Deep learning has transformed how scientists classify materials, but a new study shows just how difficult the task becomes when the most important feature is also the rarest. In research published in the Journal of Materials Science: Metallurgy, Ehsun Saeed evaluated whether a transfer-learning framework could reliably detect Martensite, a hard and brittle phase, in ultra-high carbon steel micrographs when that phase accounted for a mere 3.75 percent of the dataset. The work was conceived as a baseline diagnostic rather than a claim of industrial readiness, and its findings carry a candid message about the limits of artificial intelligence in metallurgical quality assurance.</p>
<p>The dataset consisted of 961 scanning electron microscopy micrographs spanning seven microstructural classes, drawn from the Ultra-High Carbon Steel micrograph collection maintained by the National Institute of Standards and Technology. Spheroidite dominated with 374 images, while Martensite was represented by only 36. This tenfold disparity in representation creates exactly the kind of class imbalance that undermines conventional machine learning, where optimisation naturally favours majority classes and can produce seemingly respectable accuracy while missing the minority class entirely.</p>
<p>Saeed adopted a frozen ResNet50 transfer-learning architecture, using ImageNet pre-trained weights as a fixed feature extractor and retraining only the newly added classification layers. Images were resized to 224 by 224 pixels, normalised, and augmented online with random rotations, flips, and shifts during training. A stratified train-test split of 80 to 20 preserved class proportions, leaving just seven Martensite images in the test partition. Baseline training used the Adam optimiser with categorical cross-entropy loss for 30 epochs, and two reference classifiers, a random predictor and a majority-class predictor, served as lower bounds.</p>
<p>The baseline model achieved a peak validation accuracy of 52.8 percent, well above the random baseline of 14.3 percent and the majority-class baseline of 38.9 percent. Martensite detection, however, remained modest, with precision of 0.600, recall of 0.429, and an F1-score of 0.500. Because overall accuracy can mask poor minority-class recognition, the study also reported balanced accuracy, macro-averaged metrics, and the Matthews Correlation Coefficient, all of which provide a more honest picture under severe imbalance.</p>
<p>To probe robustness, stratified five-fold cross-validation was performed, yielding an average accuracy of 0.412, a macro F1-score of 0.133 plus or minus 0.009, and an MCC of 0.134 plus or minus 0.029. These figures reveal weak multi-class generalisation and suggest that the principal challenge is not instability across partitions but genuine class overlap and insufficient minority-class representation. The gap between the single split and the cross-validated results further indicates that performance estimates are sensitive to dataset composition.</p>
<p>Two imbalance-aware strategies were then tested. Class-weighted loss, with weights inversely proportional to class frequencies, boosted Martensite recall dramatically to 0.857 but crashed overall accuracy to 29.0 percent, illustrating a flood of false positives and an overcompensation for the rarity of the minority class. Focal loss, which down-weights easily classified majority samples, fared better, delivering the strongest overall performance with 56.0 percent accuracy, 36.5 percent balanced accuracy, an MCC of 0.388, and improved Martensite precision of 0.750 while holding recall steady. The results show that loss-function design can shift the balance between sensitivity and precision, but cannot conjure data that does not exist.</p>
<p>A binary Martensite-versus-non-Martensite experiment offered further insight, achieving an AUROC of 0.908 and an MCC of 0.527, substantially stronger than the multi-class formulation. Threshold optimisation raised recall from 28.6 percent to 57.1 percent while preserving precision at 80.0 percent. This suggests that ambiguity among seven overlapping classes compounds the difficulty, and that application-specific threshold tuning can extract meaningful rare-phase discrimination from the same underlying model.</p>
<p>Interpretability and confounding analyses rounded out the study. Grad-CAM visualisations showed that activation maps concentrated on microstructural regions rather than on scale bars or image borders, and border-crop robustness tests confirmed that removing up to 15 percent of the image perimeter left performance largely unchanged. t-SNE projection of the penultimate-layer features, however, revealed that Martensite samples were dispersed throughout the feature space with no distinct cluster, overlapping heavily with Pearlite, Spheroidite, and Network microstructures. A magnification audit exposed substantial variation in imaging scale across classes, from a mean of 89 times for Network to 13,441 times for Pearlite, and magnification-only classification with a Random Forest achieved an AUROC of 0.810, indicating that imaging scale acts as a partial confounder without fully explaining model behaviour.</p>
<p>Handcrafted texture descriptors, long the workhorses of metallographic image analysis, were evaluated for comparison. Features derived from Gray-Level Co-occurrence Matrices and Local Binary Patterns, classified with Support Vector Machine and Random Forest models under stratified cross-validation, achieved performance comparable to the CNN, suggesting that informative texture signals exist in the data and that physically interpretable baselines remain valuable benchmarks in data-scarce settings.</p>
<p>The study positions itself as a reproducible baseline rather than a deployable solution, and the implications are clear. Reliable rare-phase detection in industrial metallography will require larger minority-class datasets, scale-normalised imaging protocols, quantitative microstructural descriptors fused with learned features, and specialised strategies such as few-shot learning, anomaly detection, and metric learning. Until then, the author argues, such models are best regarded as decision-support tools subject to expert review, uncertainty-aware thresholds, and external validation across instruments, compositions, and processing conditions, rather than as autonomous replacements for the trained metallurgist&#8217;s eye.</p>
<p><strong>Subject of Research:</strong> Rare-phase detection of Martensite in ultra-high carbon steel micrographs using deep learning and handcrafted texture features under severe class imbalance</p>
<p><strong>Article Title:</strong> A baseline study of rare-phase detection in ultra-high carbon steel microstructures using deep learning and handcrafted texture features</p>
<p><strong>Article References:</strong> Saeed, E. (2026). A baseline study of rare-phase detection in ultra-high carbon steel microstructures using deep learning and handcrafted texture features. <em>Journal of Materials Science: Metallurgy, 1</em>(1), Article 12. <a href="https://doi.org/10.1007/s44492-026-00012-2" rel="noopener noreferrer">https://doi.org/10.1007/s44492-026-00012-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44492-026-00012-2" rel="noopener noreferrer">10.1007/s44492-026-00012-2</a></p>
<p><strong>Keywords:</strong> ultra-high carbon steel, Martensite, deep learning, ResNet50, transfer learning, class imbalance, microstructure classification, focal loss, Grad-CAM, t-SNE, GLCM, rare-phase detection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202416</post-id>	</item>
	</channel>
</rss>
