<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>PR-AUC &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/pr-auc/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 07 Oct 2026 11:10:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>PR-AUC &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Architecture Tames Extreme Data Imbalance in Multi-Task Recommendation</title>
		<link>https://scienmag.com/new-ai-architecture-tames-extreme-data-imbalance-in-multi-task-recommendation/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 07 Oct 2026 11:10:25 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced AI models for sparse data]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[data imbalance in machine learning]]></category>
		<category><![CDATA[factorization machines]]></category>
		<category><![CDATA[gating networks]]></category>
		<category><![CDATA[H-PLE architecture for recommendation]]></category>
		<category><![CDATA[handling sparse signals in neural networks]]></category>
		<category><![CDATA[improving signal prediction in recommendation systems]]></category>
		<category><![CDATA[label imbalance]]></category>
		<category><![CDATA[machine learning for user engagement metrics]]></category>
		<category><![CDATA[Mixture of Experts]]></category>
		<category><![CDATA[multi-task learning]]></category>
		<category><![CDATA[multi-task learning challenges]]></category>
		<category><![CDATA[multi-task recommendation system]]></category>
		<category><![CDATA[negative transfer]]></category>
		<category><![CDATA[negative transfer in multi-task learning]]></category>
		<category><![CDATA[neural network architecture design for imbalanced data]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[PR-AUC]]></category>
		<category><![CDATA[progressive layered extraction]]></category>
		<category><![CDATA[rare event prediction in AI]]></category>
		<category><![CDATA[recommendation system signal prediction]]></category>
		<category><![CDATA[recommendation systems]]></category>
		<category><![CDATA[Tenrec benchmark]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=244197</guid>

					<description><![CDATA[Researchers at Keio University have developed H-PLE, a hierarchical multi-task recommendation framework that stabilizes expert routing and improves prediction of rare user engagement signals such as likes and shares under extreme label imbalance.]]></description>
										<content:encoded><![CDATA[<p>Recommendation systems face a deceptively simple question every time a user watches a video: did they like it, and did they share it? Behind that question lies one of the hardest practical problems in modern machine learning. Likes and shares are extraordinarily rare compared with ordinary views, and when a model is trained to predict several such signals at once, the abundant signal tends to drown out the scarce ones. A new study published in Applied Intelligence by Jinghao Xue of Keio University&#8217;s Graduate School of Science and Technology, together with Tianxiang Yang and Hideo Suzuki of the Department of Industrial and Systems Engineering, proposes a framework called H-PLE that confronts this imbalance head-on. The work, published on 8 September 2026 as volume 56, article 409 of the journal, demonstrates that careful architectural design can rescue the performance of sparse tasks without sacrificing the dense ones.</p>
<p>The starting point for the research is a well-known phenomenon called negative transfer. Multi-task learning rests on the appealing idea that a shared neural representation can serve several prediction tasks simultaneously, improving data efficiency because each task benefits from features learned for the others. Rich Caruana articulated this principle in 1997, and it has since become a cornerstone of industrial recommendation engines. But the bargain only holds when the tasks are roughly compatible in difficulty and data volume. When one task, such as predicting whether a user will share a video, has orders of magnitude fewer positive examples than another, such as predicting a click or a view, the shared representation is shaped almost entirely by the high-frequency task. The rare task then inherits features that are poorly suited to it, and attempts to improve it can actively degrade the common tasks.</p>
<p>Two influential architectures have dominated attempts to manage this interference. Multi-gate Mixture-of-Experts, introduced by Jiaqi Ma and colleagues at Google in 2018, maintains a pool of expert subnetworks and lets each task learn its own soft weighting over them through a gating network. Progressive Layered Extraction, or PLE, proposed by Hongyu Tang and colleagues in 2020, refined this idea by separating shared experts from task-specific experts and stacking the routing across multiple layers, so that task-specific knowledge is progressively extracted from the shared representation. Both approaches, however, rely on fully learned soft gating. As Xue and colleagues point out, under severe label imbalance the gating networks themselves become part of the problem: they are trained predominantly on gradients flowing from the high-frequency tasks, so the routing decisions drift toward serving those tasks, leaving the sparse ones with unstable expert selection and insufficiently expressive representations.</p>
<p>H-PLE, which stands for hierarchical progressive layered extraction, builds on PLE but adds three complementary design elements. The first is a hierarchical extraction structure that progressively separates shared and task-specific representations across stacked layers, giving the model a more disciplined pathway for carving out what is common from what is unique. The second, and arguably the most novel, is a raw-input-aware gating mechanism the authors call AGC. In conventional expert routing, the gate sees only learned intermediate representations, which are themselves contaminated by the imbalance problem. AGC instead injects a projection of the raw input features back into the gating decision as a residual candidate. This gives the router a stable, task-independent anchor: no matter how skewed the training gradients become, the gate can always consult an uncorrupted view of what the user and the content actually look like, which stabilizes routing precisely where supervision is sparsest.</p>
<p>The third element is a lightweight expert-level interaction module that introduces factorization-machine-style second-order feature interactions into the experts. Factorization machines, introduced by Steffen Rendle in 2010, model pairwise interactions between features through learned latent vectors, and they remain a powerful inductive bias for recommendation data, where the conjunction of two features, say a user attribute and a video genre, often carries more predictive signal than either feature alone. By embedding this bias at the expert level, H-PLE equips its subnetworks with an explicit mechanism for capturing cross-feature effects, which the authors find particularly valuable for the rare engagement tasks where every scrap of signal counts.</p>
<p>The evaluation is notable for its rigor. The researchers tested their framework on two real-world scenarios from the Tenrec benchmark, a large-scale multi-purpose recommendation dataset released by Tencent, using the QK-video and QB-video settings. They constructed binary prediction tasks for Like and Share, two engagement signals with highly imbalanced positive rates. Rather than relying on a single train-test split or a single random seed, the team ran multi-seed experiments with independent random splits, and they compared against a broad slate of modern baselines: MMoE, PLE, Cross-Stitch networks, PCGrad gradient surgery, and loss-rebalancing techniques including focal loss and class-balanced binary cross-entropy. Metrics went beyond the usual area under the ROC curve to include PR-AUC, which is far more informative under imbalance, GAUC for user-level ranking quality, calibration measures, and threshold-optimized F1 scores.</p>
<p>The results tell a story that is both encouraging and sobering. H-PLE-family variants consistently improved sparse-task ranking and minority-class retrieval over MMoE and PLE, the two architectures that dominate industrial practice. They also remained competitive with, or stronger than, the Cross-Stitch, PCGrad, and loss-rebalancing baselines. Perhaps the most striking finding concerns those rebalancing baselines themselves: focal loss and class-balanced BCE, techniques that are widely recommended for imbalanced classification, actually underperformed vanilla MMoE in the study&#8217;s most imbalanced setting, the QB-Share task. That result underscores just how difficult extreme sparsity is, and it serves as a caution against assuming that loss-level fixes can substitute for architectural ones when the imbalance is severe enough.</p>
<p>Equally important is the paper&#8217;s honesty about its own limits. The full H-PLE model was not uniformly best on every QB metric. The ablation studies, which systematically removed individual components, revealed that AGC, the raw-input-aware gating mechanism, is the most robust component under the most extreme sparsity, while the factorization-machine interaction module delivers benefits only when its interaction rank is carefully controlled. In other words, the gains come from specific, identifiable mechanisms rather than from a blanket advantage, and practitioners deploying the framework would need to tune the interaction rank rather than simply switch on every component. This kind of component-level attribution, made possible by the ablation design, is rarer in the recommendation literature than it should be, and it makes the paper&#8217;s positive claims considerably more trustworthy.</p>
<p>The practical implications reach well beyond video recommendation. Any production system that predicts multiple user behaviors, from e-commerce platforms jointly modeling clicks, add-to-carts, and purchases, to content platforms balancing watch time, comments, and subscriptions, confronts the same structural problem: dense signals dominate shared representations and starve sparse ones. The insight that gating networks are themselves vulnerable to imbalance-driven bias, and that anchoring them to raw inputs can counteract this, offers a general design principle that could be ported to other multi-task architectures. The finding that second-order interactions help only under careful rank control similarly generalizes as a reminder that inductive biases are tools with operating ranges, not free lunches.</p>
<p>The study also reflects a commendable degree of transparency about data and reproducibility. The authors state that the Tenrec dataset is available from Tencent&#8217;s public benchmark page under its own access conditions, which prevents redistribution by the authors, but that preprocessing scripts, experimental configuration files, and result summaries can be obtained from the corresponding author upon reasonable request, subject to the dataset&#8217;s terms of use. The work was supported by JST SPRING grant JPMJSP2123 and JSPS KAKENHI grant 25K17794, and the authors declare no competing interests. As recommendation systems increasingly mediate what billions of people see, ensuring that rare but meaningful signals, the shares and the passionate endorsements rather than the passive clicks, are not lost in the statistical noise becomes a matter of both engineering quality and platform health. H-PLE offers a carefully validated step in that direction, and its candid reporting of where the method falls short may prove as influential as the architecture itself.</p>
<p><strong>Subject of Research:</strong> Multi-task learning for recommendation systems under extreme label imbalance</p>
<p><strong>Article Title:</strong> H-PLE: hierarchical progressive layered extraction with raw-input-aware gating for multi-task recommendation under extreme label imbalance</p>
<p><strong>Article References:</strong> Xue, J., Yang, T., &amp; Suzuki, H. (2026). H-PLE: hierarchical progressive layered extraction with raw-input-aware gating for multi-task recommendation under extreme label imbalance. <em>Applied Intelligence, 56</em>(14), Article 409. <a href="https://doi.org/10.1007/s10489-026-07441-5" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07441-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07441-5" rel="noopener noreferrer">10.1007/s10489-026-07441-5</a></p>
<p><strong>Keywords:</strong> multi-task learning, recommendation systems, label imbalance, negative transfer, mixture-of-experts, progressive layered extraction, gating networks, factorization machines, Tenrec benchmark, PR-AUC, neural networks, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">244197</post-id>	</item>
		<item>
		<title>Graph Learning Beats Fraud by Structure, Not Smarter Score Fusion, Study Finds</title>
		<link>https://scienmag.com/graph-learning-beats-fraud-by-structure-not-smarter-score-fusion-study-finds/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 13:13:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advancements in fraud detection technology]]></category>
		<category><![CDATA[challenges in identifying synthetic identities]]></category>
		<category><![CDATA[deep learning vs rule-based fraud screening]]></category>
		<category><![CDATA[effectiveness of graph-based fraud detection]]></category>
		<category><![CDATA[financial crime]]></category>
		<category><![CDATA[fraud detection]]></category>
		<category><![CDATA[fraud prevention using graph learning]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[heterogeneous graph transformer]]></category>
		<category><![CDATA[identity linkage]]></category>
		<category><![CDATA[innovative approaches to combating financial fraud]]></category>
		<category><![CDATA[label propagation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in financial security]]></category>
		<category><![CDATA[noise robustness]]></category>
		<category><![CDATA[PR-AUC]]></category>
		<category><![CDATA[role of graph structure in financial crimes]]></category>
		<category><![CDATA[rule-based scoring]]></category>
		<category><![CDATA[score fusion]]></category>
		<category><![CDATA[structural fraud detection methods]]></category>
		<category><![CDATA[synthetic identity crime scale and impact]]></category>
		<category><![CDATA[synthetic identity fraud]]></category>
		<category><![CDATA[synthetic identity fraud detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227887</guid>

					<description><![CDATA[A controlled study of synthetic identity fraud detection shows that rule-anchored heterogeneous graph models win through complete identity relation structure rather than sophisticated score-level fusion.]]></description>
										<content:encoded><![CDATA[<p>Synthetic identity fraud has become one of the most elusive financial crimes of the digital era. Unlike conventional identity theft, in which a criminal impersonates a specific victim who eventually notices the damage and reports it, synthetic identity fraud involves assembling an entirely new persona from a blend of real, stolen, and fabricated information. Because no single person is harmed, there is no aggrieved victim to raise the alarm, and the fake identity can behave like an ordinary customer for months or years before losses surface. The scale of the problem is striking: outstanding balances linked to suspected synthetic identities in the United States reached a record 4.6 billion dollars in 2022, vehicle-rental providers reported 1.8 billion dollars in related losses in the first half of 2023 alone, and the McKinsey Global Institute estimates the crime accounts for roughly 15 percent of charge-offs in unsecured lending portfolios.</p>
<p>A new study published in Discover Informatics by Thi Khanh Hoai Nguyen and Chaochang Chiu of Yuan Ze University in Taiwan tackles a deceptively simple question: when graph neural networks outperform traditional fraud screening, where exactly does that advantage come from? Rather than assuming that deep learning automatically dominates rule-based methods, the researchers built a controlled experimental framework to measure the marginal contribution of each component in a fraud detection pipeline. Their central finding is counterintuitive and operationally important: the value of graph learning in synthetic identity fraud detection comes from the completeness of the identity relations the model can access, not from increasingly sophisticated ways of blending graph scores with rule scores.</p>
<p>The study exploits a structural weakness that fraudsters cannot easily avoid. Fabricating a fully unique identity for every fraudulent profile is expensive, so offenders routinely recycle a limited pool of identifiers, reusing the same Social Security Numbers, email addresses, and telephone numbers across many customer accounts. This reuse links otherwise unrelated profiles into a web of shared identifiers, forming the empirical basis for identity-linkage analysis. The researchers modeled this structure as a heterogeneous graph with four node types—Client, SSN, Email, and Phone—connected by typed relations. When several fraudulent clients converge on a single SSN node, a detectable fraud ring appears in the topology, which is precisely the signature that linkage-based screening exploits.</p>
<p>The data came from a PaySim-derived identity linkage dataset containing 2,433 customer profiles, of which 433 were fraudulent, a prevalence of about 17.8 percent. Across the dataset, roughly 12 to 13 percent of clients shared each identifier type with at least one other client, and fraudulent clients shared identifiers at roughly 2.6 times the rate of normal clients. At zero noise, normal clients had an average rule score of just 0.003, with only 0.25 percent sharing at least one identifier, while fraudulent clients averaged a rule score of 4.947, with 76.44 percent sharing at least one identifier. Even at the highest tested noise level, 64 percent of fraudulent clients still shared at least one identifier, explaining why a simple counting rule is such a formidable baseline.</p>
<p>The evaluation was organized in two stages under a single controlled protocol. In the first stage, four model families were compared independently: transparent rule-based scoring, non-graph machine learning such as logistic regression and gradient boosting, lightweight label propagation methods, and deep graph models including graph convolutional networks and the Heterogeneous Graph Transformer, or HGT. In the second stage, representative models were anchored to the rule core in a hybrid pipeline, so the incremental value of each component could be isolated. The rule itself is elegantly simple: each client is scored by how many other clients share its SSN, email, or phone, a metric that requires no training and is fully auditable by regulators and compliance teams.</p>
<p>The results were revealing. The rule alone achieved an average Precision-Recall Area Under the Curve, or PR-AUC, of 0.777. Non-graph machine learning models produced results almost identical to the rule, confirming that when their input features are derived from the same linkage counts, conventional learners extract nothing new. LabelSpreading, an iterative diffusion method with no deep encoder, was the strongest standalone model at 0.807, a finding the authors emphasize should elevate propagation methods to the status of serious competitors rather than weak baselines. Among deep models, R-GCN collapsed to near-random performance due to a representational failure in which embedding variance fell to approximately 8 times 10 to the minus 11, while HGT-based models performed strongly.</p>
<p>The standout hybrid was Rule + HGT+SVM, which combined the rule score with a Heterogeneous Graph Transformer encoder feeding a support vector machine classifier. It achieved the best average PR-AUC of 0.835, a gain of 0.049 over the rule alone, with a positive difference in all 20 seed-noise combinations tested. The hybrid also degraded more gracefully under corruption: from 0 to 40 percent fraud-only noise, it lost 0.092 PR-AUC compared with 0.116 for the rule alone, and at the highest noise level it retained 0.768 PR-AUC against 0.704 for the rule. In fraud screening, where reviewers can only examine a limited number of flagged profiles per day, such gains in ranking quality translate directly into additional true positives surfaced within a fixed review budget.</p>
<p>What makes the study distinctive is its forensic dissection of why the hybrid wins. Because rule and graph scores ranked clients almost identically, the researchers tested whether two more elaborate fusion mechanisms—an adaptive per-client weight and a tie-breaking scheme—could extract further gains. Neither could. The fixed-weight fusion preserved the rule&#8217;s ordering for 100 percent of pairs with different rule scores in every run, and 99.5 percent of clients shared their exact integer rule score with another client, making ties the typical case rather than the exception. Within tied groups, every fusion formula reduced to the same graph-only ordering, which was only marginally better than a random tie order at 51.1 percent accuracy. The fusion weight, in other words, was not the source of the improvement at all.</p>
<p>The real answer came from a relation-level ablation. Removing SSN, Email, or Phone individually from the graph before training consistently reduced HGT+SVM performance across all tested seed-noise configurations, with SSN removal producing the largest average degradation. Crucially, a density-matched control—randomly deleting the same number of edges while keeping all three relation types—performed worse than structural relation removal, with nominal p-values of 0.0007 or lower in the descriptive 20-pair comparison. This means a graph missing one relation entirely but retaining the others in full is a better input to the encoder than a graph that keeps all relations but thinned. What matters is having some relations represented completely rather than all relations represented but degraded, a finding that reframes how hybrid fraud systems should be designed.</p>
<p>The authors are candid about the limits of their work. The dataset is small and synthetic, the noise protocols are simpler than real adversarial behavior, and with only four independent seeds the conservative seed-level statistical tests cannot reach conventional significance thresholds, so the findings are best read as directionally consistent evidence rather than confirmatory proof. A further negative result stands out: in an inductive setting where new clients are attached to the graph only at inference, the two-layer HGT architecture collapsed to random ranking, meaning the hybrid&#8217;s advantage is currently validated only for retrospective, transductive batch screening of an existing customer book. Even so, the practical message is clear. Transparent rule-based identity-linkage scoring remains the necessary, auditable backbone of a defensible fraud pipeline, while heterogeneous graph models serve as a complementary refinement layer whose value flows from relation-aware representation—not from fancier score fusion.</p>
<p><strong>Subject of Research:</strong> Graph-based machine learning methods for detecting synthetic identity fraud in payment systems</p>
<p><strong>Article Title:</strong> Rule anchored graph learning improves synthetic identity fraud detection through multi relation structure rather than score level fusion</p>
<p><strong>Article References:</strong> Nguyen, T. K. H., &amp; Chiu, C. (2026). Rule anchored graph learning improves synthetic identity fraud detection through multi relation structure rather than score level fusion. <em>Discover Informatics, 1</em>(1), Article 22. <a href="https://doi.org/10.1007/s44564-026-00024-z" rel="noopener noreferrer">https://doi.org/10.1007/s44564-026-00024-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44564-026-00024-z" rel="noopener noreferrer">10.1007/s44564-026-00024-z</a></p>
<p><strong>Keywords:</strong> synthetic identity fraud, graph neural networks, heterogeneous graph transformer, identity linkage, fraud detection, rule-based scoring, score fusion, label propagation, PR-AUC, noise robustness, machine learning, financial crime</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227887</post-id>	</item>
		<item>
		<title>Smarter Features, Not Bigger Models, Crack Earthquake Forecasting in Central Asia</title>
		<link>https://scienmag.com/smarter-features-not-bigger-models-crack-earthquake-forecasting-in-central-asia/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:10:06 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI approaches to seismic activity]]></category>
		<category><![CDATA[Bi-LSTM]]></category>
		<category><![CDATA[CatBoost]]></category>
		<category><![CDATA[Central Asia]]></category>
		<category><![CDATA[Central Asia earthquake risk]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[class imbalance in seismic data]]></category>
		<category><![CDATA[earthquake forecasting]]></category>
		<category><![CDATA[earthquake forecasting frameworks]]></category>
		<category><![CDATA[Earthquake prediction]]></category>
		<category><![CDATA[earthquake prediction accuracy]]></category>
		<category><![CDATA[fault descriptors]]></category>
		<category><![CDATA[geophysical data analysis]]></category>
		<category><![CDATA[gradient boosting]]></category>
		<category><![CDATA[Kazakhstan]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in seismology]]></category>
		<category><![CDATA[neural network applications in earthquakes]]></category>
		<category><![CDATA[Omori decay]]></category>
		<category><![CDATA[PR-AUC]]></category>
		<category><![CDATA[predictive modeling for natural disasters]]></category>
		<category><![CDATA[seismic forecasting]]></category>
		<category><![CDATA[spatio-temporal prediction]]></category>
		<category><![CDATA[statistical evaluation of earthquake models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204452</guid>

					<description><![CDATA[A Kazakhstani research team shows that physically structured features and calibrated baselines, not model architecture, drive a tenfold improvement in macro-scale earthquake forecasting across Central Asia.]]></description>
										<content:encoded><![CDATA[<p>Earthquakes are among the most stubborn prediction problems in all of science, and a new study from Kazakhstan suggests that the path forward may lie less in exotic neural architectures and more in how the problem itself is framed. Writing in the Journal of Big Data, a team led by Marat Nurtas of the Ionosphere Institute and the International Information Technology University in Almaty reports a macro-scale earthquake forecasting framework for Central Asia that achieves a roughly tenfold improvement over a naive statistical baseline, reaching a Precision-Recall Area Under the Curve of approximately 0.451 and a Receiver Operating Characteristic Area Under the Curve of about 0.844. The striking twist is that six fundamentally different machine learning architectures, from gradient boosting to recurrent neural networks, all converged on nearly identical performance, pointing to a fundamental predictability ceiling rather than a modeling shortfall.</p>
<p>The research tackles a problem that has long plagued computational seismology: extreme class imbalance. When earthquake forecasting is cast as a grid-based classification task, the region under study is divided into spatial cells, and the model must predict whether a seismic event will occur in each cell during each time window. For Central Asia, the team discretized the territory into one-degree by one-degree grid cells and attempted to forecast earthquakes of magnitude 3.0 or greater on a weekly basis. Under this formulation, roughly 96 percent of all cell-week combinations contain no event at all, a phenomenon known as zero inflation. The true event prevalence sits at only about 4.5 percent, which means that a model doing nothing more than predicting &#8216;no earthquake&#8217; everywhere would still appear superficially accurate while being scientifically useless.</p>
<p>This imbalance has profound consequences for how forecasting models must be evaluated. Standard accuracy metrics become meaningless when negative cases dominate by more than twenty to one. The researchers therefore anchored their evaluation in the Precision-Recall Area Under the Curve, a metric that is far more sensitive to performance on the rare positive class. With a prevalence of 4.5 percent, the constant baseline for PR-AUC is 0.045, meaning any model must substantially exceed that value to demonstrate genuine predictive skill. The achieved score of 0.451 represents a tenfold improvement over this baseline, a substantial margin in a domain where even modest gains above chance are considered meaningful by the seismological community.</p>
<p>Central to the study is a carefully engineered feature space grounded in earthquake physics rather than raw statistical patterns. The framework integrates tectonic regime-conditioned normalization, which allows the model to account for the fact that different tectonic settings produce fundamentally different seismic behavior, so that features extracted from a thrust-fault environment are not treated as directly comparable to those from a strike-slip regime. It also incorporates Omori energy decay proxies, mathematical representations of the well-documented tendency of earthquake sequences to produce aftershocks whose frequency decays over time following a mainshock. These proxies give the models a physically interpretable signal about the temporal clustering of seismicity, encoding decades of seismological understanding directly into the input data.</p>
<p>Structural fault descriptors form a third pillar of the feature design. The geometry, orientation, and proximity of mapped fault systems are among the strongest known controls on where earthquakes occur, and by encoding these structural characteristics as model inputs, the framework ensures that the learning algorithms operate on geologically meaningful quantities rather than arbitrary grid statistics. The final and perhaps most consequential innovation is log-odds baseline initialization, a technique that encodes the historical cell-specific event rate directly into the learning objective. Instead of forcing each model to rediscover from scratch the simple fact that some grid cells are historically far more seismically active than others, the initialization embeds this prior knowledge into the model&#8217;s starting point, allowing learning effort to focus on deviations from the historical pattern.</p>
<p>To determine whether performance under such extreme imbalance is governed primarily by model architecture or by structured feature design, the researchers evaluated six heterogeneous architectures under a strict chronological split, ensuring that models were trained only on past data and tested on future periods, exactly as an operational forecasting system would be deployed. The architectures spanned a wide methodological range, including CatBoost and other gradient boosting methods, which excel at tabular data, and Bi-LSTM networks, a bidirectional long short-term memory architecture capable of capturing temporal dependencies in sequential data. Despite their radically different inductive biases and internal mechanics, the models converged on nearly identical PR-AUC and ROC-AUC values, a result the authors interpret as evidence that the information content of the feature space, not the capacity of the learner, is the binding constraint.</p>
<p>Equally notable is what the framework does not do. Many studies confronting severe class imbalance resort to synthetic resampling techniques, such as oversampling the rare event class or undersampling the dominant negative class, to artificially balance the training distribution. These methods can distort the learned probability calibration, producing models whose confidence scores no longer correspond to real-world event likelihoods. The Central Asia framework achieves its tenfold improvement entirely without synthetic resampling, preserving the integrity of the probability estimates. This matters enormously for practical applications, because emergency management authorities require calibrated forecasts whose stated probabilities can be trusted when weighing evacuation decisions, infrastructure inspections, and public warnings.</p>
<p>The authors argue that their findings point to the existence of a macro-scale predictability ceiling in seismic forecasting. If architecturally diverse models, given the same physically structured inputs, all plateau at the same performance level, the implication is that the remaining unpredictability reflects genuine stochasticity in the earthquake process at this spatial and temporal resolution, rather than a deficiency of current algorithms. This interpretation carries a sobering but valuable message for the field: further architectural innovation alone is unlikely to break through the ceiling, while improvements in physical understanding, richer observational data streams, and better-calibrated baselines may still push the boundary outward. It also cautions against the common practice of claiming architectural superiority from small performance differences that may fall within the noise of a shared predictability limit.</p>
<p>The work was funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan under a grant for developing a multifunctional system of ground-space monitoring and early warning of natural and technogenic emergencies, underscoring its operational motivation. For a country situated in one of the most seismically active zones of Central Asia, where the collision of the Indian and Eurasian plates drives hazardous tectonics through the Tien Shan and surrounding mountain belts, reliable macro-scale forecasting is not an academic curiosity but a matter of public safety. By demonstrating that disciplined feature engineering, physically informed priors, and rigorous baseline calibration can deliver a tenfold gain in predictive skill without exotic machinery, the Almaty team has provided both a practical forecasting tool and a methodological lesson that resonates far beyond seismology: in data-starved, imbalance-dominated problems, how you frame the question often matters more than how elaborate your model is.</p>
<p><strong>Subject of Research:</strong> Machine learning earthquake forecasting under extreme class imbalance in Central Asia</p>
<p><strong>Article Title:</strong> Macro-scale earthquake forecasting under class imbalance in Central Asia</p>
<p><strong>Article References:</strong> Nurtas, M., Nurakynov, S., Sakabekov, A., Altaibek, A., Kumarkhanova, A., &amp; Merekeyev, A. (2026). Macro-scale earthquake forecasting under class imbalance in Central Asia. <em>Journal of Big Data</em>. <a href="https://doi.org/10.1186/s40537-026-01544-z" rel="noopener noreferrer">https://doi.org/10.1186/s40537-026-01544-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40537-026-01544-z" rel="noopener noreferrer">10.1186/s40537-026-01544-z</a></p>
<p><strong>Keywords:</strong> earthquake forecasting, Central Asia, class imbalance, machine learning, CatBoost, gradient boosting, Bi-LSTM, PR-AUC, spatio-temporal prediction, Omori decay, fault descriptors, Kazakhstan</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204452</post-id>	</item>
	</channel>
</rss>
