<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>high-dimensional data &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/high-dimensional-data/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 23:41:47 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>high-dimensional data &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Clustering Framework Tames Massive High-Dimensional Data Without the Usual Bottleneck</title>
		<link>https://scienmag.com/new-clustering-framework-tames-massive-high-dimensional-data-without-the-usual-bottleneck/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:41:47 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[affinity graph]]></category>
		<category><![CDATA[anchor-based modeling]]></category>
		<category><![CDATA[clustering algorithms]]></category>
		<category><![CDATA[computational efficiency in clustering]]></category>
		<category><![CDATA[data mining]]></category>
		<category><![CDATA[document collection clustering]]></category>
		<category><![CDATA[gene expression data analysis]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[high-dimensional data clustering]]></category>
		<category><![CDATA[high-dimensional data visualization]]></category>
		<category><![CDATA[image data clustering]]></category>
		<category><![CDATA[large-scale data analysis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[manifold regularization]]></category>
		<category><![CDATA[neural network-based clustering]]></category>
		<category><![CDATA[nonnegative matrix factorization]]></category>
		<category><![CDATA[representation learning]]></category>
		<category><![CDATA[scalability]]></category>
		<category><![CDATA[scalable data mining algorithms]]></category>
		<category><![CDATA[sensor stream data segmentation]]></category>
		<category><![CDATA[subspace clustering]]></category>
		<category><![CDATA[subspace clustering techniques]]></category>
		<category><![CDATA[tensor-based clustering methods]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213439</guid>

					<description><![CDATA[Researchers have unveiled a subspace clustering framework that preserves the structural fidelity of high-dimensional data while avoiding the prohibitive cost of full sample-to-sample affinity matrices.]]></description>
										<content:encoded><![CDATA[<p>Modern data rarely arrives in tidy, well-separated groups. Images, gene expression profiles, sensor streams, and document collections all tend to live on tangled, low-dimensional structures hidden inside spaces with thousands of dimensions. Subspace clustering has emerged as one of the most powerful mathematical tools for finding those hidden structures, modeling a dataset as a union of multiple linear or affine subspaces and assigning each point to the subspace where it truly belongs. Yet the technique has long suffered from a painful trade-off: the methods that capture structure most faithfully tend to collapse under the sheer scale of real-world data. A new study published in Data Mining and Knowledge Discovery by Mo Chen, Xuesong Yin, Qi Huang, Jianhao Ding, Guodao Zhang, and Xinjun Miao now offers a way to have both fidelity and scale, presenting a framework called Large-scale Structured Subspace Clustering, or LSSC.</p>
<p>The central obstacle the researchers set out to overcome is a computational one that has shaped the field for over a decade. Classical subspace clustering methods rely on a self-representation model, in which every data point is expressed as a linear combination of all the other points. The resulting coefficient matrix encodes the relationships that define cluster membership, and it is extraordinarily informative. The problem is its size. For n data points, the full sample-to-sample affinity matrix contains n squared entries, so a dataset of one million samples would demand a matrix with a trillion coefficients. Constructing, storing, and optimizing over such an object is prohibitive on any hardware, which is precisely why many theoretically elegant subspace clustering algorithms have remained confined to benchmark datasets of a few thousand points.</p>
<p>The LSSC framework breaks this bottleneck by replacing the full self-representation with a compact sample-to-anchor representation. Instead of allowing every point to be reconstructed from every other point, the method first selects a much smaller set of anchor points that serve as representative landmarks for the data. Each sample is then expressed as a nonnegative combination of these anchors alone. Because the number of anchors grows far more slowly than the number of samples, the coefficient matrix shrinks from n squared to a size proportional to n times the anchor count, turning an intractable optimization into a manageable one. Crucially, the nonnegativity constraint on the coefficients gives the representation a natural probabilistic flavor: each weight reflects how strongly a sample belongs to the neighborhood of a given anchor, which makes the downstream clustering step far more stable.</p>
<p>Compactness alone, however, is not enough. The authors identify a second, subtler failure mode in existing scalable methods: fast representations often fail to preserve the intrinsic structural relationships of the data when distributions become complex. Two points that sit close together on the same underlying manifold may end up with very different anchor coefficients simply because of how the anchors were sampled, and the resulting affinity graph can fragment genuine clusters. LSSC counters this with a locality-aware coefficient initialization, which seeds the optimization using neighborhood information so that the reconstruction process starts from a configuration already consistent with the local geometry of the data. This initialization acts like a well-chosen starting point in a rugged optimization landscape, steering the solution toward coefficients that respect the manifold rather than artifacts of sampling.</p>
<p>On top of that initialization, the framework applies distance-weighted structure regularization, a mechanism that explicitly encourages the learned representation to remain smooth with respect to the geometry of the data. Samples that are near each other in the original feature space are pushed toward similar anchor coefficients, weaving locality information directly into the optimization objective rather than bolting it on afterward. The framework also employs coefficient regularization to keep the representation compact and well-conditioned, suppressing degenerate solutions in which a sample spreads its weight indiscriminately across many anchors. Together, these four components—anchor-based reconstruction, locality-aware initialization, distance-weighted regularization, and coefficient regularization—form a single coherent objective that learns a structured, nonnegative representation without ever forming the full sample-to-sample affinity matrix.</p>
<p>The mathematical machinery behind the approach draws on a rich lineage of research. Low-rank representation, introduced by Liu and colleagues, and sparse subspace clustering, developed by Elhamifar and Vidal, established the foundational idea that imposing structure on self-representation coefficients reveals subspace membership. Nonnegative matrix factorization, famously connected by Lee and Seung to the way humans learn parts of objects, supplies the theoretical justification for the nonnegative coefficients at the heart of LSSC. More recent work on anchor graphs and landmark-based spectral clustering demonstrated that working with a reduced set of representative points can preserve much of the quality of full spectral methods at a fraction of the cost. LSSC synthesizes these threads, but its distinctive contribution is the integration of structure-preserving regularization into the anchor-based pipeline, addressing the structural fidelity problem that earlier fast methods left unresolved.</p>
<p>To test the framework, the team ran extensive experiments on nine benchmark datasets, comparing LSSC against a battery of established clustering algorithms using four standard evaluation metrics: clustering accuracy (ACC), normalized mutual information (NMI), Purity, and the adjusted rand index (ARI). The results were striking. Among the compared methods that completed all nine datasets, LSSC achieved the best overall average rank across all four metrics, indicating that its advantage is not confined to one type of data or one evaluation criterion but reflects a consistently strong performance profile. On the large-scale datasets, where many competitors either slowed to a crawl or failed outright, LSSC maintained competitive clustering quality while demonstrating the favorable scalability that the anchor-based design was built to deliver.</p>
<p>The authors are careful to characterize the limits of their method as well as its strengths, a candor that lends the study particular scientific value. Their experiments reveal that LSSC shows sensitivity to highly imbalanced data distributions, a scenario in which some clusters contain vastly more samples than others. This vulnerability is a known hazard of anchor-based approaches, since anchors tend to be drawn disproportionately from dense regions of the data, leaving sparse clusters underrepresented in the reconstruction basis. By documenting this weakness explicitly, the researchers provide a clear empirical map of where the framework should be deployed and where practitioners should exercise caution, an honest assessment that is all too rare in a field often driven by headline benchmark numbers.</p>
<p>The implications reach well beyond the machine learning community. Subspace clustering underpins applications from image segmentation and hyperspectral image analysis to single-cell RNA sequencing, where biologists attempt to sort thousands of individual cells into functional types based on high-dimensional gene expression. A method that scales gracefully while preserving structural fidelity could make such analyses feasible on datasets that are currently out of reach, and the same logic applies to speech processing, anomaly detection, and any domain where high-dimensional observations hide low-dimensional group structure. The work, supported by research grants from Zhejiang Province and Hangzhou Dianzi University, arrives at a moment when the volume of unlabeled data is exploding faster than the tools to organize it. By showing that structure and scale need not be enemies, LSSC offers a template for the next generation of clustering algorithms: representations that are compact by design, yet rich enough to remember the geometry that made the clusters real in the first place.</p>
<p><strong>Subject of Research:</strong> Scalable anchor-based subspace clustering for large high-dimensional datasets</p>
<p><strong>Article Title:</strong> Large-scale structured subspace clustering</p>
<p><strong>Article References:</strong> Chen, M., Yin, X., Huang, Q., Ding, J., Zhang, G., &amp; Miao, X. (2026). Large-scale structured subspace clustering. <em>Data Mining and Knowledge Discovery, 40</em>(6), Article 99. <a href="https://doi.org/10.1007/s10618-026-01268-x" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01268-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01268-x" rel="noopener noreferrer">10.1007/s10618-026-01268-x</a></p>
<p><strong>Keywords:</strong> subspace clustering, anchor-based modeling, representation learning, affinity graph, manifold regularization, nonnegative matrix factorization, data mining, unsupervised learning, scalability, high-dimensional data, clustering algorithms, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213439</post-id>	</item>
		<item>
		<title>New Deterministic Algorithm Tames the Combinatorial Chaos of Regression Subset Selection</title>
		<link>https://scienmag.com/new-deterministic-algorithm-tames-the-combinatorial-chaos-of-regression-subset-selection/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:08:43 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[best subset regression]]></category>
		<category><![CDATA[bootstrap validation]]></category>
		<category><![CDATA[combinatorial optimization in linear regression]]></category>
		<category><![CDATA[deterministic algorithms]]></category>
		<category><![CDATA[deterministic methods for variable subset selection]]></category>
		<category><![CDATA[deterministic regression subset selection algorithm]]></category>
		<category><![CDATA[efficient algorithms for predictor subset identification]]></category>
		<category><![CDATA[EPR-C3 heuristic for model selection]]></category>
		<category><![CDATA[exhaustive all-subsets regression computational challenges]]></category>
		<category><![CDATA[heuristic optimization]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[high-dimensional data analysis in statistics]]></category>
		<category><![CDATA[high-dimensional predictor variable selection]]></category>
		<category><![CDATA[model selection]]></category>
		<category><![CDATA[modern approaches to regression variable selection]]></category>
		<category><![CDATA[multicollinearity]]></category>
		<category><![CDATA[multiple linear regression]]></category>
		<category><![CDATA[ordinary least squares]]></category>
		<category><![CDATA[overcoming combinatorial explosion in regression analysis]]></category>
		<category><![CDATA[scalable regression model selection techniques]]></category>
		<category><![CDATA[statistical computing]]></category>
		<category><![CDATA[statistical rigor in subset selection]]></category>
		<category><![CDATA[subset selection]]></category>
		<category><![CDATA[variance inflation factor]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212959</guid>

					<description><![CDATA[A new deterministic heuristic called EPR-C3 recovers most of the best predictor subsets found by exhaustive regression while dramatically cutting computation time in high-dimensional problems.]]></description>
										<content:encoded><![CDATA[<p>Few problems in statistics look as innocent on paper as choosing which variables to include in a multiple linear regression. Given a table of candidate predictors and an outcome of interest, the task seems to be nothing more than deciding which columns belong in the model. Yet the moment the number of candidate predictors grows, the problem explodes combinatorially. With fifty predictors there are more than a quadrillion possible subsets; with a hundred, the number of models exceeds the number of atoms in the universe by an unimaginable margin. Exhaustive all-subsets regression, the gold standard that guarantees the best model, becomes computationally impossible long before researchers reach the dimensions that modern datasets routinely present. A new study published in the International Journal of Data Science and Analytics by Jackson J. Alcázar of Universidad del Desarrollo in Chile confronts this wall directly, introducing a deterministic heuristic called EPR-C3 that promises to make high-dimensional subset selection tractable without sacrificing statistical rigor.</p>
<p>The core insight behind EPR-C3 is that the search for the best subset should not be blind. Traditional approaches to the problem fall into two broad camps, each with well-known weaknesses. Stepwise procedures, which add or remove predictors one at a time based on statistical criteria, are fast but notoriously unstable, often missing the globally best model and producing results that shift with small perturbations of the data. Penalized regression methods such as the lasso, ridge, and elastic net shrink coefficients to perform implicit selection, but they transform the problem itself, replacing the ordinary least squares formulation with a modified objective and yielding biased estimates that complicate interpretation. Genetic algorithms and other stochastic search techniques can explore large spaces, but their reliance on random initialization and random operators means that two runs of the same algorithm on the same data can produce different answers, undermining the reproducibility that scientific work demands.</p>
<p>EPR-C3 takes a different path. The method is a deterministic multi-start neighborhood-search heuristic, meaning that it systematically explores multiple regions of the subset space using structured moves rather than random jumps, and that repeated runs on identical data produce identical results. The name encodes its workflow: expansion, perturbation, and reduction form the exploratory core, while C3 refers to a three-part constraint refinement stage comprising correlation cleanup, replacement recovery, and variance inflation factor pruning. Each stage addresses a specific pathology of high-dimensional regression. Expansion grows a candidate subset by adding predictors that promise improvement. Perturbation shakes up the current solution to escape local optima, the traps where a greedy search would otherwise stall. Reduction trims away predictors that no longer earn their place.</p>
<p>The C3 refinement stage is where the method embeds statistical admissibility directly into the search. Correlation cleanup removes predictors that are so strongly correlated with others already in the model that they add redundancy rather than information. Replacement recovery attempts to swap problematic predictors for alternatives that preserve predictive power while restoring stability. Variance inflation factor pruning, the final safeguard, quantifies how much the variance of an estimated coefficient is inflated by multicollinearity and eliminates predictors whose presence destabilizes the model. By enforcing these constraints during the search rather than checking them afterward, EPR-C3 ensures that every candidate model it evaluates is statistically admissible, avoiding the embarrassing situation in which a search returns a model with excellent fit but coefficients so unstable that they are meaningless.</p>
<p>A crucial design decision sets EPR-C3 apart from penalized approaches: it preserves the native ordinary least squares formulation throughout. The final output is not a shrunken or regularized set of coefficients but an explicit regression equation of the familiar form, with interpretable coefficients estimated by classical least squares. This matters enormously for applied researchers in fields such as chemistry, medicine, and the social sciences, where the goal is often not just prediction but understanding, and where an explicit equation with transparent coefficients is the deliverable that matters. The author&#8217;s own background in chemical modeling, including prior work on predicting acid dissociation constants in nitrogen compounds, reflects this applied orientation toward interpretable equations.</p>
<p>The benchmark results are striking. In analyses comparing EPR-C3 against exhaustive all-subsets regression, the heuristic recovered a high proportion of the Top-K admissible subsets, the best-ranked models that exhaustive search identifies, while dramatically reducing computational cost in highly combinatorial problems. In other words, when the exhaustive answer is computable, EPR-C3 usually finds it; when exhaustive search is impossible, EPR-C3 still delivers high-quality admissible models in a fraction of the time. The study goes further and proposes an empirical utility threshold, a practical guideline that tells analysts when the heuristic search becomes preferable to exhaustive enumeration, giving practitioners a principled rule for choosing between the two strategies rather than guessing.</p>
<p>The comparative evaluation is unusually thorough for a methods paper. EPR-C3 was benchmarked against stepwise procedures, penalized-regression preselection, filter and wrapper screening approaches, genetic algorithms, and branch-and-bound subset selection, the latter being the classical exact method that prunes the search space using bounds on the objective. Across these comparisons, a consistent pattern emerged: embedding admissibility constraints within the search workflow improves the tradeoff between recovery of the best models and computational efficiency. Competing methods that search first and check statistical validity later, or that abandon the least squares framework entirely, tend to either waste effort on inadmissible candidates or return models that require post-hoc repair. EPR-C3&#8217;s constraint-aware architecture avoids both failure modes.</p>
<p>Reproducibility and robustness received dedicated attention. Bootstrap analyses, in which the method was run on resampled versions of the data, showed close agreement with exhaustive selection in best-subset recovery, out-of-bag performance, and predictor-inclusion concentration, the degree to which the same predictors are selected consistently across resamples. This concentration measure is particularly important because instability of variable selection is one of the most cited criticisms of heuristic and stepwise methods. A seed-package analysis further confirmed that the deterministic nature of the algorithm holds in practice, with repeated runs converging on the same solutions. An applied case study demonstrated that EPR-C3 could reproduce a previously published multiple linear regression equation, recovering the same model with reduced runtime, providing a concrete demonstration that the method delivers on real problems rather than only on synthetic benchmarks.</p>
<p>The significance of this work extends beyond the specific algorithm. Variable selection remains one of the most contested practices in applied statistics, with methodologists repeatedly warning that automated selection can produce misleading inference, inflated significance, and models that fail to replicate. By making the search deterministic, constraint-aware, and anchored in the classical least squares framework, EPR-C3 offers a middle path between the computational impossibility of exhaustive search and the statistical compromises of penalization and greedy stepwise selection. The method is implemented in the author-maintained MLR-X project and publicly available, lowering the barrier for adoption. As datasets in genomics, chemometrics, economics, and the behavioral sciences continue to swell in dimensionality, tools that combine computational tractability with statistical admissibility and exact reproducibility are likely to find eager audiences. The study, published as volume 22, article 311 of the journal and supported by ANID-FONDECYT grant funding, suggests that the future of high-dimensional model selection may lie not in abandoning classical regression but in searching its vast model space more intelligently.</p>
<p><strong>Subject of Research:</strong> Deterministic constraint-aware subset selection in high-dimensional multiple linear regression</p>
<p><strong>Article Title:</strong> Epr-c3: a deterministic constraint-aware heuristic for high-dimensional subset selection in multiple linear regression</p>
<p><strong>Article References:</strong> Epr-c3: a deterministic constraint-aware heuristic for high-dimensional subset selection in multiple linear regression. (n.d.). <a href="https://doi.org/10.1007/s41060-026-01298-0" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01298-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01298-0" rel="noopener noreferrer">10.1007/s41060-026-01298-0</a></p>
<p><strong>Keywords:</strong> subset selection, multiple linear regression, ordinary least squares, multicollinearity, variance inflation factor, heuristic optimization, best-subset regression, statistical computing, high-dimensional data, model selection, deterministic algorithms, bootstrap validation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212959</post-id>	</item>
		<item>
		<title>New Ensemble Feature Selection Method Reaches Near-Perfect IoT Intrusion Detection</title>
		<link>https://scienmag.com/new-ensemble-feature-selection-method-reaches-near-perfect-iot-intrusion-detection/</link>
		
		<dc:creator><![CDATA[Hailey Crawford]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 17:54:42 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[anomaly detection in IoT]]></category>
		<category><![CDATA[boosting-based feature ranking]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[cybersecurity]]></category>
		<category><![CDATA[cybersecurity for connected devices]]></category>
		<category><![CDATA[dimensionality reduction]]></category>
		<category><![CDATA[ensemble feature selection]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[feature ranking]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[high-dimensional traffic data analysis]]></category>
		<category><![CDATA[imbalanced dataset handling]]></category>
		<category><![CDATA[intrusion detection]]></category>
		<category><![CDATA[IoT intrusion detection]]></category>
		<category><![CDATA[IoT network security]]></category>
		<category><![CDATA[IoT security]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning cybersecurity]]></category>
		<category><![CDATA[mutual information]]></category>
		<category><![CDATA[optimized feature subset selection]]></category>
		<category><![CDATA[ranking with boosting]]></category>
		<category><![CDATA[real-time IoT threat identification]]></category>
		<category><![CDATA[scalable intrusion detection methods]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=207415</guid>

					<description><![CDATA[Researchers at Manipur University have developed EF²RB, an ensemble feature selection framework that boosts IoT intrusion detection accuracy to between 95 and 100 percent on benchmark datasets.]]></description>
										<content:encoded><![CDATA[<p>Researchers at Manipur University in India have unveiled a new machine learning framework that dramatically improves the way security systems detect intrusions in Internet of Things networks, achieving detection accuracy between 95 and 100 percent on several benchmark datasets. The method, called EF²RB, tackles one of the most stubborn problems in modern cybersecurity: the sheer size and messiness of the traffic data that billions of connected devices generate every second. By combining multiple feature selection techniques with an innovative boosting-style ranking mechanism, the team showed that a carefully chosen handful of features can outperform far larger feature sets, cutting computational cost while tightening detection performance.</p>
<p>The Internet of Things presents security analysts with an unusually hostile data environment. Unlike traditional enterprise networks, IoT ecosystems mix smart thermostats, cameras, medical sensors, and industrial controllers, each producing traffic with distinct statistical fingerprints. The resulting datasets are high-dimensional, often containing hundreds of numerical and categorical features, and severely imbalanced, with benign traffic dwarfing attack traffic by orders of magnitude. Attacks themselves range from botnet command-and-control traffic to application-layer exploits, so the signals that betray an intrusion vary enormously in character. Machine learning classifiers trained on such data struggle not because the information is absent but because it is buried under redundant, irrelevant, and noisy variables that dilute the learning signal.</p>
<p>Feature selection is the discipline of separating signal from this noise. Rather than transforming features into new abstract dimensions, as feature extraction methods such as autoencoders do, feature selection identifies and retains the original attributes that carry the most discriminative power, preserving interpretability. Filter methods, which rank features using statistical measures such as mutual information and correlation without consulting any classifier, are prized for their speed and scalability. Their weakness is that any single ranking criterion has blind spots; a measure that captures linear relationships may miss nonlinear ones, and vice versa. Ensemble feature selection addresses this by aggregating the verdicts of several ranking algorithms, in the same spirit that ensemble classifiers combine many weak learners into one strong predictor.</p>
<p>EF²RB, developed by Chandam Chinglensana Singh, Nazrul Hoque, and Khumukcham Robindro Singh, extends this ensemble philosophy with a concept the authors call Ranking with Boosting. The framework begins by partitioning the feature space into manageable segments, a decision that proved essential for scalability. Each partition is then evaluated by multiple filter ranking algorithms, including methods built on mutual information such as mRMR, JoMIC, and MIFS-ND, alongside correlation-based ranking. A consensus voting mechanism then reconciles the individual rankings to produce a common feature subset, and a correlation-pruning stage removes redundant variables that carry overlapping information. An iterative refinement loop, functioning as the ranker booster, progressively sharpens the subset across rounds, analogous to how boosting algorithms iteratively focus on the hardest examples.</p>
<p>The ablation experiments conducted by the team offer a revealing anatomy of the framework. Using the RT-IoT2022 dataset alongside three non-IoT benchmarks from the UCI Repository, namely Mice Protein Expression, Spambase, and Optical Recognition of Handwritten Digits, the researchers disabled each component in turn while holding the rest constant. The consensus mechanism emerged as the single most influential element: switching it off on RT-IoT2022 shrank the selected feature set from 21 features to just 5 and dragged accuracy down from roughly 98.94 percent to 94.77 percent. Correlation pruning, by contrast, mainly trims redundancy, expanding the retained features from 21 to 23 when disabled while leaving accuracy almost unchanged.</p>
<p>The most striking ablation result concerned feature-space partitioning. When the researchers allowed the ranking algorithms to process the entire feature space at once, runtimes exploded to extraordinary levels, with mRMR requiring more than 20,500 seconds, JoMIC nearly 50,000 seconds, and MIFS-ND more than 9,600 seconds on a single dataset. Worse, the consensus mechanism failed to produce any common feature subset at all, which prevented the classification stage from running. On the handwritten digits dataset, disabling partitioning likewise halted the pipeline entirely. These findings underline that for genuinely high-dimensional data, how you divide the problem can matter as much as the algorithms you apply to it. Iterative feature selection, interestingly, contributed only marginal gains, suggesting the framework remains robust even without that refinement layer.</p>
<p>Beyond the ablation study, the team embedded EF²RB in a full IoT intrusion detection system and tested it with baseline machine learning classifiers on high-dimensional IoT intrusion datasets that include Bot-IoT, Edge-IIoTset, NSL-KDD, UNSW-NB15, N-BaIoT, and Mu-IoT, alongside classical benchmarks. The resulting detector consistently achieved high performance, with accuracy between 95 and 100 percent on certain datasets, even though the classifier worked from a dramatically reduced feature subset. Because filter methods dominate the framework, the heavy computation happens once, offline, before deployment, leaving the live detector fast enough for resource-constrained IoT gateways that cannot afford heavyweight deep learning inference.</p>
<p>The comparison with existing approaches is instructive. Prior ensemble methods such as IDS-EFS and rank aggregation schemes in software defect prediction have shown that combining filters improves stability, while metaheuristic wrappers guided by multiple rankers have pushed accuracy on imbalanced data. EF²RB distinguishes itself by uniting partitioning, multi-ranker consensus, redundancy pruning, and iterative boosting in one pipeline, and by demonstrating each component&#8217;s contribution through systematic ablation rather than reporting only aggregate results. The authors have also released their implementation publicly on GitHub, which lowers the barrier for security teams and researchers to reproduce, audit, and adapt the framework for their own network environments.</p>
<p>The broader significance lies in what this means for defending the rapidly expanding IoT attack surface. Botnet campaigns such as Mirai demonstrated years ago that insecure connected devices can be weaponized at internet scale, and detection systems have struggled to keep pace with both the volume and heterogeneity of device traffic. A feature selection method that is dataset-agnostic, as the UCI benchmark results indicate, offers security engineers a reusable tool rather than a bespoke fix: the same pipeline that compresses IoT intrusion data can also condense spam indicators or biomedical measurements. As regulatory pressure and liability concerns push manufacturers toward hardened devices, methods like EF²RB supply the monitoring side of that equation, turning enormous, noisy traffic streams into compact feature sets that lightweight classifiers can act on in near real time.</p>
<p>Limitations remain, and the authors are candid about them. The evaluation relied on established benchmark datasets rather than newly generated traffic, so real-world deployment performance will depend on how faithfully those datasets mirror live networks, which constantly evolve as attackers adapt. The framework&#8217;s runtime, though far better with partitioning, still depends on computationally expensive mutual information estimates, which may need further optimization for streaming or federated deployments. Yet the core result stands: disciplined consensus among diverse ranking perspectives, guided by a boosting-inspired refinement, can extract a small, potent feature subset from the chaos of high-dimensional IoT data. In a field where every percentage point of detection accuracy translates into intercepted attacks, that is a consequential advance.</p>
<p><strong>Subject of Research:</strong> Ensemble filter feature selection with ranker boosting for high-dimensional IoT intrusion detection</p>
<p><strong>Article Title:</strong> EF&#040;^{2}&#041;RB: Ensemble of filter feature selection methods with ranker booster for classification of high-dimensional IoT intrusion data</p>
<p><strong>Article References:</strong> Chinglensana Singh, C., Hoque, N., &amp; Robindro Singh, K. (2026). EF$$^{2}$$RB: Ensemble of filter feature selection methods with ranker booster for classification of high-dimensional IoT intrusion data. <em>Knowledge and Information Systems, 68</em>(1), Article 252. <a href="https://doi.org/10.1007/s10115-026-02871-6" rel="noopener noreferrer">https://doi.org/10.1007/s10115-026-02871-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10115-026-02871-6" rel="noopener noreferrer">10.1007/s10115-026-02871-6</a></p>
<p><strong>Keywords:</strong> feature selection, feature ranking, dimensionality reduction, ensemble learning, IoT security, intrusion detection, machine learning, mutual information, high-dimensional data, class imbalance, cybersecurity, ranking with boosting</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">207415</post-id>	</item>
		<item>
		<title>Simple Fix Guarantees Connected Graphs for Robust Text Clustering</title>
		<link>https://scienmag.com/simple-fix-guarantees-connected-graphs-for-robust-text-clustering/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:47:22 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[document clustering]]></category>
		<category><![CDATA[Document similarity measurement]]></category>
		<category><![CDATA[Ensuring connected graphs for reliable topic discovery]]></category>
		<category><![CDATA[Epsilon-threshold graph construction]]></category>
		<category><![CDATA[graph connectivity]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[Hyperparameter sensitivity in graph construction]]></category>
		<category><![CDATA[incremental graph construction]]></category>
		<category><![CDATA[k-nearest neighbor graphs]]></category>
		<category><![CDATA[K-nearest-neighbor graph fragility]]></category>
		<category><![CDATA[Laplacian eigenmaps]]></category>
		<category><![CDATA[Machine Learning journal]]></category>
		<category><![CDATA[minimum spanning tree]]></category>
		<category><![CDATA[MTEB]]></category>
		<category><![CDATA[Near-duplicate detection in text datasets]]></category>
		<category><![CDATA[Neighborhood graph in text mining]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[Robust text clustering techniques]]></category>
		<category><![CDATA[Semi-supervised label propagation]]></category>
		<category><![CDATA[SentenceTransformer embeddings]]></category>
		<category><![CDATA[SentenceTransformers]]></category>
		<category><![CDATA[spectral clustering]]></category>
		<category><![CDATA[Spectral clustering challenges]]></category>
		<category><![CDATA[text embeddings]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204096</guid>

					<description><![CDATA[Researchers have developed an incremental k-nearest-neighbor graph construction that guarantees connectivity by design, dramatically improving spectral clustering of text embeddings at low sparsity while cutting build time and memory costs.]]></description>
										<content:encoded><![CDATA[<p>One of the quiet workhorses of modern text mining is a deceptively simple data structure: a neighborhood graph, in which every document becomes a node and edges link each document to the items most similar to it. These graphs underpin topic discovery, near-duplicate detection, semi-supervised label propagation, and the retrieval indices used in retrieval-augmented generation. Yet, according to a new study published in the journal Machine Learning, this foundational step is far more fragile than most practitioners realize. On realistic text datasets, the standard k-nearest-neighbor graphs that power spectral clustering can shatter into many disconnected components at the sparsity levels people actually use in practice, silently degrading results and making the whole pipeline hypersensitive to a single hyperparameter.</p>
<p>The study, led by Marko Pranjić and Boshko Koloski of the Jožef Stefan Institute together with Nada Lavrač, Senja Pollak, and Marko Robnik-Šikonja of the University of Ljubljana, quantifies just how badly things can go wrong. Working with the well-known 20 Newsgroups collection of roughly 20,000 documents, the researchers encoded the test split of 7,532 documents into 384-dimensional vectors using the SentenceTransformer model all-MiniLM-L12-v2 and measured cosine distance between the embeddings. When they built epsilon-threshold graphs, connecting every pair of documents closer than a given distance, they found that even the minimal distance required to keep the graph connected demanded around one million edges, and increasing that distance by only five percent inflated the edge count by more than sixty percent. Sparser is better for memory and speed, but sparser also means disconnected.</p>
<p>The k-nearest-neighbor approach fared somewhat better, since every document is linked to its k most similar documents, and tuning k to obtain a connected graph proved easier than tuning a distance threshold. But connectedness is never guaranteed. Theoretical work cited in the study shows that a k-NN graph is only asymptotically connected with high probability when k exceeds roughly 5.1774 times the natural logarithm of the number of points, which for as few as 300 data points already means k above 30, far beyond the small values used in real pipelines. In the experiments, some dataset partitions within the TwentyNewsgroups benchmark exhibited disconnected components even at k equals 15, and a Reddit sentence-clustering dataset remained disconnected at k equals 20. The high-dimensional geometry of embeddings makes matters worse: a phenomenon known as hubness means a few points appear in a disproportionate number of neighbor lists while others become isolated anti-hubs, skewing degree distributions and aggravating disconnection.</p>
<p>Why does disconnection matter so much? In spectral clustering, each connected component of the graph can only be assigned to a single cluster. When the number of components equals or exceeds the number of desired clusters, the clustering becomes trivial, and no similarity-based criterion can rescue it. The same fragility propagates beyond clustering: in item-based recommender systems, unreachable items simply cannot be recommended, and in label propagation, an isolated component of documents can never receive a propagated label. Connectivity, in other words, determines reachability, and reachability determines whether the downstream algorithm can function at all.</p>
<p>The researchers&#8217; solution is elegantly minimal. Instead of letting every node search for neighbors across the entire dataset simultaneously, their incremental construction inserts documents one at a time, and each new node is linked to its k nearest neighbors among the nodes already present in the graph. That single restriction, searching only among previously inserted nodes, guarantees a connected graph for any value of k. The proof is a short induction: the first insertion creates a single connected component, and every subsequent node attaches to k existing members of that component, extending it without ever splitting it. Because the first node of a new region can only attach to earlier nodes, the graph remains one piece by construction, for every k and every insertion order.</p>
<p>The modification also confers a practical bonus that is increasingly relevant in the era of streaming data. In a standard k-NN graph, adding a single new document can trigger a large reconfiguration, since many existing nodes might count the newcomer among their nearest neighbors, effectively forcing a rebuild. The incremental construction induces only local changes, restricted to the row and column of the newly added node in the adjacency matrix. New documents can therefore be inserted without reconstructing the existing graph, and the last added nodes can even be deleted, properties the authors argue could be exploited in applications where data arrives continuously or becomes invalidated over time.</p>
<p>To test whether the guaranteed connectivity translates into better clustering, the team evaluated spectral clustering with Laplacian eigenmaps across eleven sentence- and paragraph-level clustering tasks drawn from six dataset sources in the Massive Text Embedding Benchmark, covering 182 distinct clustering problems in total. Performance was measured with V-measure, the harmonic combination of homogeneity and completeness. The results show a clear pattern: in the low-k regime, where standard k-NN graphs are most prone to fragmentation, the incremental approach consistently outperformed the standard construction, achieving near-top scores already at k equals 3 on many datasets, while matching the standard graph at larger k where both methods operate on connected structures. The largest gains appeared on TwentyNewsgroups, precisely where disconnection was most prevalent.</p>
<p>The team then compared their method against the most obvious alternative: repairing the standard k-NN graph by augmenting it with a minimum spanning tree, a global structure that stitches disconnected pieces together. Under matched MST augmentation, the incremental graph still held the edge in the sparse regime, improving on k-NN plus MST by an average of 2.5 V-measure points for k between 1 and 3, and by 3.8 points at k equals 1. A Bayesian signed-rank analysis with a region of practical equivalence confirmed that the MST is practically beneficial only in the sparsest regime and practically irrelevant from k equals 6 onward. Crucially, the exact MST computation requires forming the dense N-by-N distance matrix, an O(N-squared) step. The incremental construction avoids that matrix entirely because it attains connectivity through the insertion order rather than through added global edges, and it proved between 23 and 44 times faster to build than the MST-augmented alternatives, with peak memory dropping from as much as 20.2 gigabytes to as little as 1.2 gigabytes on the largest evaluated partitions, run on a CPU node of the Vega EuroHPC system.</p>
<p>Two natural objections receive careful treatment in the paper. First, because the incremental graph depends on the order in which nodes are inserted, is the method stable? Across ten randomized orderings per dataset, the standard deviation of clustering performance rarely exceeded one percent and was often below half a percent, and the structural statistics of the resulting graphs, including transitivity, assortativity, and label homophily, varied only in the third decimal place. Even adversarial orderings, in which the first inserted nodes were deliberately concentrated near the embedding centroid or drawn entirely from a single class, changed V-measure by at most 0.4 points relative to random ordering. Second, does the approach depend on the specific embedding model? Tests with larger models such as bge-base-en-v1.5 and all-mpnet-base-v2 showed that bigger encoders consistently improve results, though the benefit of denser graphs and larger models proved dataset-dependent, with Reddit clustering the notable exception.</p>
<p>The authors are candid about limitations. The five largest Reddit paragraph-level partitions exceeded the 32-bit index range of the sparse eigensolver and were excluded from some recomputed comparisons, and the robustness study covers concentrated initial seeds but not fully temporally ordered document streams, since the benchmark snapshots carry no timestamps. Future directions include replacing exact nearest-neighbor search with approximate methods, coupling the incremental graph with efficient eigenvector update techniques, and deploying the construction in temporal community detection. Still, the message is striking: a one-line change to how documents are inserted into a similarity graph eliminates a failure mode that has quietly lurked in spectral text clustering, delivering guaranteed connectivity, faster builds, lower memory, and better low-k performance, all without needing a proper metric, global graph features, or any repair step at all.</p>
<p><strong>Subject of Research:</strong> Incremental k-nearest-neighbor graph construction for robust spectral clustering of text embeddings</p>
<p><strong>Article Title:</strong> Incremental Graph Construction Enables Robust Spectral Clustering of Texts</p>
<p><strong>Article References:</strong> Incremental Graph Construction Enables Robust Spectral Clustering of Texts. (n.d.). <a href="https://doi.org/10.1007/s10994-026-07158-z" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07158-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07158-z" rel="noopener noreferrer">10.1007/s10994-026-07158-z</a></p>
<p><strong>Keywords:</strong> spectral clustering, incremental graph construction, k-nearest neighbor graphs, text embeddings, Laplacian eigenmaps, graph connectivity, minimum spanning tree, Machine Learning journal, SentenceTransformers, MTEB, document clustering, high-dimensional data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204096</post-id>	</item>
		<item>
		<title>Information Theory Meets Machine Learning to Catch Industrial Cyberattacks</title>
		<link>https://scienmag.com/information-theory-meets-machine-learning-to-catch-industrial-cyberattacks/</link>
		
		<dc:creator><![CDATA[Teresa Odom]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 22:23:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[anomaly detection in industrial environments]]></category>
		<category><![CDATA[anomaly detection in power grids]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[applying information theory to anomaly detection]]></category>
		<category><![CDATA[chemical plant cybersecurity]]></category>
		<category><![CDATA[cyberattack prevention in manufacturing]]></category>
		<category><![CDATA[data-driven cybersecurity methods]]></category>
		<category><![CDATA[early detection of industrial cyber threats]]></category>
		<category><![CDATA[Gaussian kernel]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[Industrial control system cybersecurity]]></category>
		<category><![CDATA[industrial control systems]]></category>
		<category><![CDATA[industrial cybersecurity]]></category>
		<category><![CDATA[information content]]></category>
		<category><![CDATA[information entropy]]></category>
		<category><![CDATA[information theory applications in machine learning]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for cyberattack detection]]></category>
		<category><![CDATA[machine learning techniques for industrial safety]]></category>
		<category><![CDATA[one-class support vector machine]]></category>
		<category><![CDATA[One-Class Support Vector Machine limitations]]></category>
		<category><![CDATA[SWaT dataset]]></category>
		<category><![CDATA[WADI dataset]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=199204</guid>

					<description><![CDATA[Researchers in Xi'an have developed an information-theory-based anomaly detection method that outperforms state-of-the-art algorithms on critical industrial control system benchmarks.]]></description>
										<content:encoded><![CDATA[<p>Industrial control systems quietly run the modern world. They purify drinking water, route electricity through power grids, manage chemical plants, and keep assembly lines moving. When something goes wrong in these systems—whether through mechanical failure or a deliberate cyberattack—the consequences can cascade from a single factory floor to entire cities. Detecting anomalies in these environments before they escalate is therefore one of the most consequential challenges in modern cybersecurity. A new study published in Applied Intelligence by researchers at Xi&#8217;an University of Posts and Telecommunications introduces a method that could significantly sharpen that detection capability, and it does so by borrowing one of the oldest and most elegant ideas in science: information theory.</p>
<p>The research, led by Zhongmin Wang, Zhongjian Yuan, Cong Gao, and Yanping Chen, addresses a long-standing weakness in a classical machine learning technique known as the One-Class Support Vector Machine, or OCSVM. The OCSVM has been a workhorse of anomaly detection since its introduction in the early 2000s. Its appeal lies in its ability to learn from a single class of data—normally, it needs to see only examples of healthy behavior to build a model of what &#8216;normal&#8217; looks like. Anything that falls outside that learned boundary is flagged as an anomaly. This is crucial in industrial settings, where attack examples are rare, dangerous to stage, and endlessly varied, making conventional supervised learning impractical.</p>
<p>Yet the OCSVM carries two stubborn Achilles&#8217; heels. First, its performance depends heavily on the choice of kernel function and, in particular, on the parameters of the Gaussian kernel that defines how similarity between data points is measured. Getting these parameters right often requires laborious trial and error or expensive cross-validation, and a poor choice can gut the model&#8217;s accuracy. Second, industrial sensor data is increasingly high-dimensional, with hundreds or thousands of measurements streaming from pumps, valves, and controllers. As dimensionality rises, data points spread farther apart, distributions become sparse, and the notion of distance itself loses meaning—a phenomenon statisticians call the curse of dimensionality. The Gaussian kernel, which relies on Euclidean distances, struggles to capture the true distribution of normal samples in this sparse high-dimensional space, and detection performance degrades.</p>
<p>The Xi&#8217;an team&#8217;s answer, which they call Information Clustering One-Class Support Vector Machine (IC-OCSVM), tackles both problems in a single framework built on information-theoretic foundations. The method begins with a preprocessing stage that treats each feature dimension of the industrial data as if it were a separate information system. For every dimension, the researchers compute the Shannon information entropy—a measure of the uncertainty or unpredictability carried by that variable. Dimensions with low information content contribute little to distinguishing between samples and are treated as redundant and removed. This entropy-based pruning reduces dimensionality in a principled way, concentrating the model&#8217;s attention on the sensors and signals that actually carry meaningful variability, and simultaneously softening the effects of sparsity before the learning stage even begins.</p>
<p>The second and more novel stage replaces the conventional Gaussian kernel entirely. Instead of measuring similarity through Euclidean distance, IC-OCSVM introduces the concept of information content to characterize the differences between samples. Information content, a concept descending from Claude Shannon&#8217;s mathematical theory of communication, quantifies how surprising or informative one sample is relative to another. The researchers design an explicit mapping function based on this information-content measure, which serves the same mathematical role as the implicit feature mapping performed by a kernel trick—but with a crucial advantage. Because the mapping is explicit and constructed directly from information theory, the method no longer depends on the delicate selection of Gaussian kernel parameters. The model essentially builds its own geometry from the information structure of the data rather than relying on a pre-chosen distance metric.</p>
<p>This design choice has a second benefit that matters greatly for industrial deployments. By measuring the distance between samples through information content rather than raw coordinate differences, the constructed OCSVM is far less susceptible to the sparsity problem that plagues high-dimensional data. Samples that would appear arbitrarily far apart in a high-dimensional Euclidean space may share substantial information structure, allowing the model to recognize the common signature of normal behavior even when the raw feature space is vast and thinly populated. In effect, the method asks a more meaningful question of the data: not &#8216;how far apart are these points?&#8217; but &#8216;how much do these observations tell us about each other?&#8217;</p>
<p>To test the approach, the team turned to two of the most demanding and widely respected benchmark datasets in industrial control system security: SWaT and WADI. SWaT, the Secure Water Treatment testbed developed at the Singapore University of Technology and Design, simulates a full-scale water purification process complete with realistic cyberattack scenarios, while WADI extends the same experimental philosophy to water distribution networks. Both datasets feature multivariate time series from dozens of sensors and actuators, injected attacks of varying sophistication, and the noisy, correlated measurements that make real industrial anomaly detection so difficult. The researchers compared IC-OCSVM against a roster of state-of-the-art anomaly detection algorithms, including deep-learning approaches built on autoencoders, generative adversarial networks, and graph neural networks.</p>
<p>The results were striking. IC-OCSVM achieved an F1-score of 85.91 percent on SWaT and 68.28 percent on WADI, outperforming the best competing baseline by 3.2 and 3.3 percentage points respectively. In a field where incremental gains of a fraction of a percentage point frequently justify publication, improvements of this size—particularly on the notoriously difficult WADI dataset—are significant. The F1-score, which balances precision and recall into a single number, is especially meaningful in industrial security, where a detector that cries wolf too often wastes operator attention and one that stays silent too long allows attacks to proceed. IC-OCSVM&#8217;s edge on both fronts suggests that the information-theoretic framing genuinely captures structure that Gaussian-kernel and deep-learning baselines miss.</p>
<p>The implications extend well beyond water treatment. Any setting where anomalies must be learned from normal data alone—power grid monitoring, manufacturing quality control, aircraft engine health tracking, building automation—faces the same twin burdens of kernel tuning and high-dimensional sparsity. A method that sidesteps kernel parameter selection removes a costly and error-prone step from the deployment pipeline, while entropy-based dimension reduction offers a computationally light alternative to heavyweight deep architectures. Notably, IC-OCSVM achieves its results without the massive training datasets and GPU resources that deep learning methods typically demand, which could make sophisticated anomaly detection accessible to smaller operators and resource-constrained facilities that cannot maintain large labeled datasets or dedicated machine learning infrastructure.</p>
<p>The work also represents a broader and somewhat counterintuitive trend in machine learning research: the return of classical theory to solve problems that modern deep learning has struggled with. Shannon&#8217;s information theory, formulated in the 1940s, provides tools that are interpretable, mathematically grounded, and robust in ways that black-box neural networks often are not. By fusing information-theoretic feature analysis with the boundary-learning power of support vector machines, the Xi&#8217;an researchers have demonstrated that careful mathematical design can still beat brute-force complexity in the right domain. As industrial systems become ever more connected and the attack surface for critical infrastructure continues to expand, tools like IC-OCSVM point toward a future in which the sentinels guarding our water, power, and factories are built not just on more data, but on a deeper understanding of what information itself reveals.</p>
<p><strong>Subject of Research:</strong> A hybrid information clustering and one-class support vector machine method for anomaly detection in industrial control systems</p>
<p><strong>Article Title:</strong> A hybrid method integrating information clustering and one-class support vector machine for industrial anomaly detection</p>
<p><strong>Article References:</strong> Wang, Z., Yuan, Z., Gao, C., &amp; Chen, Y. (2026). A hybrid method integrating information clustering and one-class support vector machine for industrial anomaly detection. <em>Applied Intelligence, 56</em>(14), Article 418. <a href="https://doi.org/10.1007/s10489-026-07447-z" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07447-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07447-z" rel="noopener noreferrer">10.1007/s10489-026-07447-z</a></p>
<p><strong>Keywords:</strong> anomaly detection, industrial control systems, one-class support vector machine, information entropy, information content, SWaT dataset, WADI dataset, industrial cybersecurity, machine learning, high-dimensional data, Gaussian kernel, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">199204</post-id>	</item>
		<item>
		<title>Gene Selection Gets Smarter: Co-expression Networks Meet Genetic Algorithms</title>
		<link>https://scienmag.com/gene-selection-gets-smarter-co-expression-networks-meet-genetic-algorithms/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:18:41 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[bioinformatics feature selection methods]]></category>
		<category><![CDATA[biomarker discovery]]></category>
		<category><![CDATA[cancer classification]]></category>
		<category><![CDATA[co-expression networks]]></category>
		<category><![CDATA[computational biology data challenges]]></category>
		<category><![CDATA[dimensionality reduction in genomics]]></category>
		<category><![CDATA[disease classification gene markers]]></category>
		<category><![CDATA[gene co-expression network analysis]]></category>
		<category><![CDATA[gene feature selection]]></category>
		<category><![CDATA[genetic algorithms]]></category>
		<category><![CDATA[genetic algorithms for feature selection]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[high-throughput sequencing data analysis]]></category>
		<category><![CDATA[information-theoretic genetic operators]]></category>
		<category><![CDATA[integrating biology and evolutionary mathematics]]></category>
		<category><![CDATA[machine learning in biomedical data]]></category>
		<category><![CDATA[multi-objective optimization]]></category>
		<category><![CDATA[mutual information]]></category>
		<category><![CDATA[noise reduction in genetic datasets]]></category>
		<category><![CDATA[NSGA-II]]></category>
		<category><![CDATA[Precision medicine]]></category>
		<category><![CDATA[Weighted Non-dominated Sorting Genetic Algorithm]]></category>
		<category><![CDATA[WGCNA]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196239</guid>

					<description><![CDATA[A new hybrid algorithm called CJWGA combines gene co-expression networks with enhanced genetic operators to select small, accurate gene subsets from high-dimensional medical data.]]></description>
										<content:encoded><![CDATA[<p>Modern medicine is drowning in data, and a new study argues that the way out is not more computing power but a smarter partnership between biology and evolutionary mathematics. In research published in the Journal of Advanced Research, a team led by Zhilin Wang, Weiping Ding, Jinquan Zhang, Ali Asghar Heidari, Mingjing Wang and Huiling Chen introduces a feature selection framework called CJWGA, a Weighted Non-dominated Sorting Genetic Algorithm that combines gene co-expression networks with information-theoretic genetic operators. The method is designed to tackle one of the most stubborn problems in computational biology: how to find the handful of genes that truly matter for disease classification inside datasets containing thousands of candidate features, most of which are noise, redundancy, or statistical distraction.</p>
<p>The scale of the problem is easy to underestimate. High-throughput sequencing and mass spectrometry now allow laboratories to measure the expression of every gene in the human genome across hundreds of samples at once. A dataset might record 10,000 genes while including fewer than a hundred patients. This imbalance creates what statisticians call the curse of dimensionality: the number of possible feature subsets grows as two to the power of n, so for a dataset with 10,000 genes the search space is astronomically larger than anything a brute-force enumeration could ever cover. Worse, adding features does not reliably improve a model. Extra genes can introduce redundancy and noise, causing classifiers to overfit the training data while performing poorly on patients they have never seen. Running times grow as well, because computational complexity rises steadily with the number of features examined.</p>
<p>Existing feature selection strategies fall into three broad families, each with well-known trade-offs. Filtering methods, which rank genes using statistical measures such as mutual information, are fast and scalable but blind to the interactions between features. Wrapper methods, which evaluate subsets by feeding them to a classifier, capture those nonlinear relationships but at a punishing computational cost. Embedded methods such as LASSO regression and tree-based models select features during training, but none of these approaches ask the deeper biological question: which genes actually work together, and which modules of co-regulated genes drive the disease being studied? The new framework was built precisely to fill that gap, treating the biology of gene cooperation as the starting point rather than an afterthought.</p>
<p>The first stage of CJWGA relies on Weighted Gene Co-expression Network Analysis, or WGCNA, a technique originally proposed by Zhang and Horvath that constructs a weighted network linking genes whose expression levels rise and fall together across samples. Genes are not loners; they participate in biological processes through intricate webs of interaction, and WGCNA captures those relationships from a systems perspective. The pipeline begins with Z-score normalization of expression values, followed by a Pearson correlation matrix that is then transformed into a weighted adjacency matrix using a soft thresholding exponent chosen so the network follows a scale-free topology, in which a few highly connected hub genes dominate while most genes have few connections. A Topological Overlap Measure, which accounts for shared neighbors, is then fed into hierarchical clustering to identify modules of functionally related genes, with module eigengenes derived by principal component analysis.</p>
<p>But the authors recognized that relying on a single eigengene per module throws away too much information. A lone principal component cannot reflect the diversity of functions within a module, and it can be biased by outlier expression patterns. Their answer is a preprocessing step called IMGCNet, which uses conditional mutual information to rank genes within each module by how much extra information they carry about the disease label, given the eigengene is already known. A higher conditional mutual information value means a gene retains a strong dependency on the phenotype even after controlling for what the module representative already explains. Larger modules are allowed to retain more genes and smaller modules fewer, through a descending allocation rule that preserves the biological representativeness of each module without letting small, specialized groups flood the analysis.</p>
<p>The second stage hands the modules to an enhanced version of NSGA-II, the classic multi-objective genetic algorithm that balances competing goals by evolving a population of candidate solutions toward a Pareto front. Here the two objectives are minimizing the number of selected genes and maximizing classification accuracy, formalized with a binary decision vector over features and evaluated with a K-Nearest Neighbor classifier on a 70-30 train-test split. Crucially, the researchers designed a hierarchical encoding scheme: the first layer of each chromosome encodes which modules are selected, and the second layer encodes which genes within each chosen module survive. This two-layer structure preserves the biological meaning of the modularization rather than flattening it back into a flat string of bits.</p>
<p>The heart of the contribution lies in two new operators. The Combined Information Entropy Crossover Operator, or CIECO, computes a joint mutual information score across the genes selected in both parents, those selected in neither, and those selected in only one. The resulting value, transformed through a probabilistic function, decides whether crossover should prune doubly-selected genes, promote single-selected ones, or hold steady. When the score is positive, unselected genes carry little information and conservative trimming is favored; when it is negative, redundant double selections are removed and a few unselected genes are introduced to seek greater information content. The Joint Adaptive Mutation Operator, or JAMO, then fine-tunes individual genes using an adaptive rate that depends on iteration progress, the proportion of genes already selected in the module, and the ratio of joint mutual information between selected and unselected genes, with an exponent parameter that keeps the balance under control.</p>
<p>The experimental evaluation covered eight publicly available gene expression datasets, including Brain_Tumor1, Brain_Tumor2, CNS, Leukemia, Leukemia1, Leukemia2, Lung_Cancer and Prostate_Tumor, all with more than 5,000 features and sample sizes between 50 and 203. Against three specialist algorithms, FQEISS, WMOSS and WQEISS, CJWGA achieved the lowest classification error on the CNS, Leukemia, Leukemia1, Leukemia2 and Prostate_Tumor datasets, while selecting the smallest feature subsets on six of the eight datasets. On Inverted Generational Distance, a standard measure of how well a computed Pareto front approximates the true optimum, CJWGA scored zero, meaning perfect overlap with the reference front, on six datasets. Ablation experiments confirmed that both new operators contribute measurably: removing the crossover operator or the mutation operator individually degraded either accuracy or subset compactness. Parameter sweeps established that a crossover proportion of 0.2 and a mutation exponent of 3 offered the most robust results. Because joint mutual information is computed only within compact modules rather than across the entire feature space, the framework retains reasonable scalability even as dataset dimensionality grows.</p>
<p>The implications reach beyond benchmark tables. A feature selection method that respects gene co-expression relationships can point clinicians toward biologically meaningful biomarkers rather than statistical artifacts, a prerequisite for precision medicine where a compact, interpretable gene panel must support diagnosis and treatment decisions. The authors caution, however, that systematic biological interpretation of the selected genes remains future work, and they note that the framework could be extended to dimensionality reduction problems well outside genomics. As high-throughput biology continues to generate data faster than medicine can absorb it, tools like CJWGA suggest that the path forward lies in algorithms that speak both languages fluently: the language of information theory and the language of biological networks. The study is available as open access, supported by the National Natural Science Foundation of China and several provincial research programs.</p>
<p><strong>Subject of Research:</strong> Gene feature selection in high-dimensional medical gene expression data using co-expression networks and genetic algorithms</p>
<p><strong>Article Title:</strong> Advancing Gene Feature Selection: A Synergistic Approach with Co-expression Networks and Genetic Algorithms</p>
<p><strong>Article References:</strong> Wangy, Z., Ding, W., Zhang, J., Heidari, A. A., Wang, M., &amp; Chen, H. (2026). Advancing Gene Feature Selection: A Synergistic Approach with Co-expression Networks and Genetic Algorithms. <em>Journal of Advanced Research</em>. <a href="https://doi.org/10.1016/j.jare.2026.08.064" rel="noopener noreferrer">https://doi.org/10.1016/j.jare.2026.08.064</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.jare.2026.08.064" rel="noopener noreferrer">10.1016/j.jare.2026.08.064</a></p>
<p><strong>Keywords:</strong> gene feature selection, co-expression networks, WGCNA, genetic algorithms, multi-objective optimization, NSGA-II, mutual information, bioinformatics, cancer classification, precision medicine, high-dimensional data, biomarker discovery</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196239</post-id>	</item>
	</channel>
</rss>
