<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>real-time data streaming analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/real-time-data-streaming-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 09:12:56 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>real-time data streaming analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Graph-Based Algorithm Tames Streaming Features in High-Dimensional Data</title>
		<link>https://scienmag.com/new-graph-based-algorithm-tames-streaming-features-in-high-dimensional-data/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 09:12:56 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[classification accuracy]]></category>
		<category><![CDATA[data mining]]></category>
		<category><![CDATA[dimensionality reduction]]></category>
		<category><![CDATA[feature redundancy reduction]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[feature selection in gene expression and spam filtering]]></category>
		<category><![CDATA[graph-based feature selection]]></category>
		<category><![CDATA[graph-based scoring]]></category>
		<category><![CDATA[group streaming feature selection]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[high-dimensional data algorithms]]></category>
		<category><![CDATA[Laplacian score]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[modular sensor data processing]]></category>
		<category><![CDATA[multi-task Lasso]]></category>
		<category><![CDATA[multi-view learning systems]]></category>
		<category><![CDATA[real-time data streaming analysis]]></category>
		<category><![CDATA[scalable machine learning methods]]></category>
		<category><![CDATA[sparse reconstruction]]></category>
		<category><![CDATA[sparse reconstruction in machine learning]]></category>
		<category><![CDATA[streaming features]]></category>
		<category><![CDATA[Streaming high-dimensional data analysis]]></category>
		<category><![CDATA[structural relationship preservation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234418</guid>

					<description><![CDATA[A new three-stage algorithm called FS-GSR combines graph-based scoring with sparse reconstruction to select compact, high-performing feature subsets from high-dimensional data whose features arrive in streaming groups.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models are often drowning in data, but the problem is not always the sheer volume of samples. Increasingly, it is the flood of features—the individual measurable properties of each data point—that threatens to overwhelm algorithms. In domains ranging from gene expression profiling to spam filtering, features do not arrive all at once. They stream in sequentially, often in predefined groups, and waiting for the complete feature set before analysis begins is computationally prohibitive. A new study published in Applied Intelligence proposes a method designed precisely for this challenging scenario, and its results suggest a significant step forward in how machines can learn from data that never stops arriving.</p>
<p>The method, called FS-GSR—Feature Selection with Group Streaming via Graph-based Scoring and Sparse Reconstruction—was developed by Tianyuan Jia of The Hong Kong Polytechnic University. It tackles a core weakness of existing streaming feature selection techniques: most process features one at a time, ignoring the structural relationships and redundancies that exist both within and between groups of features. When features arrive in batches, as they do in multi-view learning systems or modular sensor deployments, sequential individual processing can disrupt group structures and degrade performance. FS-GSR instead operates at the group level, updating the selected feature subset in real time as each new batch arrives.</p>
<p>At the heart of the approach is a three-stage framework that functions like a computational funnel. The first stage performs a hybrid, graph-guided filtering of each incoming feature group. The algorithm constructs a similarity graph over the samples, connecting points that are close in the feature space, and computes a Laplacian Score for every feature. This score measures how smoothly a feature varies across locally similar samples—features that preserve the intrinsic geometric structure of the data receive low scores, while noisy, irregular features are penalized. In parallel, a supervised F-score evaluates each feature&#8217;s discriminative power by comparing between-class variance to within-class variance. The two measures are blended with equal weight, and only the top fraction of features by combined score survives this initial screening.</p>
<p>The second stage refines this candidate pool by preserving structure at a finer granularity. The algorithm performs a spectral decomposition of the graph Laplacian, extracting ten eigenvectors that form a low-dimensional embedding of the sample manifold. Each surviving feature is then evaluated by how well it can reconstruct this structural representation, using a closed-form projection solution that avoids costly iterative optimization. Features that best approximate the manifold&#8217;s skeleton are retained, ensuring that the selected subset captures not just individual relevance but the collective geometry of the data. Because the reconstruction coefficients have an explicit analytical solution, this stage adds minimal computational overhead.</p>
<p>The third and final stage is where FS-GSR departs most sharply from its predecessors. All candidate features accumulated from the first two stages are merged into a global pool, and a multi-task Lasso regression is applied across the entire set simultaneously, using one-hot encoded class labels as the response. The L2,1-norm penalty drives the coefficients of redundant features to exactly zero, yielding a compact, non-overlapping final feature set. The authors also provide theoretical backing for this stage, formulating it as a row-support recovery problem in high-dimensional multivariate regression. Under standard conditions on the covariance structure, they prove that the sparse reconstruction can recover the true set of informative features with high probability, provided the effective sample complexity exceeds a specific threshold.</p>
<p>The computational design is notable for its scalability. By confining expensive cubic operations to the fixed sample space rather than the growing feature dimension, the algorithm&#8217;s cost per streaming group scales linearly with the number of incoming features. Because the first two stages aggressively compress each group before global optimization, the feature pool fed into the multi-task Lasso remains small, preventing the exponential blow-up that plagues conventional structure-preserving methods. Scalability experiments on the Ovarian Cancer dataset confirmed near-linear growth in running time as feature dimension increased, and revealed a V-shaped trade-off in execution time as the number of streaming groups varied—very few large groups incur heavy intra-group matrix operations, while very many small groups accumulate loop overhead.</p>
<p>Empirically, the method was tested on five benchmark datasets spanning biomedical, image, and text domains, including Colon (62 samples, 2,000 features), Ovarian Cancer (253 samples, 15,154 features), MLL, COIL20, and PCMAC. These datasets deliberately covered contrasting correlation structures: the Ovarian Cancer data exhibited strong intra-group correlations of up to 0.88, while the PCMAC text dataset showed extremely sparse correlations of roughly 0.025. Against four established baselines—OGFS, Group-SAOLA, Fast-OSFS, and alpha-investing—FS-GSR consistently achieved the highest or near-highest classification accuracy and F1 scores under both KNN and SVM classifiers. On the Ovarian Cancer dataset, it reached accuracy and F1 values above 0.99.</p>
<p>Perhaps the most striking result concerned feature compression. In one experimental configuration, FS-GSR retained an average of just 8.12 features from the 15,154-dimensional Ovarian Cancer dataset, while simultaneously improving KNN accuracy from 0.9845 to 0.9960 and SVM accuracy from 0.9885 to 0.9980—a 63 percent compression relative to the two-stage variant alone. Ablation studies confirmed that each stage contributes distinct benefits: the first enables efficient local filtering, the second improves structural consistency, and the third enforces global sparsity and discriminative refinement. The authors note that the degree of compression adapts to each dataset&#8217;s intrinsic correlation structure, with strongly correlated data requiring more retained features to preserve associative signals.</p>
<p>Parameter sensitivity analyses reinforced the method&#8217;s practicality. Classification accuracy remained stable across broad plateaus for the retention ratios governing the first two stages, and the mixing weight balancing structural and discriminative scoring performed robustly between 0.2 and 0.8, peaking near 0.5. The embedding dimension saturated quickly, with accuracy plateauing once it exceeded five, justifying the default setting of ten. This robustness means practitioners can deploy the method without exhaustive grid searches, using a simple boundary-driven tuning strategy on validation data.</p>
<p>The implications extend well beyond the benchmark datasets. Streaming, group-wise feature arrival is characteristic of stepwise sensor deployment, progressive module activation in multi-view systems, and sequential extraction pipelines in bioinformatics and text mining. By combining graph-based manifold learning with sparse global reconstruction in an online framework, FS-GSR offers a template for learning systems that must remain both accurate and parsimonious as data evolves. The author identifies future extensions toward multi-label settings, incomplete data systems, and broader generalization, suggesting that the challenge of features that never stop flowing is one the field is only beginning to master.</p>
<p><strong>Subject of Research:</strong> Group streaming feature selection for high-dimensional data using graph-based scoring and sparse reconstruction</p>
<p><strong>Article Title:</strong> FS-GSR: Graph-based scoring and sparse reconstruction for group streaming feature selection</p>
<p><strong>Article References:</strong> Jia, T. (2026). FS-GSR: Graph-based scoring and sparse reconstruction for group streaming feature selection. <em>Applied Intelligence, 56</em>(15), Article 445. <a href="https://doi.org/10.1007/s10489-026-07498-2" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07498-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07498-2" rel="noopener noreferrer">10.1007/s10489-026-07498-2</a></p>
<p><strong>Keywords:</strong> feature selection, streaming features, graph-based scoring, sparse reconstruction, multi-task Lasso, Laplacian score, machine learning, high-dimensional data, data mining, classification accuracy, dimensionality reduction, Applied Intelligence</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234418</post-id>	</item>
	</channel>
</rss>
