<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>spectral clustering &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/spectral-clustering/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 00:54:13 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>spectral clustering &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Safeguarded Graph Learning Method Promises Clustering That Can Only Get Better, Never Worse</title>
		<link>https://scienmag.com/new-safeguarded-graph-learning-method-promises-clustering-that-can-only-get-better-never-worse/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 00:54:13 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced clustering algorithms for medical imaging and customer data]]></category>
		<category><![CDATA[Applied Intelligence]]></category>
		<category><![CDATA[benchmark datasets]]></category>
		<category><![CDATA[clustering safety guarantee]]></category>
		<category><![CDATA[clustering safety net]]></category>
		<category><![CDATA[data fusion]]></category>
		<category><![CDATA[graph fusion]]></category>
		<category><![CDATA[graph learning]]></category>
		<category><![CDATA[guaranteed performance in clustering algorithms]]></category>
		<category><![CDATA[machine learning theory]]></category>
		<category><![CDATA[mathematically safe clustering methods]]></category>
		<category><![CDATA[maximin optimization]]></category>
		<category><![CDATA[multi-modal data clustering]]></category>
		<category><![CDATA[multi-view clustering]]></category>
		<category><![CDATA[multi-view data analysis]]></category>
		<category><![CDATA[multi-view data integration]]></category>
		<category><![CDATA[performance guarantees in unsupervised learning]]></category>
		<category><![CDATA[reliability in data science clustering]]></category>
		<category><![CDATA[robust graph-based clustering techniques]]></category>
		<category><![CDATA[safe machine learning]]></category>
		<category><![CDATA[safe maximin graph learning]]></category>
		<category><![CDATA[spectral clustering]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=229879</guid>

					<description><![CDATA[Researchers in Guangzhou have developed a graph-based multi-view clustering method built on maximin optimization that comes with a theorem guaranteeing its results never fall below a baseline clustering performance.]]></description>
										<content:encoded><![CDATA[<p>Clustering algorithms are the workhorses of modern data science, quietly sorting everything from medical images to customer records into meaningful groups without any labels to guide them. Yet for all their power, these algorithms share a frustrating weakness: there is rarely any guarantee that a sophisticated new method will actually outperform a simpler baseline. A team of researchers in Guangzhou, China, has now tackled this problem head-on, developing a multi-view clustering technique that comes with a mathematical safety net. The work, published in Applied Intelligence by Naiyao Liang, Zuyuan Yang, Dan Xiang, Mingyang Liu, Jianbin Xiong, and Junjie Yang, introduces a framework called safe maximin multiple graph learning, which is designed so that its clustering results can never fall below the performance of a reference baseline.</p>
<p>The problem the team addresses stems from the way modern data arrives. A single object in the real world is rarely described by just one kind of information. A news article, for instance, may be represented by its text, the images it contains, and the links pointing to it. A patient may be characterized by blood tests, scan results, and clinical notes. Each of these descriptions is called a view, and multi-view clustering aims to combine them so that the resulting grouping of data points is more accurate than what any single view could achieve alone. When the data is abundant and the views are informative, this fusion works beautifully. But when some views are noisy, redundant, or misleading, the fusion process can actively degrade the outcome, producing clusters that are worse than those obtained from a single, well-behaved view.</p>
<p>Graph-based methods have become one of the most popular families of techniques for this fusion task. In these approaches, each view of the data is converted into a graph, a mathematical structure in which nodes represent data points and weighted edges encode how similar two points are according to that particular view. The algorithm then learns a consensus graph, or a set of fused graphs, that ideally captures the shared structure across all views while discarding view-specific noise. Spectral clustering or related procedures are then applied to the learned graph structure to produce the final partition. Over the past decade, researchers have proposed a rich variety of such methods, including parameter-free auto-weighted multiple graph learning, self-weighted multiview clustering, and graph-based multi-view clustering frameworks that construct a single high-quality consensus graph directly from the raw data.</p>
<p>What has been largely missing from this literature, the authors argue, is the notion of safety. In the emerging field of safe machine learning, an algorithm is considered safe if its performance is guaranteed not to be worse than that of a specified baseline method. This idea has been explored in safe classification, safe weakly supervised learning, and, more recently, in safe multi-view clustering, where researchers have developed theorems and algorithms ensuring that adding more views to a model cannot cause performance degradation. Tang and Liu, for example, presented deep safe multi-view clustering at the Conference on Computer Vision and Pattern Recognition in 2022, explicitly aiming to reduce the risk of performance loss as the number of views increases. Tao and colleagues earlier proposed reliable multi-view clustering with similar motivations. However, in the specific and widely used setting of graph-based multi-view clustering, safety guarantees have remained rare, limiting how confidently these methods can be deployed in sensitive applications.</p>
<p>The new method fills this gap by introducing a maximizing performance gain strategy directly into the graph learning process. Rather than fusing multiple view-specific graphs in a way that simply optimizes an average objective, the proposed model formulates the fusion as a maximin optimization problem. The term maximin refers to a classical decision-theoretic principle: choose the action that maximizes the minimum possible gain. In the clustering context, this means the algorithm seeks a fused graph representation that maximizes the worst-case improvement over the baseline clustering result across the candidate solutions considered during learning. By explicitly optimizing the lower bound of performance gain rather than an expected or average gain, the model builds conservatism into the fusion itself, steering the learned graphs toward solutions that are robust even under unfavorable conditions in individual views.</p>
<p>The theoretical heart of the paper is a theorem that the authors develop to certify the safety of the proposed model. The theorem establishes that the clustering result obtained by the safe maximin multiple graph learning framework is guaranteed to be at least as good as the baseline clustering result, in the sense of the performance measure adopted by the framework. This is a meaningful departure from most graph-based multi-view clustering methods, where any claim of superiority rests entirely on empirical comparison. Here, the safety property is proven rather than merely observed, which means practitioners can adopt the method knowing that, under the conditions of the theorem, the fused solution will not underperform the reference. The authors complement the theorem with a remark that analyzes the safety of the proposed algorithm itself, ensuring that the numerical procedure used to solve the optimization problem preserves the safety guarantee established at the model level.</p>
<p>Solving the maximin formulation is not trivial, because such problems involve nested optimization: an inner problem that evaluates the worst-case scenario and an outer problem that maximizes over it. The researchers developed an effective algorithm tailored to the structure of their objective, iteratively updating the learned graphs and the fusion weights so that the maximin criterion is satisfied. The computational machinery draws on established optimization tools, and the authors acknowledge the MOSEK optimization software in their materials, suggesting that conic optimization solvers play a role in the implementation. The algorithmic design matters because a safety guarantee that only holds for the exact solution of the model is of limited practical value; the authors&#8217; remark on the algorithm&#8217;s safety addresses precisely this concern, connecting the discrete iterations of the solver to the theoretical property of the continuous model.</p>
<p>To evaluate the framework, the team conducted experiments comparing their method against state-of-the-art multi-view clustering approaches. The experimental results, as reported in the paper, indicate that the proposed method achieves safe clustering results, meaning the empirical outcomes are consistent with the theoretical guarantee, while also obtaining competitive clustering performance relative to leading alternatives. The datasets used in the study include well-known multi-view benchmarks referenced in the paper&#8217;s data notes, such as the 3Sources news dataset, collections distributed with the COMIC framework, multi-view datasets from the ELKI project repository, and data associated with reliable multi-view clustering and consistent graph learning research. These benchmarks span text, image, and heterogeneous sources, providing a reasonable testbed for assessing whether the safety mechanism comes at an unacceptable cost in raw accuracy.</p>
<p>The significance of this work extends beyond a single algorithm. Safe learning has become an increasingly urgent concern as machine learning systems are deployed in domains where failure carries real consequences, from medical imaging to autonomous perception. The reference list of the paper reflects this broader context, citing work on reconciling privacy and accuracy in AI for medical imaging, trusted multi-view classification with evidential fusion, and safe multi-view graph convolutional networks for semi-supervised classification. By bringing formal safety guarantees into graph-based multi-view clustering, one of the most widely used paradigms in unsupervised learning, the study signals a shift in how the community thinks about model evaluation. Accuracy alone, the authors suggest, is an incomplete criterion; a method should also be judged by whether it can be trusted not to make things worse.</p>
<p>There are, of course, practical considerations that will shape adoption. The maximin formulation adds computational overhead compared with simpler fusion schemes, and the safety guarantee is defined relative to a chosen baseline and performance measure, so the strength of the guarantee depends on how meaningful that baseline is for a given application. The authors note that data supporting the findings are available from the corresponding author upon reasonable request, which should facilitate independent verification. The work was supported by the National Natural Science Foundation of China and several Guangdong provincial research programs, reflecting sustained institutional investment in safe and reliable machine learning. As multi-view data continues to proliferate across science, industry, and medicine, frameworks like safe maximin multiple graph learning point toward a future in which unsupervised algorithms are not just powerful but provably dependable, offering users a floor of performance beneath which their results cannot fall.</p>
<p><strong>Subject of Research:</strong> Safe graph-based multi-view clustering using maximin optimization with guaranteed performance over a baseline</p>
<p><strong>Article Title:</strong> Safe maximin multiple graph learning for multi-view clustering</p>
<p><strong>Article References:</strong> Liang, N., Yang, Z., Xiang, D., Liu, M., Xiong, J., &amp; Yang, J. (2026). Safe maximin multiple graph learning for multi-view clustering. <em>Applied Intelligence, 56</em>(15), Article 456. <a href="https://doi.org/10.1007/s10489-026-07307-w" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07307-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07307-w" rel="noopener noreferrer">10.1007/s10489-026-07307-w</a></p>
<p><strong>Keywords:</strong> multi-view clustering, safe machine learning, graph learning, maximin optimization, graph fusion, unsupervised learning, spectral clustering, clustering safety guarantee, Applied Intelligence, machine learning theory, data fusion, benchmark datasets</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">229879</post-id>	</item>
		<item>
		<title>Speeding Up Graph Clustering: A New Survey Maps the Fast Lane</title>
		<link>https://scienmag.com/speeding-up-graph-clustering-a-new-survey-maps-the-fast-lane/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 00:13:34 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[advancements in graph clustering techniques]]></category>
		<category><![CDATA[anchor points]]></category>
		<category><![CDATA[bipartite graph]]></category>
		<category><![CDATA[community detection in social networks]]></category>
		<category><![CDATA[computational challenges in graph clustering]]></category>
		<category><![CDATA[density peaks clustering]]></category>
		<category><![CDATA[eigenvalue decomposition]]></category>
		<category><![CDATA[eigenvalue decomposition in clustering]]></category>
		<category><![CDATA[fast graph clustering algorithms]]></category>
		<category><![CDATA[graph clustering]]></category>
		<category><![CDATA[Graph clustering scalability]]></category>
		<category><![CDATA[graph cut]]></category>
		<category><![CDATA[image segmentation using graph methods]]></category>
		<category><![CDATA[irregular shape data segmentation]]></category>
		<category><![CDATA[label propagation]]></category>
		<category><![CDATA[large-scale data clustering]]></category>
		<category><![CDATA[multi-view clustering]]></category>
		<category><![CDATA[non-negative matrix factorization]]></category>
		<category><![CDATA[scalable machine learning]]></category>
		<category><![CDATA[similarity graph construction]]></category>
		<category><![CDATA[spectral clustering]]></category>
		<category><![CDATA[spectral embedding in graph clustering]]></category>
		<category><![CDATA[survey of scalable graph clustering methods]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224494</guid>

					<description><![CDATA[A new comprehensive survey in Vicinagearth organizes the rapidly growing field of fast graph clustering, from anchor-based bipartite graphs to multi-view acceleration, revealing how researchers are making large-scale unsupervised learning computationally feasible.]]></description>
										<content:encoded><![CDATA[<p>Graph clustering has long been one of the quiet workhorses of modern data science. Unlike feature-driven methods such as k-means, which struggle when data do not form neat convex blobs, graph clustering treats data points as vertices connected by edges and groups them according to the structure of those connections. That makes it remarkably good at discovering clusters of arbitrary shapes, from tangled social communities to irregular image segments. But there is a catch: as datasets have ballooned into millions or billions of points, the classical machinery of graph clustering has begun to buckle under its own computational weight. A comprehensive new survey, published in the open-access journal Vicinagearth by Jingjing Xue, Liyin Xing, Feiping Nie, Xuelong Li and colleagues at Northwestern Polytechnical University and China Telecom&#8217;s TeleAI, now offers the first systematic map of the fast graph clustering landscape, cataloguing the tricks that researchers have devised to make these methods scale.</p>
<p>The core problem, the authors explain, lies in the standard three-step pipeline that most graph clustering algorithms follow. First, an n-by-n similarity graph is constructed, where n is the number of data points. Second, the algorithm performs an eigenvalue decomposition (EVD) of that graph to obtain a spectral embedding. Third, the continuous embedding is discretized, typically with k-means or spectral rotation, to produce final cluster assignments. Each step is expensive. Building the graph costs on the order of n squared operations, and the eigenvalue decomposition on a dense graph costs on the order of n cubed. For a dataset with a million points, those scalings translate into computations that are simply out of reach. The survey&#8217;s central contribution is to organize the many acceleration strategies that have emerged into a coherent taxonomy, splitting them into single-view and multi-view families, and within each family into large graph methods and bipartite graph methods.</p>
<p>Large graph methods attack the problem head-on while keeping the full graph. Within the graph cut tradition, the survey distinguishes approximate and non-approximate approaches. Approximate methods, such as the Nyström technique, KASP and power iteration clustering, speed up the spectral embedding step by working on sampled columns of the similarity matrix or by iteratively estimating eigenvectors. These methods inherit a weakness: because they relax the discrete clustering problem and then post-process the continuous solution, the final result can drift far from the true optimum of the original cut objective, and the post-processing itself is not unique. Non-approximate methods take a different route, solving the graph cut model directly and thereby avoiding the eigenvalue decomposition altogether. Algorithms such as Fast-CD, which uses coordinate descent to update the cluster indicator matrix without any auxiliary variables, reduce the time complexity from n cubed to the number of edges, or n squared in the worst dense case. Related models like SBMC and EBMC add balanced regularization terms that prevent the trivial solution of isolating a few objects as a cluster, and they run in linear time.</p>
<p>The second major family, graph density methods, sidesteps the need to specify the number of clusters in advance, a decisive advantage when that number is unknown. Density-based spatial clustering of applications with noise, or DBSCAN, and density peaks clustering, or DPC, both identify clusters by separating high-density regions from low-density valleys in the data space. Their Achilles heel is a time complexity of n squared and a proliferation of parameters that weakens generalization. The survey documents a wave of accelerations built on k-nearest-neighbor searches, sampling strategies and grid-based indexing. Methods such as KNN-DBSCAN restrict density calculations to each point&#8217;s top-k neighbors, achieving linear computational and storage costs, while FastDPeak, DenPEHC and ADPC-KNN accelerate the density peak search or combine it with k-means to improve scalability. Distributed implementations on MapReduce-style frameworks extend these ideas to truly massive datasets.</p>
<p>Bipartite graph methods represent the survey&#8217;s most conceptually elegant acceleration strategy. Rather than wrestling with the full n-by-n graph, these methods select a small set of m representative anchor points and construct a compact n-by-m bipartite graph that records each sample&#8217;s affinity to each anchor. The full graph can then be approximated as a product involving the bipartite graph, and crucially, spectral analysis can be performed on a small m-by-m matrix instead of the original one, cutting the complexity from n cubed to n times m squared. The survey carefully dissects how anchors are generated, comparing random selection, k-means strategies, balanced k-means based hierarchical k-means (BKHK), variance-based de-correlation anchor selection (VDA), anchor learning with graph (ALG) and directly alternate sampling (DAS). BKHK stands out for producing stable, representative anchors efficiently, while VDA and ALG avoid random initialization entirely and better capture the intrinsic structure of the data. When the input is already a graph, such as a social network or citation network, label propagation algorithms can learn the compact bipartite representation directly from the edge structure.</p>
<p>On top of this bipartite scaffolding, the survey identifies three distinct clustering strategies. Graph cut methods such as FSC, LSC and FNC either apply singular value decomposition to the small anchor matrix or optimize the bipartite cut model directly, with algorithms like FDBC and GCSED pushing complexity down to linear time. Co-clustering methods exploit the duality between samples and features, grouping rows and columns of a data matrix simultaneously, which is particularly valuable for text and gene expression data; approaches range from Dhillon&#8217;s bipartite spectral graph partitioning to non-negative matrix tri-factorization and information-theoretic schemes. Label transmission methods go one step further: instead of updating an n-by-c label matrix through every iteration, they transmit label information from the m anchors to all samples through the relation Y = BU, so that only an m-by-c matrix needs updating. The FCAG algorithm achieves this while avoiding trivial solutions without extra parameters, and its iteration cost is entirely independent of the number of samples.</p>
<p>The multi-view setting, where the same entities are described by several heterogeneous feature sets or graphs, multiplies the computational burden, and the survey shows how the same two families of accelerations carry over. Early fusion methods first merge all views into a single weighted fusion graph and then cluster it; late fusion methods cluster each view separately and align the resulting embeddings. In both cases, the eigenvalue decomposition of Laplacian matrices remains the bottleneck, so fast multi-view algorithms either bypass it entirely or adopt anchors. Algorithms such as OMSC, FMVPG and FMDC obtain discrete cluster indicators directly through non-negative embeddings, spectral rotation or two-step optimization, avoiding EVD and reaching linear complexity. Anchor-based multi-view methods, including SFMC, BIGMC, FMCNOF and EMKMC, construct per-view bipartite graphs, often sharing a common anchor set selected by k-means on the union of views, and fuse them as weighted sums. Tensor-based approaches like TBGL and SWAGL add low-rank tensor regularization to capture complementary structure across views, at a steep price in running time.</p>
<p>The survey is not purely theoretical. The authors benchmark dozens of representative algorithms on standard single-view datasets, including face image collections such as AR, Face and Umist, the emotion recognition set CK, the speech dataset Isolet and the gesture dataset Palm, and on multi-view benchmarks including MSRC, ORL, YaleB, Wikipedia articles, Caltech101, Scene, Digit and MNIST, measuring accuracy and normalized mutual information over twenty repeated runs. Their findings are refreshingly candid: no single algorithm dominates across all data, so matching the method to the dataset remains essential. In running time, FCAG, Fast-CD and KASP emerge as the three fastest single-view methods, each for a different reason: KASP performs eigenvalue decomposition on small sampled matrices, Fast-CD solves the cut model directly without EVD, and FCAG updates only small label matrices through anchor guidance. In the multi-view experiments, MVFCAG excels on scene images, TBGL and SFMC lead on text and handwritten digits, and non-negative matrix factorization methods suit face images, while TBGL&#8217;s tensor machinery makes it the slowest of the bunch.</p>
<p>The survey closes with a sober look at what remains unsolved. Many fast algorithms trade accuracy for speed and can falter on noisy or incomplete data; robustness to outliers is still an open challenge. Parameter sensitivity, whether in the number of anchors, the sparsity of the graph or density thresholds, often demands domain expertise, motivating the search for parameter-free formulations. Dynamic graphs that evolve over time, and semi-supervised settings where scarce labels could guide clustering, both represent fertile ground for future work. Yet the trajectory is clear: fast graph clustering now delivers real-time or near real-time results on large-scale data with modest memory footprints, making it deployable on embedded systems and edge platforms and compatible with parallel and GPU computing. From disrupting criminal networks in social media to analyzing surveillance video and intelligence documents, the methods catalogued in this survey are quietly reshaping what is computationally possible in unsupervised learning.</p>
<p><strong>Subject of Research:</strong> Fast graph clustering algorithms for large-scale single-view and multi-view data</p>
<p><strong>Article Title:</strong> A comprehensive survey of fast graph clustering</p>
<p><strong>Article References:</strong> Xue, J., Xing, L., Wang, Y., Fan, X., Kong, L., Zhang, Q., Nie, F., &amp; Li, X. (2024). A comprehensive survey of fast graph clustering. <em>Vicinagearth, 1</em>(1), Article 7. <a href="https://doi.org/10.1007/s44336-024-00008-3" rel="noopener noreferrer">https://doi.org/10.1007/s44336-024-00008-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-024-00008-3" rel="noopener noreferrer">10.1007/s44336-024-00008-3</a></p>
<p><strong>Keywords:</strong> graph clustering, spectral clustering, bipartite graph, anchor points, graph cut, density peaks clustering, multi-view clustering, eigenvalue decomposition, non-negative matrix factorization, label propagation, scalable machine learning, unsupervised learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224494</post-id>	</item>
		<item>
		<title>New Tensor Method Cleans Up Messy Multi-View Data in One Efficient Step</title>
		<link>https://scienmag.com/new-tensor-method-cleans-up-messy-multi-view-data-in-one-efficient-step/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 17:08:16 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced data mining methods]]></category>
		<category><![CDATA[alternating optimization]]></category>
		<category><![CDATA[cross-view consistency]]></category>
		<category><![CDATA[data mining]]></category>
		<category><![CDATA[efficient data clustering techniques]]></category>
		<category><![CDATA[handling sensor failures and privacy restrictions]]></category>
		<category><![CDATA[incomplete multi-view clustering]]></category>
		<category><![CDATA[incomplete multi-view data]]></category>
		<category><![CDATA[low-frequency nuclear norm]]></category>
		<category><![CDATA[low-rank approximation]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[missing data]]></category>
		<category><![CDATA[missing data handling in machine learning]]></category>
		<category><![CDATA[multi-channel dataset integration]]></category>
		<category><![CDATA[multi-view clustering]]></category>
		<category><![CDATA[multi-view data fusion]]></category>
		<category><![CDATA[multi-view data imputation]]></category>
		<category><![CDATA[one-step optimization]]></category>
		<category><![CDATA[spectral clustering]]></category>
		<category><![CDATA[tensor learning]]></category>
		<category><![CDATA[tensor low-frequency learning]]></category>
		<category><![CDATA[tensor-based data analysis]]></category>
		<category><![CDATA[TLF-IMVC method]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217314</guid>

					<description><![CDATA[A new one-step tensor learning method uses a novel low-frequency nuclear norm to jointly filter noise and cluster incomplete multi-view data efficiently.]]></description>
										<content:encoded><![CDATA[<p>Modern datasets rarely arrive through a single channel. A video clip may be described by its pixels, its audio track, its subtitles, and the metadata surrounding it; a patient may be characterized by imaging scans, laboratory tests, and clinical notes. Each of these parallel descriptions is called a view, and combining them so that a machine learning algorithm can group similar examples together is the task of multi-view clustering. In the real world, however, this task is complicated by a stubborn inconvenience: some views are simply missing. Sensor failures, incomplete records, privacy restrictions, and expensive measurement procedures all conspire to leave gaps in the data, and the field that grapples with this problem is known as incomplete multi-view clustering. A newly published study in the journal Data Mining and Knowledge Discovery proposes a method called one-step Tensor Low-Frequency learning for Incomplete Multi-View Clustering, or TLF-IMVC, that addresses several long-standing weaknesses of existing approaches at once.</p>
<p>The work, authored by Lisha Zhao and Hongwei Ge of Jiangnan University and Shuzhi Su of Anhui University of Science and Technology, targets three problems that the authors identify as persistent obstacles in the literature. The first is computational cost: many tensor-based methods are powerful but slow. The second is a limited ability to characterize higher-order consistency, meaning the structured agreement that should exist across all views simultaneously rather than merely pairwise. The third, and perhaps most subtle, is the degradation that arises from two-step optimization frameworks, in which an algorithm first repairs or completes the missing data and then performs clustering on the result as if the two operations were independent. TLF-IMVC is designed to dissolve that separation, fusing the steps into a single joint optimization.</p>
<p>To understand why the two-step design is problematic, it helps to picture what happens when an algorithm fills in missing views before clustering. The imputation stage makes decisions based on incomplete information, and whatever errors it introduces are then frozen into the completed dataset. The clustering stage has no way to push back and correct those errors, because it only ever sees the finished product. Errors compound rather than cancel. By contrast, a one-step framework lets the clustering objective directly shape how missing information is recovered, so that the two processes inform each other continuously during optimization. The authors of the new study adopt this philosophy, joining the estimation of cluster structure and the exploitation of cross-view relationships inside a unified mathematical formulation.</p>
<p>The mathematical heart of the method is the tensor, a higher-dimensional generalization of the matrix. Where a matrix is a grid of numbers with rows and columns, a tensor can stack many such grids into a three-dimensional or higher-order object. In multi-view clustering, a natural construction is to take the spectral embeddings, the low-dimensional coordinate representations produced by spectral clustering, for each view and stack them across views to form a tensor. This stacked object encodes cross-view correlations as higher-order structure. Tensor-based methods have attracted considerable attention precisely because they can model these higher-order relationships, which pairwise matrix techniques inevitably flatten away. TLF-IMVC builds its core representation on exactly such stacked spectral embedding tensors.</p>
<p>What distinguishes the new approach is the particular structure it imposes on that tensor. The method jointly models two properties: low-frequency structure and low-rank structure. The low-rank assumption, familiar from decades of work on robust principal component analysis and related techniques, holds that the essential information in the data lives in a small number of dominant patterns, so that the tensor can be well approximated by one with far fewer degrees of freedom. The low-frequency assumption is newer in this context and more evocative. It treats the tensor as a signal that can be decomposed, in effect, into components of varying frequency, with genuine cluster-consistent information concentrated in the smooth, slowly varying low-frequency components while noise and artifacts accumulate in the high-frequency ones.</p>
<p>To make this idea operational, the authors introduce what they call a Tensor Low-Frequency nuclear norm. A nuclear norm is a standard device in low-rank optimization: it is a convex surrogate for the rank of a matrix or tensor, and minimizing it encourages solutions to be low-rank without requiring one to know the rank in advance. The novel norm proposed in this study adds a discriminative frequency-domain constraint, meaning it does not treat all components of the tensor equally. Instead, it selectively preserves the informative low-frequency components, which carry the consistent cross-view signal, and suppresses the noisy high-frequency components, which tend to encode corruption and missing-view artifacts. The result is a regularizer that simultaneously encourages low-rank structure and frequency-domain cleanliness, filtering the data representation as part of the optimization itself rather than as a separate preprocessing step.</p>
<p>Solving the resulting optimization problem is nontrivial, since it involves coupled variables, tensor operations, and non-smooth regularizers. The authors develop an efficient alternating optimization algorithm, a strategy in which the variables are updated in turn, each optimized while the others are held fixed, cycling until convergence. Alternating schemes of this kind are the workhorses of multi-view clustering research, and their theoretical behavior under nonconvex settings has been studied extensively in the optimization literature the paper builds upon. The practical payoff claimed for the new algorithm is efficiency: by operating on the stacked spectral embeddings and exploiting the joint one-step formulation, TLF-IMVC avoids the heavy cost of completing raw missing views while still enforcing consistency across all of them.</p>
<p>The empirical evaluation spans eight benchmark datasets, on which the authors report extensive experiments comparing TLF-IMVC against state-of-the-art incomplete multi-view clustering methods. According to the study, the results demonstrate the effectiveness and competitiveness of the proposed approach, supporting the central claims that joint one-step optimization, tensor-based higher-order modeling, and frequency-domain regularization combine into a method that outperforms alternatives that address these issues in isolation. The datasets used in the experiments are publicly available from standard repositories, and the analysis code and supplementary materials are available from the corresponding author upon reasonable request, a transparency measure that should make it straightforward for other researchers to verify and extend the results.</p>
<p>The significance of the work lies less in any single technical gadget than in the direction it points. Incomplete multi-view clustering has become a crowded subfield, with a steady stream of methods based on matrix completion, graph learning, contrastive prediction, anchor graphs, and deep generative networks, and tensor frameworks have featured prominently among them. What TLF-IMVC contributes is a specific answer to the question of what structure the fused representation should obey. By identifying low-frequency content as the signature of genuine cross-view consistency and encoding that intuition directly into a tensor nuclear norm, the method offers a principled filter that is learned jointly with the clustering rather than bolted on beforehand. If the frequency-domain perspective proves portable, it could inform the design of future methods well beyond the specific algorithm introduced here.</p>
<p>For practitioners, the practical appeal is the combination of accuracy and efficiency in a setting that is all too common. Data pipelines in medicine, multimedia analysis, and sensor networks routinely produce multi-view collections with missing entries, and methods that are either too slow or too brittle to handle the gaps end up shelved. A one-step framework that jointly handles recovery and clustering, grounded in a theoretically motivated norm, addresses both concerns. The study, published in volume 40 of Data Mining and Knowledge Discovery as article 108, arrived after peer review beginning in April 2026 and appearing in September 2026, and was supported by funding from the National Natural Science Foundation of China and provincial science foundations. As datasets grow richer and messier in parallel, techniques of this kind, which extract clean, consistent structure from noisy, incomplete observations, are likely to become an increasingly standard part of the machine learning toolkit.</p>
<p><strong>Subject of Research:</strong> One-step tensor low-frequency learning for clustering data with incomplete multiple views</p>
<p><strong>Article Title:</strong> One-step tensor low-frequency learning for incomplete multi-view clustering</p>
<p><strong>Article References:</strong> Zhao, L., Ge, H., &amp; Su, S. (2026). One-step tensor low-frequency learning for incomplete multi-view clustering. <em>Data Mining and Knowledge Discovery, 40</em>(6), Article 108. <a href="https://doi.org/10.1007/s10618-026-01263-2" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01263-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01263-2" rel="noopener noreferrer">10.1007/s10618-026-01263-2</a></p>
<p><strong>Keywords:</strong> incomplete multi-view clustering, tensor learning, low-frequency nuclear norm, spectral clustering, low-rank approximation, one-step optimization, cross-view consistency, data mining, machine learning, alternating optimization, missing data, unsupervised learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217314</post-id>	</item>
		<item>
		<title>Simple Fix Guarantees Connected Graphs for Robust Text Clustering</title>
		<link>https://scienmag.com/simple-fix-guarantees-connected-graphs-for-robust-text-clustering/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:47:22 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[document clustering]]></category>
		<category><![CDATA[Document similarity measurement]]></category>
		<category><![CDATA[Ensuring connected graphs for reliable topic discovery]]></category>
		<category><![CDATA[Epsilon-threshold graph construction]]></category>
		<category><![CDATA[graph connectivity]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[Hyperparameter sensitivity in graph construction]]></category>
		<category><![CDATA[incremental graph construction]]></category>
		<category><![CDATA[k-nearest neighbor graphs]]></category>
		<category><![CDATA[K-nearest-neighbor graph fragility]]></category>
		<category><![CDATA[Laplacian eigenmaps]]></category>
		<category><![CDATA[Machine Learning journal]]></category>
		<category><![CDATA[minimum spanning tree]]></category>
		<category><![CDATA[MTEB]]></category>
		<category><![CDATA[Near-duplicate detection in text datasets]]></category>
		<category><![CDATA[Neighborhood graph in text mining]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[Robust text clustering techniques]]></category>
		<category><![CDATA[Semi-supervised label propagation]]></category>
		<category><![CDATA[SentenceTransformer embeddings]]></category>
		<category><![CDATA[SentenceTransformers]]></category>
		<category><![CDATA[spectral clustering]]></category>
		<category><![CDATA[Spectral clustering challenges]]></category>
		<category><![CDATA[text embeddings]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204096</guid>

					<description><![CDATA[Researchers have developed an incremental k-nearest-neighbor graph construction that guarantees connectivity by design, dramatically improving spectral clustering of text embeddings at low sparsity while cutting build time and memory costs.]]></description>
										<content:encoded><![CDATA[<p>One of the quiet workhorses of modern text mining is a deceptively simple data structure: a neighborhood graph, in which every document becomes a node and edges link each document to the items most similar to it. These graphs underpin topic discovery, near-duplicate detection, semi-supervised label propagation, and the retrieval indices used in retrieval-augmented generation. Yet, according to a new study published in the journal Machine Learning, this foundational step is far more fragile than most practitioners realize. On realistic text datasets, the standard k-nearest-neighbor graphs that power spectral clustering can shatter into many disconnected components at the sparsity levels people actually use in practice, silently degrading results and making the whole pipeline hypersensitive to a single hyperparameter.</p>
<p>The study, led by Marko Pranjić and Boshko Koloski of the Jožef Stefan Institute together with Nada Lavrač, Senja Pollak, and Marko Robnik-Šikonja of the University of Ljubljana, quantifies just how badly things can go wrong. Working with the well-known 20 Newsgroups collection of roughly 20,000 documents, the researchers encoded the test split of 7,532 documents into 384-dimensional vectors using the SentenceTransformer model all-MiniLM-L12-v2 and measured cosine distance between the embeddings. When they built epsilon-threshold graphs, connecting every pair of documents closer than a given distance, they found that even the minimal distance required to keep the graph connected demanded around one million edges, and increasing that distance by only five percent inflated the edge count by more than sixty percent. Sparser is better for memory and speed, but sparser also means disconnected.</p>
<p>The k-nearest-neighbor approach fared somewhat better, since every document is linked to its k most similar documents, and tuning k to obtain a connected graph proved easier than tuning a distance threshold. But connectedness is never guaranteed. Theoretical work cited in the study shows that a k-NN graph is only asymptotically connected with high probability when k exceeds roughly 5.1774 times the natural logarithm of the number of points, which for as few as 300 data points already means k above 30, far beyond the small values used in real pipelines. In the experiments, some dataset partitions within the TwentyNewsgroups benchmark exhibited disconnected components even at k equals 15, and a Reddit sentence-clustering dataset remained disconnected at k equals 20. The high-dimensional geometry of embeddings makes matters worse: a phenomenon known as hubness means a few points appear in a disproportionate number of neighbor lists while others become isolated anti-hubs, skewing degree distributions and aggravating disconnection.</p>
<p>Why does disconnection matter so much? In spectral clustering, each connected component of the graph can only be assigned to a single cluster. When the number of components equals or exceeds the number of desired clusters, the clustering becomes trivial, and no similarity-based criterion can rescue it. The same fragility propagates beyond clustering: in item-based recommender systems, unreachable items simply cannot be recommended, and in label propagation, an isolated component of documents can never receive a propagated label. Connectivity, in other words, determines reachability, and reachability determines whether the downstream algorithm can function at all.</p>
<p>The researchers&#8217; solution is elegantly minimal. Instead of letting every node search for neighbors across the entire dataset simultaneously, their incremental construction inserts documents one at a time, and each new node is linked to its k nearest neighbors among the nodes already present in the graph. That single restriction, searching only among previously inserted nodes, guarantees a connected graph for any value of k. The proof is a short induction: the first insertion creates a single connected component, and every subsequent node attaches to k existing members of that component, extending it without ever splitting it. Because the first node of a new region can only attach to earlier nodes, the graph remains one piece by construction, for every k and every insertion order.</p>
<p>The modification also confers a practical bonus that is increasingly relevant in the era of streaming data. In a standard k-NN graph, adding a single new document can trigger a large reconfiguration, since many existing nodes might count the newcomer among their nearest neighbors, effectively forcing a rebuild. The incremental construction induces only local changes, restricted to the row and column of the newly added node in the adjacency matrix. New documents can therefore be inserted without reconstructing the existing graph, and the last added nodes can even be deleted, properties the authors argue could be exploited in applications where data arrives continuously or becomes invalidated over time.</p>
<p>To test whether the guaranteed connectivity translates into better clustering, the team evaluated spectral clustering with Laplacian eigenmaps across eleven sentence- and paragraph-level clustering tasks drawn from six dataset sources in the Massive Text Embedding Benchmark, covering 182 distinct clustering problems in total. Performance was measured with V-measure, the harmonic combination of homogeneity and completeness. The results show a clear pattern: in the low-k regime, where standard k-NN graphs are most prone to fragmentation, the incremental approach consistently outperformed the standard construction, achieving near-top scores already at k equals 3 on many datasets, while matching the standard graph at larger k where both methods operate on connected structures. The largest gains appeared on TwentyNewsgroups, precisely where disconnection was most prevalent.</p>
<p>The team then compared their method against the most obvious alternative: repairing the standard k-NN graph by augmenting it with a minimum spanning tree, a global structure that stitches disconnected pieces together. Under matched MST augmentation, the incremental graph still held the edge in the sparse regime, improving on k-NN plus MST by an average of 2.5 V-measure points for k between 1 and 3, and by 3.8 points at k equals 1. A Bayesian signed-rank analysis with a region of practical equivalence confirmed that the MST is practically beneficial only in the sparsest regime and practically irrelevant from k equals 6 onward. Crucially, the exact MST computation requires forming the dense N-by-N distance matrix, an O(N-squared) step. The incremental construction avoids that matrix entirely because it attains connectivity through the insertion order rather than through added global edges, and it proved between 23 and 44 times faster to build than the MST-augmented alternatives, with peak memory dropping from as much as 20.2 gigabytes to as little as 1.2 gigabytes on the largest evaluated partitions, run on a CPU node of the Vega EuroHPC system.</p>
<p>Two natural objections receive careful treatment in the paper. First, because the incremental graph depends on the order in which nodes are inserted, is the method stable? Across ten randomized orderings per dataset, the standard deviation of clustering performance rarely exceeded one percent and was often below half a percent, and the structural statistics of the resulting graphs, including transitivity, assortativity, and label homophily, varied only in the third decimal place. Even adversarial orderings, in which the first inserted nodes were deliberately concentrated near the embedding centroid or drawn entirely from a single class, changed V-measure by at most 0.4 points relative to random ordering. Second, does the approach depend on the specific embedding model? Tests with larger models such as bge-base-en-v1.5 and all-mpnet-base-v2 showed that bigger encoders consistently improve results, though the benefit of denser graphs and larger models proved dataset-dependent, with Reddit clustering the notable exception.</p>
<p>The authors are candid about limitations. The five largest Reddit paragraph-level partitions exceeded the 32-bit index range of the sparse eigensolver and were excluded from some recomputed comparisons, and the robustness study covers concentrated initial seeds but not fully temporally ordered document streams, since the benchmark snapshots carry no timestamps. Future directions include replacing exact nearest-neighbor search with approximate methods, coupling the incremental graph with efficient eigenvector update techniques, and deploying the construction in temporal community detection. Still, the message is striking: a one-line change to how documents are inserted into a similarity graph eliminates a failure mode that has quietly lurked in spectral text clustering, delivering guaranteed connectivity, faster builds, lower memory, and better low-k performance, all without needing a proper metric, global graph features, or any repair step at all.</p>
<p><strong>Subject of Research:</strong> Incremental k-nearest-neighbor graph construction for robust spectral clustering of text embeddings</p>
<p><strong>Article Title:</strong> Incremental Graph Construction Enables Robust Spectral Clustering of Texts</p>
<p><strong>Article References:</strong> Incremental Graph Construction Enables Robust Spectral Clustering of Texts. (n.d.). <a href="https://doi.org/10.1007/s10994-026-07158-z" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07158-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07158-z" rel="noopener noreferrer">10.1007/s10994-026-07158-z</a></p>
<p><strong>Keywords:</strong> spectral clustering, incremental graph construction, k-nearest neighbor graphs, text embeddings, Laplacian eigenmaps, graph connectivity, minimum spanning tree, Machine Learning journal, SentenceTransformers, MTEB, document clustering, high-dimensional data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204096</post-id>	</item>
	</channel>
</rss>
