<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Machine Learning journal &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/machine-learning-journal/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 00:23:31 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Machine Learning journal &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Teaching Graph AI to Forget: SURGE Rewrites the Rules of Machine Unlearning</title>
		<link>https://scienmag.com/teaching-graph-ai-to-forget-surge-rewrites-the-rules-of-machine-unlearning/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 00:23:31 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[CCPA data privacy]]></category>
		<category><![CDATA[challenges of model retraining for data deletion]]></category>
		<category><![CDATA[citation analysis data privacy]]></category>
		<category><![CDATA[data deletion]]></category>
		<category><![CDATA[fraud detection data removal]]></category>
		<category><![CDATA[GDPR]]></category>
		<category><![CDATA[GDPR compliance in AI]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[graph unlearning]]></category>
		<category><![CDATA[Jensen-Shannon divergence]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[Machine Learning journal]]></category>
		<category><![CDATA[machine unlearning]]></category>
		<category><![CDATA[Machine unlearning in graph neural networks]]></category>
		<category><![CDATA[membership inference]]></category>
		<category><![CDATA[molecular discovery data security]]></category>
		<category><![CDATA[neural network model erasure]]></category>
		<category><![CDATA[ogbn-arxiv]]></category>
		<category><![CDATA[privacy]]></category>
		<category><![CDATA[recommendation system data privacy]]></category>
		<category><![CDATA[response field]]></category>
		<category><![CDATA[scalable machine unlearning techniques]]></category>
		<category><![CDATA[structural perturbation in neural networks]]></category>
		<category><![CDATA[SURGE framework for model forgetting]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=245886</guid>

					<description><![CDATA[Researchers have developed SURGE, a graph unlearning framework that treats data deletion requests as structural perturbations and repairs graph neural networks through response-field sensing and adaptive distillation.]]></description>
										<content:encoded><![CDATA[<p>Every time a user invokes the right to be forgotten under Europe&#8217;s GDPR or California&#8217;s CCPA, companies face an uncomfortable technical reality: their machine learning models may still carry traces of the very data they have been ordered to erase. For graph neural networks, the models that power recommendation engines, fraud detection, citation analysis, and molecular discovery, the problem is even harder. Information in a graph does not sit in isolated slots; it flows through connections, so removing a single node or edge can ripple across the entire learned representation. A new study published in the journal Machine Learning proposes a framework called SURGE, which treats deletion requests not as surgical cuts to model parameters but as structural perturbations whose effects can be sensed, repaired, and distilled across the whole network.</p>
<p>The work, authored by Chaofan Shen, Mingyu Wang, and Jing Zhang of Southeast University&#8217;s School of Cyber Science and Engineering in Nanjing, addresses a gap that has frustrated researchers since machine unlearning emerged as a field. Full retraining from scratch after every deletion request is the gold standard for forgetting, but it is prohibitively expensive for large graphs and models that may need to serve thousands of erasure requests. Approximate unlearning methods promise speed, yet existing approaches fall into camps with distinct weaknesses. Some partition the graph so that deleting a shard leaves the rest of the model untouched, but this disrupts the very structure that makes graph learning powerful. Others approximate removal directly in parameter space, editing weights with influence functions or closed-form corrections, but they can leave residual traces of deleted data. Still others rely on fixed neighborhoods when updating predictions, which cannot capture the fact that some nodes are far more sensitive to a deletion than others.</p>
<p>SURGE&#8217;s central insight is conceptual: a deletion request is fundamentally a prediction-level structural perturbation. When an edge disappears from a citation network, the predictions for nodes near that edge shift in ways that depend on how information propagates through the graph. Rather than guessing which parameters to edit, SURGE asks how the model&#8217;s outputs would change if the requested data were genuinely removed, and then repairs the model to match that ideal. The framework unfolds in three stages that the authors describe as Sense, Repair, and Distill.</p>
<p>In the sensing stage, SURGE estimates a continuous response field over all nodes in the graph. Instead of treating every node identically, the response field quantifies how strongly each node&#8217;s prediction is affected by the structural change introduced by the deletion. Nodes directly connected to removed data register strong responses; distant nodes register weak ones. This continuous, per-node characterization is what allows the method to express heterogeneous node sensitivity, something the authors argue fixed-neighborhood schemes cannot do. The response field effectively becomes a map of where forgetting must be deep and where it can be gentle.</p>
<p>The repair stage then solves for corrected logits, the raw pre-softmax outputs of the classifier, across the graph. Crucially, this is formulated as a graph-regularized optimization problem that is solved without backpropagation. Rather than running gradient descent through the network, SURGE draws on techniques from numerical linear algebra for sparse systems, computing corrections that respect both the deletion constraints and the smoothness structure of the graph. The graph regularizer ensures that corrected predictions remain coherent with their neighbors, preventing the kind of fragmented, inconsistent outputs that can arise when unlearning methods patch predictions locally. Because no gradients flow through the deep network, the repair step avoids the cost and instability of fine-tuning while still producing outputs that approximate what full retraining would yield.</p>
<p>The final stage, distillation, transfers the corrected knowledge back into the model. Knowledge distillation, a technique originally popularized for compressing large teacher networks into smaller students, is repurposed here as a forgetting mechanism. But SURGE adapts it: the distillation is response-adaptive, meaning the influence of the corrected teacher predictions on each node is weighted by that node&#8217;s position in the response field. Nodes that were strongly perturbed by the deletion receive aggressive correction, while nodes barely affected are nudged only lightly, preserving their original utility. This residual distillation is what gives the method its name and its balance between two competing goals: erasing the influence of deleted data and retaining accuracy on everything that remains.</p>
<p>The empirical case for SURGE rests on an unusually rigorous evaluation design. The authors tested the framework across seven standard benchmarks, including widely used citation corpora such as Cora, CiteSeer, and PubMed alongside larger graphs, and across multiple graph neural network backbones, including graph convolutional networks and graph attention networks. They evaluated three distinct deletion types: node unlearning, edge unlearning, and feature unlearning, covering the full range of erasure requests a deployed system might receive. To eliminate luck from the comparison, they ran a matched 18-cell evaluation spanning deletion types, datasets, backbones, and ten random seeds per configuration, comparing SURGE against three established baselines: ETR, an erase-then-rectify parameter editing approach; MEGU, a mutual-evolution unlearning method; and IDEA, a framework for certified graph unlearning.</p>
<p>The results are striking. Across the matched evaluation, SURGE achieved the highest mean micro-F1 score of 0.8632, indicating that models unlearned with SURGE retained the most predictive power on retained data. It also achieved the best retraining faithfulness, measured as a test Jensen-Shannon divergence of just 0.0243 against models retrained from scratch, meaning its post-deletion predictions stayed closest to the gold standard of full retraining. In other words, SURGE did not trade forgetting quality for utility or vice versa; it led on both axes simultaneously. The Jensen-Shannon divergence metric, rooted in Shannon entropy theory, provides a symmetric measure of how two probability distributions diverge, making it a natural yardstick for how faithfully an unlearned model mimics its fully retrained counterpart. An additional ablation study on ogbn-arxiv, a large-scale graph from the Open Graph Benchmark with millions of nodes, further supports the method&#8217;s scalability, suggesting the approach is not confined to small academic datasets.</p>
<p>The broader significance of this work lies at the intersection of privacy regulation and the economics of deployed AI. Membership inference attacks, demonstrated in influential security research over the past decade, can often detect whether a specific record was part of a model&#8217;s training data, turning residual traces into genuine privacy liabilities. As graph neural networks increasingly underpin systems that process personal relationships, transactions, and social connections, the ability to provably and efficiently remove an individual&#8217;s data becomes a compliance necessity rather than an academic curiosity. SURGE&#8217;s response-field formulation offers a template that could generalize: sense how a data change propagates, repair predictions to match the ideal, and distill the repair back into the model, all without the expense of retraining.</p>
<p>There remain open questions, as with any approximate unlearning method. The framework&#8217;s guarantees are empirical rather than cryptographic, and the field continues to debate what level of assurance suffices for regulatory compliance. Certified unlearning approaches offer stronger theoretical promises but often at greater cost or with more restrictive assumptions. Still, the combination of leading utility, near-retraining faithfulness, backbone-agnostic design, and demonstrated scalability on a graph the size of ogbn-arxiv positions SURGE as one of the most complete graph unlearning solutions reported to date. Funded by the Basic Research Program of Jiangsu, the study signals that the era of treating forgetting as an afterthought in graph machine learning is coming to an end. As deletion requests multiply and regulators sharpen their scrutiny, methods that let models forget precisely, quickly, and faithfully may become as fundamental to trustworthy AI as the training algorithms themselves.</p>
<p><strong>Subject of Research:</strong> Efficient removal of nodes, edges, and features from trained graph neural networks without full retraining</p>
<p><strong>Article Title:</strong> SURGE: Structural Perturbation Response Field Guided Residual Distillation for Graph Unlearning</p>
<p><strong>Article References:</strong> Shen, C., Wang, M., &amp; Zhang, J. (2026). SURGE: Structural Perturbation Response Field Guided Residual Distillation for Graph Unlearning. <em>Machine Learning, 115</em>(10), Article 239. <a href="https://doi.org/10.1007/s10994-026-07176-x" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07176-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07176-x" rel="noopener noreferrer">10.1007/s10994-026-07176-x</a></p>
<p><strong>Keywords:</strong> graph unlearning, graph neural networks, machine unlearning, knowledge distillation, privacy, GDPR, data deletion, response field, Jensen-Shannon divergence, ogbn-arxiv, membership inference, Machine Learning journal</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">245886</post-id>	</item>
		<item>
		<title>Simple Fix Guarantees Connected Graphs for Robust Text Clustering</title>
		<link>https://scienmag.com/simple-fix-guarantees-connected-graphs-for-robust-text-clustering/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:47:22 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[document clustering]]></category>
		<category><![CDATA[Document similarity measurement]]></category>
		<category><![CDATA[Ensuring connected graphs for reliable topic discovery]]></category>
		<category><![CDATA[Epsilon-threshold graph construction]]></category>
		<category><![CDATA[graph connectivity]]></category>
		<category><![CDATA[high-dimensional data]]></category>
		<category><![CDATA[Hyperparameter sensitivity in graph construction]]></category>
		<category><![CDATA[incremental graph construction]]></category>
		<category><![CDATA[k-nearest neighbor graphs]]></category>
		<category><![CDATA[K-nearest-neighbor graph fragility]]></category>
		<category><![CDATA[Laplacian eigenmaps]]></category>
		<category><![CDATA[Machine Learning journal]]></category>
		<category><![CDATA[minimum spanning tree]]></category>
		<category><![CDATA[MTEB]]></category>
		<category><![CDATA[Near-duplicate detection in text datasets]]></category>
		<category><![CDATA[Neighborhood graph in text mining]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[Robust text clustering techniques]]></category>
		<category><![CDATA[Semi-supervised label propagation]]></category>
		<category><![CDATA[SentenceTransformer embeddings]]></category>
		<category><![CDATA[SentenceTransformers]]></category>
		<category><![CDATA[spectral clustering]]></category>
		<category><![CDATA[Spectral clustering challenges]]></category>
		<category><![CDATA[text embeddings]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204096</guid>

					<description><![CDATA[Researchers have developed an incremental k-nearest-neighbor graph construction that guarantees connectivity by design, dramatically improving spectral clustering of text embeddings at low sparsity while cutting build time and memory costs.]]></description>
										<content:encoded><![CDATA[<p>One of the quiet workhorses of modern text mining is a deceptively simple data structure: a neighborhood graph, in which every document becomes a node and edges link each document to the items most similar to it. These graphs underpin topic discovery, near-duplicate detection, semi-supervised label propagation, and the retrieval indices used in retrieval-augmented generation. Yet, according to a new study published in the journal Machine Learning, this foundational step is far more fragile than most practitioners realize. On realistic text datasets, the standard k-nearest-neighbor graphs that power spectral clustering can shatter into many disconnected components at the sparsity levels people actually use in practice, silently degrading results and making the whole pipeline hypersensitive to a single hyperparameter.</p>
<p>The study, led by Marko Pranjić and Boshko Koloski of the Jožef Stefan Institute together with Nada Lavrač, Senja Pollak, and Marko Robnik-Šikonja of the University of Ljubljana, quantifies just how badly things can go wrong. Working with the well-known 20 Newsgroups collection of roughly 20,000 documents, the researchers encoded the test split of 7,532 documents into 384-dimensional vectors using the SentenceTransformer model all-MiniLM-L12-v2 and measured cosine distance between the embeddings. When they built epsilon-threshold graphs, connecting every pair of documents closer than a given distance, they found that even the minimal distance required to keep the graph connected demanded around one million edges, and increasing that distance by only five percent inflated the edge count by more than sixty percent. Sparser is better for memory and speed, but sparser also means disconnected.</p>
<p>The k-nearest-neighbor approach fared somewhat better, since every document is linked to its k most similar documents, and tuning k to obtain a connected graph proved easier than tuning a distance threshold. But connectedness is never guaranteed. Theoretical work cited in the study shows that a k-NN graph is only asymptotically connected with high probability when k exceeds roughly 5.1774 times the natural logarithm of the number of points, which for as few as 300 data points already means k above 30, far beyond the small values used in real pipelines. In the experiments, some dataset partitions within the TwentyNewsgroups benchmark exhibited disconnected components even at k equals 15, and a Reddit sentence-clustering dataset remained disconnected at k equals 20. The high-dimensional geometry of embeddings makes matters worse: a phenomenon known as hubness means a few points appear in a disproportionate number of neighbor lists while others become isolated anti-hubs, skewing degree distributions and aggravating disconnection.</p>
<p>Why does disconnection matter so much? In spectral clustering, each connected component of the graph can only be assigned to a single cluster. When the number of components equals or exceeds the number of desired clusters, the clustering becomes trivial, and no similarity-based criterion can rescue it. The same fragility propagates beyond clustering: in item-based recommender systems, unreachable items simply cannot be recommended, and in label propagation, an isolated component of documents can never receive a propagated label. Connectivity, in other words, determines reachability, and reachability determines whether the downstream algorithm can function at all.</p>
<p>The researchers&#8217; solution is elegantly minimal. Instead of letting every node search for neighbors across the entire dataset simultaneously, their incremental construction inserts documents one at a time, and each new node is linked to its k nearest neighbors among the nodes already present in the graph. That single restriction, searching only among previously inserted nodes, guarantees a connected graph for any value of k. The proof is a short induction: the first insertion creates a single connected component, and every subsequent node attaches to k existing members of that component, extending it without ever splitting it. Because the first node of a new region can only attach to earlier nodes, the graph remains one piece by construction, for every k and every insertion order.</p>
<p>The modification also confers a practical bonus that is increasingly relevant in the era of streaming data. In a standard k-NN graph, adding a single new document can trigger a large reconfiguration, since many existing nodes might count the newcomer among their nearest neighbors, effectively forcing a rebuild. The incremental construction induces only local changes, restricted to the row and column of the newly added node in the adjacency matrix. New documents can therefore be inserted without reconstructing the existing graph, and the last added nodes can even be deleted, properties the authors argue could be exploited in applications where data arrives continuously or becomes invalidated over time.</p>
<p>To test whether the guaranteed connectivity translates into better clustering, the team evaluated spectral clustering with Laplacian eigenmaps across eleven sentence- and paragraph-level clustering tasks drawn from six dataset sources in the Massive Text Embedding Benchmark, covering 182 distinct clustering problems in total. Performance was measured with V-measure, the harmonic combination of homogeneity and completeness. The results show a clear pattern: in the low-k regime, where standard k-NN graphs are most prone to fragmentation, the incremental approach consistently outperformed the standard construction, achieving near-top scores already at k equals 3 on many datasets, while matching the standard graph at larger k where both methods operate on connected structures. The largest gains appeared on TwentyNewsgroups, precisely where disconnection was most prevalent.</p>
<p>The team then compared their method against the most obvious alternative: repairing the standard k-NN graph by augmenting it with a minimum spanning tree, a global structure that stitches disconnected pieces together. Under matched MST augmentation, the incremental graph still held the edge in the sparse regime, improving on k-NN plus MST by an average of 2.5 V-measure points for k between 1 and 3, and by 3.8 points at k equals 1. A Bayesian signed-rank analysis with a region of practical equivalence confirmed that the MST is practically beneficial only in the sparsest regime and practically irrelevant from k equals 6 onward. Crucially, the exact MST computation requires forming the dense N-by-N distance matrix, an O(N-squared) step. The incremental construction avoids that matrix entirely because it attains connectivity through the insertion order rather than through added global edges, and it proved between 23 and 44 times faster to build than the MST-augmented alternatives, with peak memory dropping from as much as 20.2 gigabytes to as little as 1.2 gigabytes on the largest evaluated partitions, run on a CPU node of the Vega EuroHPC system.</p>
<p>Two natural objections receive careful treatment in the paper. First, because the incremental graph depends on the order in which nodes are inserted, is the method stable? Across ten randomized orderings per dataset, the standard deviation of clustering performance rarely exceeded one percent and was often below half a percent, and the structural statistics of the resulting graphs, including transitivity, assortativity, and label homophily, varied only in the third decimal place. Even adversarial orderings, in which the first inserted nodes were deliberately concentrated near the embedding centroid or drawn entirely from a single class, changed V-measure by at most 0.4 points relative to random ordering. Second, does the approach depend on the specific embedding model? Tests with larger models such as bge-base-en-v1.5 and all-mpnet-base-v2 showed that bigger encoders consistently improve results, though the benefit of denser graphs and larger models proved dataset-dependent, with Reddit clustering the notable exception.</p>
<p>The authors are candid about limitations. The five largest Reddit paragraph-level partitions exceeded the 32-bit index range of the sparse eigensolver and were excluded from some recomputed comparisons, and the robustness study covers concentrated initial seeds but not fully temporally ordered document streams, since the benchmark snapshots carry no timestamps. Future directions include replacing exact nearest-neighbor search with approximate methods, coupling the incremental graph with efficient eigenvector update techniques, and deploying the construction in temporal community detection. Still, the message is striking: a one-line change to how documents are inserted into a similarity graph eliminates a failure mode that has quietly lurked in spectral text clustering, delivering guaranteed connectivity, faster builds, lower memory, and better low-k performance, all without needing a proper metric, global graph features, or any repair step at all.</p>
<p><strong>Subject of Research:</strong> Incremental k-nearest-neighbor graph construction for robust spectral clustering of text embeddings</p>
<p><strong>Article Title:</strong> Incremental Graph Construction Enables Robust Spectral Clustering of Texts</p>
<p><strong>Article References:</strong> Incremental Graph Construction Enables Robust Spectral Clustering of Texts. (n.d.). <a href="https://doi.org/10.1007/s10994-026-07158-z" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07158-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07158-z" rel="noopener noreferrer">10.1007/s10994-026-07158-z</a></p>
<p><strong>Keywords:</strong> spectral clustering, incremental graph construction, k-nearest neighbor graphs, text embeddings, Laplacian eigenmaps, graph connectivity, minimum spanning tree, Machine Learning journal, SentenceTransformers, MTEB, document clustering, high-dimensional data</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204096</post-id>	</item>
	</channel>
</rss>
