<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>adaptive clustering in data streams &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/adaptive-clustering-in-data-streams/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 06 Oct 2026 11:12:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>adaptive clustering in data streams &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Streaming Algorithm Maps Ever-Changing Data Schemas in Real Time</title>
		<link>https://scienmag.com/new-streaming-algorithm-maps-ever-changing-data-schemas-in-real-time/</link>
		
		<dc:creator><![CDATA[Gavin Prescott]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 11:12:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive clustering in data streams]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[cluster evolution]]></category>
		<category><![CDATA[clustering methods for evolving data schemas]]></category>
		<category><![CDATA[coresets]]></category>
		<category><![CDATA[data streams]]></category>
		<category><![CDATA[dynamic and unknown data domain analysis]]></category>
		<category><![CDATA[dynamic stream clustering algorithms]]></category>
		<category><![CDATA[handling evolving data attributes]]></category>
		<category><![CDATA[IoT]]></category>
		<category><![CDATA[k-means]]></category>
		<category><![CDATA[machine learning for streaming data schema adaptation]]></category>
		<category><![CDATA[Min-Hash]]></category>
		<category><![CDATA[novel algorithms for streaming data structure profiling]]></category>
		<category><![CDATA[real-time analytics]]></category>
		<category><![CDATA[real-time data format identification in big data]]></category>
		<category><![CDATA[Real-time data stream schema mapping]]></category>
		<category><![CDATA[real-time schema detection in IoT sensor data]]></category>
		<category><![CDATA[schema profiling]]></category>
		<category><![CDATA[schemaless data]]></category>
		<category><![CDATA[sliding window]]></category>
		<category><![CDATA[stream clustering]]></category>
		<category><![CDATA[unknown data formats in streaming data]]></category>
		<category><![CDATA[unsupervised clustering of chaotic data streams]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=241078</guid>

					<description><![CDATA[Researchers at the University of Bologna have developed DSC+, a two-phase streaming clustering algorithm that profiles the evolving schemas of heterogeneous data streams in real time without prior knowledge of the attribute domain.]]></description>
										<content:encoded><![CDATA[<p>Every second, millions of devices, applications, and sensors fire off messages into data streams, and almost none of them agree on a common format. A humidity sensor sends one set of fields, a weather station another, and a logistics platform something else entirely. For engineers trying to make sense of this flood, knowing what kinds of data are actually flowing through the pipes is a surprisingly hard problem. A team of researchers at the University of Bologna has now unveiled an algorithm, called DSC+ for Dynamic Stream Clustering, that can profile the structure of these chaotic streams on the fly, without ever knowing in advance what shapes the data will take.</p>
<p>The work, published in the journal Knowledge and Information Systems, tackles a setting the authors call a Dynamic and Unknown Domain, or DUD. It is dynamic because the schemas of incoming messages, meaning the set of attributes each message carries, evolve constantly in number, similarity, and distribution. It is unknown because there is no fixed dictionary of attributes; new fields can appear at any moment with names nobody has seen before. In such an environment, the classical assumption underlying most clustering algorithms, that data points live in a fixed-dimensional space, simply breaks down.</p>
<p>The goal of schema profiling is to give practitioners a summarized view of the message structures within a sliding window of the stream. That view is built by clustering similar schemas together and representing each group by its centroid, in the spirit of the k-means algorithm. The resulting profile helps analysts understand which kinds of messages dominate the traffic, and it supports downstream tasks such as formulating queries over schemaless data. But recomputing a full clustering from scratch every time the window slides is both computationally prohibitive and analytically jarring, since profiles computed independently in successive windows would be hard to compare over time.</p>
<p>DSC+ addresses these constraints with a two-phase design executed at every slide of the window. In the first phase, the schemas of incoming messages are pre-aggregated into compact data structures called schema features, or SFs. Rather than storing every individual schema, an SF records, for each attribute, the empirical probability that it appears within the group of schemas it summarizes. These structures are additive and subtractive, which means they can be merged, split, and trimmed as old messages expire from the window and new ones arrive, all without ever reconstructing the original messages.</p>
<p>To decide which schemas should be grouped together in the first place, the algorithm borrows a trick from document comparison: Min-Hash. Each schema is passed through a set of hash functions that produce a short identifier, and the mathematics of the technique guarantees that two schemas with high Jaccard similarity, meaning they share many attributes, are likely to receive the same identifier. The researchers adapted the method so that the hash indexes remain stable even as the attribute domain grows without bound, a property the original Min-Hash formulation lacks. Schemas sharing an identifier are folded into a single schema feature, producing a pane-level coreset, and these are in turn combined into a fixed-size window-level coreset that feeds the clustering stage.</p>
<p>The second phase is where DSC+ departs most sharply from prior work. Instead of rerunning k-means on each window, the algorithm incrementally updates the previous clustering result through a set of formalized rules that capture five distinct evolution phenomena. Sliding occurs when the schemas within a cluster gradually change, shifting the centroid. Fading-in happens when a new source starts emitting recognizably different schemas, and fading-out when an existing source goes silent. Splitting arises when one cluster grows so internally diverse that it should become two, and merging when two clusters drift so close together that they should become one.</p>
<p>Each rule is triggered by statistical signals rather than guesswork. The algorithm monitors the scattering of each cluster, the average distance between its members and its centroid, and the separation between pairs of clusters, the distance between their centroids. When a robust z-score, computed with the median absolute deviation, flags a cluster as an outlier in scattering, a split is executed by locally running k-means with two centers. When a pair of clusters overlaps enough and their separation becomes anomalously low, they are merged. New schema features are simply assigned to the nearest centroid, and empty clusters are deleted. If the quality of the updated profile, measured by the Simplified Silhouette, drops by more than what the worst possible sudden change in the data could plausibly explain, a full reclustering is triggered as a safety net.</p>
<p>The experimental evaluation, run on both synthetic datasets engineered to simulate each evolution phenomenon and on four real-world datasets, including application logs from a consulting company, sensor data from a precision agriculture project, and network traffic logs from the Suricata monitoring tool, shows the approach outperforming the closest state-of-the-art competitors. Against the baseline of rerunning the full clustering algorithm, OMRk++, on every window, DSC+ produced profiles that were actually slightly better in terms of silhouette, an effect the authors attribute to the pre-aggregation step filtering out noise from individual schemas. Compared with the streaming competitors CSCS and FEAC-S, DSC+ tracked changes in the number of clusters far more faithfully while executing faster in most configurations.</p>
<p>The efficiency gains stem from the fact that the expensive clustering operates on a coreset capped at a fixed maximum size, so execution time scales sublinearly with the volume of incoming data. In tests where the window length grew from ten thousand to one million schemas, DSC+ sustained increasing stream rates while a competitor&#8217;s throughput actually declined, because that competitor struggled to compress large windows into its fixed-size summary. The researchers also tuned the key parameters, including the coreset size and the number of hash functions, showing that accuracy plateaus quickly as the coreset grows, which means substantial data reduction comes at almost no cost in profile quality.</p>
<p>The authors are candid about limitations. The current implementation runs on a single server, and a distributed version exploiting the additivity of schema features is left for future work. The system also lacks explicit mechanisms for handling outlier schemas that resemble nothing seen before, although such cases did not arise in the real datasets tested. Still, the contribution marks a clear step forward: a clustering method that embraces a variable and unknown attribute domain, tracks the birth, death, splitting, and merging of clusters in real time, and does so within a bounded memory footprint. For the growing ranks of organizations drowning in heterogeneous IoT and log data, DSC+ offers a way to finally see, in real time, what their streams are actually made of.</p>
<p><strong>Subject of Research:</strong> Real-time schema profiling of heterogeneous data streams using dynamic stream clustering with variable-k k-means</p>
<p><strong>Article Title:</strong> Dynamic stream clustering for real-time schema profiling with DSC+</p>
<p><strong>Article References:</strong> Dynamic stream clustering for real-time schema profiling with DSC+. (n.d.). <a href="https://doi.org/10.1007/s10115-026-02889-w" rel="noopener noreferrer">https://doi.org/10.1007/s10115-026-02889-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10115-026-02889-w" rel="noopener noreferrer">10.1007/s10115-026-02889-w</a></p>
<p><strong>Keywords:</strong> schema profiling, data streams, stream clustering, k-means, coresets, Min-Hash, IoT, big data, sliding window, cluster evolution, schemaless data, real-time analytics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">241078</post-id>	</item>
	</channel>
</rss>
