<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>CLMSynth &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/clmsynth/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 11:57:45 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>CLMSynth &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Open-Source Tool Turns Cluster-Label Agreement Into a Tunable Dial for Benchmarking</title>
		<link>https://scienmag.com/new-open-source-tool-turns-cluster-label-agreement-into-a-tunable-dial-for-benchmarking/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 11:57:45 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adjusted Rand index]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[benchmarking clustering with adjustable label accuracy]]></category>
		<category><![CDATA[CLMSynth]]></category>
		<category><![CDATA[cluster-label alignment measurement]]></category>
		<category><![CDATA[cluster-label matching]]></category>
		<category><![CDATA[cluster-structure and label relationship]]></category>
		<category><![CDATA[clustering]]></category>
		<category><![CDATA[clustering algorithm benchmarking]]></category>
		<category><![CDATA[data generator for clustering experiments]]></category>
		<category><![CDATA[dataset label agreement control]]></category>
		<category><![CDATA[evaluating clustering algorithm robustness]]></category>
		<category><![CDATA[k-means]]></category>
		<category><![CDATA[label noise]]></category>
		<category><![CDATA[machine learning evaluation]]></category>
		<category><![CDATA[Matthews Correlation Coefficient]]></category>
		<category><![CDATA[open-source Python clustering tools]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[reproducible synthetic dataset creation]]></category>
		<category><![CDATA[synthetic benchmark data for machine learning]]></category>
		<category><![CDATA[synthetic data]]></category>
		<category><![CDATA[synthetic data generation for clustering]]></category>
		<category><![CDATA[tunable cluster-structure simulation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222498</guid>

					<description><![CDATA[Researchers have developed CLMSynth, an open-source Python package that generates synthetic labels with precisely controlled agreement to known cluster structures, letting scientists stress-test clustering benchmarks and validation metrics under configurable real-world conditions.]]></description>
										<content:encoded><![CDATA[<p>Every benchmark in machine learning rests on a quiet assumption that rarely gets examined: that the labels attached to a dataset actually line up with the structure hidden inside the data. When researchers test a clustering algorithm, they typically measure how well the groups the algorithm discovers agree with pre-assigned class labels, treating that agreement as a fixed property of the dataset. A new open-source Python package called CLMSynth, described in the journal SoftwareX by Arootin Gharibian and Miloš Kudělka of VŠB–Technical University of Ostrava, flips that assumption on its head. Instead of accepting whatever agreement exists between clusters and labels, the tool lets researchers dial in exactly how much agreement they want, generating synthetic label assignments that hit a requested target with reproducible precision.</p>
<p>The motivation comes from a gap in how synthetic benchmark data are built. Existing generators for clustering benchmarks, such as the multidimensional dataset generator MDCGen and the Framework for Benchmarking Clustering Algorithms, concentrate on shaping the geometry of clusters themselves, and they generally assume that cluster structure and class labels are equivalent. Label-noise generators, meanwhile, work in the opposite direction, injecting corruption into labels that already exist. CLMSynth addresses what the authors call the unsupervised counterpart of this problem: it takes a dataset whose ground-truth clusters are already known and independently controls the relationship between newly generated labels and that existing structure, without touching the clusters at all.</p>
<p>At its core, the software is a label allocation engine. Given a dataset with clusters, the user specifies how many synthetic labels to create, how balanced or skewed their distribution should be, and how closely each label should match the cluster assignment. The generator then solves for the label assignment that best matches a user-specified cluster–label matching measure within the configuration&#8217;s constraints. If the requested target is feasible, the delivered value matches it within a user-declared tolerance; if it is not, the solver reports the closest achievable value or refuses the configuration outright. This inversion, turning a validation metric from a fixed output into a flexible target, is the tool&#8217;s central technical contribution.</p>
<p>The mathematics behind the solver is elegant in its economy. For a binary Matthews correlation coefficient, once the cluster size, label size, and total number of points are fixed, the only free variable in the entire confusion matrix is the number of true positives. Expanding the MCC formula reveals that the denominator is constant and that the required true-positive count is an affine function of the target MCC, which means an exact label count per cluster can be computed by rounding to the nearest valid integer. For multi-label, multi-cluster cases, where a single-pair MCC becomes hard to interpret, the tool uses a Hungarian-matched implementation of Gorodkin&#8217;s RK statistic, and the adjusted Rand index is available as an alternative target that scales better on very large datasets because it compares point pairs rather than searching over label-cluster correspondences.</p>
<p>Configuration happens through six generative conditions that together emulate the messiness of real-world data. Users define the label space and its balance, choosing between deterministic skews and distributions sampled at random. Four matching modes then govern the baseline relationship: a perfect bijection between clusters and labels, a random permutation that serves as a lower bound on any validation measure, a single mode that places one label inside one cluster only, and a custom mode that permits many-to-one, partial, and overlapping assignments through a rule matrix. Because real data are rarely tidy, syndromes with overlapping and unbalanced symptoms being a canonical example, these modes deliberately span the space between the upper and lower bounds of what any clustering algorithm could score.</p>
<p>Noise is treated as an experimental variable rather than an afterthought. Points left unclaimed by the matching rules can receive structured competing noise, where a chosen share of one cluster&#8217;s unclaimed points is assigned a specific competing label, or they can be distributed through spillover rules that return each label to its target count, concentrate leftovers on a designated label, or assign them uniformly at random. The distinction matters because structured and unstructured noise are not interchangeable, particularly when labels sit at the boundary of adjacent clusters, such as an adjacent diagnostic category or a threshold applied to a biomarker that varies continuously within a cluster. A further configuration controls centroid proximity, allowing labels to favor either the core or the boundary of a cluster, independently of the overall matching level.</p>
<p>The authors demonstrate the tool&#8217;s expressive range on a single dataset of four clusters, generating fifteen distinct cluster–label matching outcomes that illustrate every configuration option, from perfect alignment through dominant-minority label balances, core versus boundary placement, and competing noise at a cluster&#8217;s edge, down to an unsolvable configuration where the solver gracefully returns the nearest feasible value. An interactive command-line wizard guides users through building YAML configuration files, with troubleshooting designed to catch incompatible inputs before they produce unintended results. The package, released under the MIT license and installable with a single pip command, runs on any platform supporting Python 3.12 through 3.14 and depends only on standard scientific libraries including numpy, pandas, scipy, and scikit-learn.</p>
<p>One demonstration in particular shows why the tool could change how benchmarks are read. Working with a benchmark dataset of fifteen clusters, the researchers allocated one label per cluster but varied the amount of competing noise inside a central supercluster of eight tightly packed groups. As the requested agreement level rose, the k-means algorithm&#8217;s preferred number of clusters flipped from eight to fifteen: below a certain threshold of label agreement, the benchmark reported the supercluster, and above it, the algorithm resolved the fine structure. In other words, whether a clustering method appears to find the coarse or the fine partition depends not on the algorithm but on how strongly the projected labels demand that resolution. The switch occurred between requested correlation values of 0.45 and 0.50, a range that a fixed benchmark dataset could never have exposed.</p>
<p>The implications reach beyond clustering research. External validation of clusters through label recovery has faced growing criticism, especially in domains like medicine where class labels may only partially coincide with underlying structure, and global metrics frequently penalize algorithms that isolate one meaningful cluster while treating the rest of the data as noise. By making the degree, direction, and spatial placement of labels into configurable inputs, CLMSynth lets researchers stress-test their metrics across all possible cases rather than being limited to whatever real-world data happen to contain. Because label placement is metric-invariant, two datasets can carry identical cluster–label matching values while differing completely in where labels sit within clusters, enabling controlled comparisons that real data cannot provide.</p>
<p>The tool has honest limitations. The ceiling of the multi-class MCC solution cannot be precomputed from a configuration alone, so targets set above that ceiling produce identical outcomes accompanied by a warning, and the MCC target is searched over a coarse grid to keep computation tractable, which can yield slightly suboptimal solutions. Still, the authors position CLMSynth as a complement rather than a replacement: where label-noise tools like SYNLABEL vary labels against the feature space, CLMSynth varies labels against a fixed cluster partition, and together the two approaches cover controlled label synthesis for both supervised and unsupervised benchmarking. For a field increasingly aware that its benchmarks encode hidden assumptions, a tool that makes those assumptions explicitly adjustable is a quietly powerful addition to the toolbox.</p>
<p><strong>Subject of Research:</strong> A synthetic data generator for controlling cluster–label matching in clustering benchmark evaluation</p>
<p><strong>Article Title:</strong> A synthetic data generator for cluster and label matching evaluation</p>
<p><strong>Article References:</strong> Gharibian, A., &amp; Kudělka, M. (2026). A synthetic data generator for cluster and label matching evaluation. <em>SoftwareX, 36</em>, Article 103077. <a href="https://doi.org/10.1016/j.softx.2026.103077" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103077</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103077" rel="noopener noreferrer">10.1016/j.softx.2026.103077</a></p>
<p><strong>Keywords:</strong> CLMSynth, synthetic data, clustering, benchmarking, label noise, cluster-label matching, adjusted Rand index, Matthews correlation coefficient, k-means, open-source software, machine learning evaluation, Python</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222498</post-id>	</item>
	</channel>
</rss>
