<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>consensus clustering &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/consensus-clustering/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 26 Sep 2026 22:32:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>consensus clustering &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Ensemble Clustering Offers Psychologists a Fix for Unstable Subgroups</title>
		<link>https://scienmag.com/ensemble-clustering-offers-psychologists-a-fix-for-unstable-subgroups/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 22:32:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[Behavior Research Methods]]></category>
		<category><![CDATA[biological markers and symptom profiles clustering]]></category>
		<category><![CDATA[cluster analysis]]></category>
		<category><![CDATA[code-based clustering solutions]]></category>
		<category><![CDATA[consensus clustering]]></category>
		<category><![CDATA[ensemble clustering]]></category>
		<category><![CDATA[ensemble clustering techniques]]></category>
		<category><![CDATA[improving clustering reliability in psychology]]></category>
		<category><![CDATA[latent class analysis]]></category>
		<category><![CDATA[limitations of traditional clustering methods]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning and cluster stability]]></category>
		<category><![CDATA[open-access tutorials for clustering]]></category>
		<category><![CDATA[psychiatry]]></category>
		<category><![CDATA[psychology]]></category>
		<category><![CDATA[psychology and psychiatry data analysis]]></category>
		<category><![CDATA[R programming]]></category>
		<category><![CDATA[replication crisis]]></category>
		<category><![CDATA[reproducibility in psychological research]]></category>
		<category><![CDATA[stability]]></category>
		<category><![CDATA[stability and robustness in clustering]]></category>
		<category><![CDATA[unstable subgroups in psychological research]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216761</guid>

					<description><![CDATA[A new tutorial in Behavior Research Methods shows that ensemble clustering can dramatically stabilise subgroup identification in psychology and psychiatry, where single-run methods collapse under even tiny data perturbations.]]></description>
										<content:encoded><![CDATA[<p>Cluster analysis has quietly shaped psychology and psychiatry for more than eight decades, promising to sort messy human data into meaningful subgroups without the distortions of human bias. The technique dates back to 1939, when R.C. Tryon introduced the term to identify clusters of variables, and it has since grown into a cornerstone of research on symptom profiles, disease classification, treatment response, and biological markers. More than 60,000 clustering articles have been published to date, with over 5,000 appearing in 2024 alone. Yet a new open-access tutorial in Behavior Research Methods, led by Caroline X. Gao of the University of Melbourne and Orygen together with an international team, argues that the field&#8217;s most widely used clustering practices are quietly failing, and it offers a practical, code-heavy roadmap for a more reliable alternative: ensemble clustering.</p>
<p>The core problem, the authors explain, is that traditional single-run clustering methods lack stability, robustness, and generalisability. Stability refers to whether results remain consistent across repeated runs; robustness describes resilience to noise; generalisability concerns whether findings hold in different datasets. All three are chronically weak in conventional approaches. This is not simply a matter of sloppy implementation. Kleinberg&#8217;s impossibility theorem shows that no clustering method can simultaneously satisfy all desirable properties, and the optimisation problem underlying clustering has no exact solution. Algorithms therefore rely on assumptions and trade-offs, and their stochastic optimisers mean that minor perturbations of the data, different starting points, or a different choice of algorithm can produce substantially different subgroups from the same dataset.</p>
<p>The downstream consequence is a replication crisis for subgroup research. Identified subtypes often vary dramatically across studies, and when broadly similar patterns do emerge, they frequently reflect arbitrary cut-offs along a single severity dimension rather than genuine heterogeneity. In mental health research, the stakes are particularly high: despite decades of effort, current evidence does not support reproducible, biologically valid, and clinically agreed-upon neuroimaging subtypes for mental illnesses. By contrast, fields such as oncology have made real translational progress with tumour subtyping, largely by adopting ensemble clustering methods that combine solutions from different data distributions, algorithms, and optimisation runs into a single consensus.</p>
<p>Ensemble clustering works much like ensemble methods in supervised machine learning, where predictions from many models are pooled to reduce the weaknesses of any single model. The clustering version follows a two-stage framework. The first stage generates many base clusterings, or BCs, by running different algorithms, random seeds, parameter settings, or resampled subsets of observations and variables. The second stage aggregates these partitions into a consensus, defined not as complete agreement but as the stable grouping pattern that recurs across solutions. The tutorial stresses that good BCs should be both diverse and high quality, a difficult balance because aggressive sub-sampling increases diversity while degrading individual partition quality, and optimising quality tends to drive all BCs toward the same answer.</p>
<p>The consensus stage is where the methodological richness lies. The simplest approach is majority voting, but because clustering algorithms return arbitrary cluster labels, BCs must first be relabelled and aligned, and voting fails entirely when BCs are highly diverse. More sophisticated techniques summarise the BCs before running a second-stage clustering model. Cluster labels can be treated as categorical variables and clustered with finite mixture models or k-modes. They can be binary-encoded into a cluster association matrix for k-means-based consensus clustering, or converted into bipartite graphs for graph partitioning methods such as the hypergraph-partitioning algorithm. Alternatively, a co-association matrix records how often each pair of observations is grouped together across BCs, serving directly as a distance matrix for hierarchical clustering in evidence accumulation clustering, or transformed into weighted graphs for algorithms such as the cluster-based similarity partitioning algorithm. Newer options include the link-based clustering ensemble, the meta-clustering algorithm, adaptive clustering ensemble, and even deep learning approaches.</p>
<p>Choosing the right components depends on the data. K-means is the most common BC algorithm because it is computationally cheap, though it assumes spherical clusters and is sensitive to outliers; model-based methods such as latent class analysis may better suit psychological data where subgroups plausibly arise from latent generative processes. The co-association matrix remains the most popular consensus tool because heatmaps of it visually display cluster stability and help select the number of clusters. Ensemble size matters too: performance generally improves with more BCs, but benchmarking studies suggest quality plateaus between roughly 50 and 100 base clusterings. A sampling proportion of about 80 percent is typically recommended, capturing most data characteristics while preserving variation among BCs.</p>
<p>A major practical payoff of the ensemble approach is a principled way to choose the number of clusters. Instead of relying on internal validity indices such as Silhouette, Calinski-Harabasz, Dunn, or Gap statistics, which often disagree with one another, researchers can inspect the consensus matrix. When the chosen number of clusters matches the true structure, the matrix displays clear square-shaped blocks along the antidiagonal, indicating pairs of observations that are consistently grouped together or apart. This visual pattern can be quantified through the cumulative distribution function of consensus indices, the delta area plot, the cophenetic correlation coefficient, and the proportion of ambiguous clustering score, though the authors caution that PAC carries an inherent bias toward larger cluster counts because it ignores null reference distributions.</p>
<p>To make the method concrete, the tutorial walks through a full R implementation using packages such as diceR, ConsensusClusterPlus, and clue, with step-by-step code provided in supplementary materials. Using a simulated dataset of 600 observations with 12 variables generated from three multivariate normal clusters, the authors demonstrate the dice function, combining k-means, Gaussian mixture models, and Ward&#8217;s hierarchical clustering with the link-based consensus function across 100 BCs. They also cover pre-registration of analysis plans, data assessment, dimensionality reduction via PCA or alternatives such as UMAP and NMF, imputation for missing data introduced by resampling, trimming and reweighting of poor-quality BCs, permutation tests to confirm solutions did not arise by chance, and cross-validation frameworks for tuning hyper-parameters in the absence of ground truth.</p>
<p>The most striking evidence comes from simulations built on three real-world psychological datasets covering quality of life, alexithymic traits, and cannabis expectancies. The team perturbed each dataset by randomly removing between 0.25 and 20 percent of the data, refilling gaps with multiple imputation, and repeated the process 100 times per condition. Even tiny perturbations of 0.25 percent produced substantially different results under single-run methods; hierarchical clustering in one case performed barely better than random assignment. Ensemble clustering consistently showed the slowest decay in stability, measured by the Rand index, and much narrower variability across perturbed dataset pairs, particularly for higher-dimensional data.</p>
<p>The authors are careful about limits. Ensemble methods inherit the assumptions of their base algorithms, so stable solutions can still miss non-convex or arbitrary cluster shapes; stability and internal validity do not guarantee external validity, and even valid, generalisable solutions may lack practical utility if they merely partition severity. Ordinal Likert-scale items and latent constructs can distort distance matrices, and longitudinal data require specialised Bayesian ensemble pipelines. Computational cost is real, though manageable with binary association matrices and high-performance computing. Still, the tutorial&#8217;s message is unambiguous: minor perturbations of psychological data can upend single-run clustering, ensemble methods tame that instability, and with mature R tooling now available, the authors advocate widespread adoption to finally turn heterogeneous psychological data into subgroups researchers can trust.</p>
<p><strong>Subject of Research:</strong> Ensemble clustering methods for improving the stability and reproducibility of subgroup identification in psychological and psychiatric research</p>
<p><strong>Article Title:</strong> Ensemble clustering: A practical tutorial</p>
<p><strong>Article References:</strong> Gao, C. X., Wang, S., Zhu, Y., Ziou, M., Teo, S. M., Smith, C. L., Chiu, D., Talhouk, A., Wang, M., Yu, W., Cotton, S. M., &amp; Dwyer, D. (2026). Ensemble clustering: A practical tutorial. <em>Behavior Research Methods, 58</em>(10), Article 286. <a href="https://doi.org/10.3758/s13428-026-03158-y" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03158-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03158-y" rel="noopener noreferrer">10.3758/s13428-026-03158-y</a></p>
<p><strong>Keywords:</strong> ensemble clustering, cluster analysis, consensus clustering, psychiatry, psychology, machine learning, unsupervised learning, R programming, replication crisis, latent class analysis, stability, Behavior Research Methods</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216761</post-id>	</item>
		<item>
		<title>Tumors in the Same Dog Are Molecularly Worlds Apart, Landmark Study Shows</title>
		<link>https://scienmag.com/tumors-in-the-same-dog-are-molecularly-worlds-apart-landmark-study-shows/</link>
		
		<dc:creator><![CDATA[William Thompson]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 01:01:24 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[breast cancer subtypes]]></category>
		<category><![CDATA[canine mammary tumor molecular classification]]></category>
		<category><![CDATA[canine mammary tumors]]></category>
		<category><![CDATA[canine tumor heterogeneity]]></category>
		<category><![CDATA[Comparative Oncology]]></category>
		<category><![CDATA[comparative oncology studies in dogs]]></category>
		<category><![CDATA[consensus clustering]]></category>
		<category><![CDATA[gene co-expression network]]></category>
		<category><![CDATA[high]]></category>
		<category><![CDATA[impact of tumor heterogeneity on diagnosis and treatment]]></category>
		<category><![CDATA[molecular]]></category>
		<category><![CDATA[molecular differences in tumors from the same dog]]></category>
		<category><![CDATA[molecular subgroups of canine mammary tumors]]></category>
		<category><![CDATA[RNA sequencing]]></category>
		<category><![CDATA[RNA sequencing in veterinary oncology]]></category>
		<category><![CDATA[synchronous tumor development in dogs]]></category>
		<category><![CDATA[synchronous tumors]]></category>
		<category><![CDATA[transcriptomic analysis of canine cancers]]></category>
		<category><![CDATA[Transcriptomics]]></category>
		<category><![CDATA[tumor biology in canine mammary tumors]]></category>
		<category><![CDATA[tumor genetic diversity within individual dogs]]></category>
		<category><![CDATA[tumor heterogeneity]]></category>
		<category><![CDATA[veterinary oncology tumor research]]></category>
		<category><![CDATA[veterinary pathology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200312</guid>

					<description><![CDATA[A new transcriptomic study of 179 canine mammary tumors shows that synchronous tumors within the same dog are molecularly heterogeneous and often fall into distinct expression-based clusters resembling human breast cancer subtypes.]]></description>
										<content:encoded><![CDATA[<p>When a dog develops mammary tumors, veterinarians almost never find just one. Roughly half to sixty percent of dogs diagnosed with canine mammary tumors carry multiple, physically distinct growths at the same time, a phenomenon known as synchronous tumor development. For decades, the assumption has been that tumors arising in the same animal, sharing the same genetic background, hormones, and environment, would at least resemble one another biologically. A new transcriptomic study from Norwegian researchers dismantles that assumption with striking clarity, showing that even tumors sitting side by side within a single dog can be molecular strangers.</p>
<p>The study, published in the journal Veterinary Oncology, was led by Ingrid Marie Moberg and colleagues at the Norwegian University of Life Sciences, Oslo University Hospital, and the University of Oslo. The team analyzed RNA sequencing data from 179 canine mammary tumors, drawn from a cohort of naturally occurring tumors removed from companion dogs during routine surgery. Their goal was twofold: to define molecular subgroups of canine mammary tumors without reference to histology, and then to ask whether tumors from the same dog share those molecular identities. The answer to the second question, in most cases, was no.</p>
<p>To classify the tumors, the researchers first reduced their high-dimensional gene expression data using principal component analysis, then applied unsupervised consensus clustering, an algorithm that groups samples based purely on expression similarity and stability across repeated resampling. This revealed five robust transcriptomic clusters. Crucially, the clusters did not map cleanly onto histological diagnosis. Each cluster contained a mixture of benign and malignant tumors, confirming that the microscopic appearance of these tumors tells only part of the story of their underlying biology.</p>
<p>Gene set enrichment analysis then revealed that the five clusters bear a remarkable resemblance to molecular subtypes long recognized in human breast cancer. Three of the clusters showed enrichment of hormone-related gene programs, including estrogen response, and were found to harbor the transcription factor GATA3, a classic marker of luminal breast cancer in humans. One cluster combined cell-cycle proliferation with immune and interferon signaling, evoking the aggressive basal-like or triple-negative phenotype, while a fifth cluster was defined by epithelial-mesenchymal transition, angiogenesis, and inflammatory modules, reminiscent of the claudin-low subtype. The parallel with human breast cancer taxonomy is not merely cosmetic; it strengthens the case for dogs as a comparative model in which the biology of mammary cancer can be studied in a spontaneously arising disease.</p>
<p>Beyond clustering, the team constructed a gene co-expression network using the hCoCena framework, identifying ten modules of genes that are expressed together and that encode distinct biological processes. One module captured hormone signaling, two captured immune and interferon responses, three reflected metabolism and proliferation through glycolysis, PI3K-AKT-mTOR signaling, oxidative phosphorylation, and MYC targets, and others traced epithelial differentiation through WNT signaling, epithelial-mesenchymal transition, cell-cycle regulation through E2F targets and the G2M checkpoint, and tumor microenvironment features such as angiogenesis. Transcription factor enrichment within the modules pointed to GATA3 as a regulator of the hormonal program and E2F1 as a driver of the proliferative module, providing candidate master switches behind the observed phenotypes.</p>
<p>The heart of the study, however, lies in its analysis of synchronous tumors. The researchers focused on 45 dogs that each carried exactly two tumors, yielding 90 paired samples classified as benign-benign, malignant-benign, or malignant-malignant. When they compared cluster assignments within each pair, concordance was low. Only about 45 percent of tumor pairs landed in the same transcriptomic cluster overall, with malignant-malignant pairs showing the highest agreement at 56 percent, benign-benign pairs at 44 percent, and mixed malignant-benign pairs at just 36 percent. In other words, the majority of dogs carried two tumors with fundamentally different molecular identities.</p>
<p>To quantify this divergence at the level of individual genes, the team calculated intraclass correlation coefficients for every gene across the paired samples. Genes were scored as low, moderate, or high in correlation between a dog&#8217;s two tumors. The results were unambiguous: between roughly 76 and 90 percent of genes showed low correlation across all diagnostic categories, and fewer than one percent of genes were highly correlated within any category. This pattern of widespread discordance held regardless of whether both tumors were benign, both malignant, or one of each, indicating that molecular independence between synchronous tumors is the norm rather than the exception.</p>
<p>A small set of exceptions proved informative. The analysis identified a handful of genes, including OMD and EN1, whose expression is strongly correlated within certain categories of synchronous tumor pairs and which have been reported as prognostic markers in human cancers. The authors suggest that these genes may point to shared disease processes or protective mechanisms against malignant transformation, and that they deserve further investigation, potentially at the DNA level, to uncover any genetic factors underlying their coordinated behavior. Meanwhile, examination of module-level variation showed that programs linked to cell-cycle activity, hormone signaling, and immune responses fluctuated most between paired tumors, while modules tied to epithelial differentiation and metabolism remained comparatively stable within individuals.</p>
<p>The clinical implications of this work reach in two directions at once. For veterinary medicine, the findings suggest that each tumor in a multi-tumor patient should be evaluated as a biologically independent lesion rather than assumed to be representative of its neighbors. Because molecular subtypes in human breast cancer drive dramatically different treatment decisions, from endocrine therapy for hormone receptor-positive disease to chemotherapy and PARP inhibitors for triple-negative tumors, a biology-driven approach could eventually refine the limited therapeutic options currently available for canine patients. The authors note, for instance, that identifying hormone-positive subtypes could inform whether ovariohysterectomy offers real benefit at the time of tumor removal, a decision veterinarians currently make without molecular guidance.</p>
<p>For comparative oncology, the study reinforces the value of the canine model in a way that human cohorts cannot easily replicate. Synchronous bilateral breast cancer occurs in only around one percent of human patients, whereas synchronous mammary tumors affect the majority of affected dogs. Studying these paired tumors within the same genetic background eliminates many host-specific confounders, offering a natural experiment in tumor evolution and inter-individual heterogeneity. The authors acknowledge limitations inherent to bulk RNA sequencing, which averages expression across all cells in a tissue and may reflect tissue composition as much as tumor biology, and they caution that a single RNA sample may not represent an entire tumor. Even so, their conclusion stands firm: histopathology alone does not capture the molecular reality of canine mammary tumors, and expression-based profiling offers a more faithful map of the biological terrain that clinicians and researchers alike will need to navigate.</p>
<p><strong>Subject of Research:</strong> Transcriptomic heterogeneity of synchronous canine mammary tumors and their molecular resemblance to human breast cancer subtypes</p>
<p><strong>Article Title:</strong> High molecular heterogeneity in synchronous canine mammary tumors detected by transcriptomic analysis</p>
<p><strong>Article References:</strong> Moberg, I. M., Murphy, S. L., Hansen, N., Borge, K. S., Gunnes, G., Sørlie, T., Lingaas, F., Bergholtz, H., &amp; Solbakken, M. H. (2026). High molecular heterogeneity in synchronous canine mammary tumors detected by transcriptomic analysis. <em>Veterinary Oncology, 3</em>(1), Article 11. <a href="https://doi.org/10.1186/s44356-026-00066-3" rel="noopener noreferrer">https://doi.org/10.1186/s44356-026-00066-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s44356-026-00066-3" rel="noopener noreferrer">10.1186/s44356-026-00066-3</a></p>
<p><strong>Keywords:</strong> canine mammary tumors, transcriptomics, RNA sequencing, tumor heterogeneity, synchronous tumors, breast cancer subtypes, gene co-expression network, consensus clustering, comparative oncology, veterinary pathology, High, molecular</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200312</post-id>	</item>
	</channel>
</rss>
