<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>expectation maximization &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/expectation-maximization/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 00:45:44 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>expectation maximization &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Poisson-Based Algorithm Finds Hidden Groups in Count Data While Ignoring the Noise</title>
		<link>https://scienmag.com/new-poisson-based-algorithm-finds-hidden-groups-in-count-data-while-ignoring-the-noise/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 00:45:44 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[3CPO clustering method]]></category>
		<category><![CDATA[clustering]]></category>
		<category><![CDATA[column selection]]></category>
		<category><![CDATA[count data]]></category>
		<category><![CDATA[count data analysis]]></category>
		<category><![CDATA[count data vs. traditional tabular data]]></category>
		<category><![CDATA[data mining]]></category>
		<category><![CDATA[document word frequency clustering]]></category>
		<category><![CDATA[domain-aware clustering algorithms]]></category>
		<category><![CDATA[expectation maximization]]></category>
		<category><![CDATA[gene expression]]></category>
		<category><![CDATA[hidden group detection in count matrices]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[minimum description length]]></category>
		<category><![CDATA[noise reduction in count data]]></category>
		<category><![CDATA[open-access data mining research]]></category>
		<category><![CDATA[outlier detection]]></category>
		<category><![CDATA[Poisson distribution]]></category>
		<category><![CDATA[Poisson-based clustering algorithm]]></category>
		<category><![CDATA[RNA molecule count analysis]]></category>
		<category><![CDATA[statistical modeling for count data]]></category>
		<category><![CDATA[subspace clustering]]></category>
		<category><![CDATA[text clustering]]></category>
		<category><![CDATA[unsupervised learning for non-negative integers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211694</guid>

					<description><![CDATA[Researchers have developed 3CPO, a Poisson-based clustering algorithm that groups count data accurately while automatically identifying which columns of a matrix are relevant, outperforming standard methods across text, biology, and economics data sets.]]></description>
										<content:encoded><![CDATA[<p>Count data are everywhere. Whenever a matrix records how many times something happened — how often a person chose one option over another, how many words appear in a document, how many RNA molecules a cell produces — the result is a table of non-negative integers with properties that ordinary clustering tools struggle to respect. A new open-access study in Data Mining and Knowledge Discovery by Collin Leiber, Kai Puolamäki and Heikki Mannila, researchers at Aalto University and the University of Helsinki, introduces an algorithm called 3CPO that clusters such data using a statistically sound Poisson model while simultaneously deciding which columns of the matrix actually matter for the grouping.</p>
<p>The core problem the authors tackle is that count matrices behave differently from typical tabular data. In a standard spreadsheet, each column describes a different trait — a birth year, a height, a gender — and the columns do not even share a common data type. In a count matrix, by contrast, every entry counts occurrences of the same kind of event, and both rows and columns share a common domain. Generic clustering algorithms, which usually assume continuous features or require pre-processing such as normalization, often produce unreliable results on raw counts. Earlier work, notably a widely cited 2010 paper by O&#8217;Hara and Kotze, showed that log-transforming count data is frequently unsuitable, which limits the usefulness of traditional pipelines that depend on such transformations.</p>
<p>3CPO — short for Clustering and Column selection of Count Data using a Poisson-based Optimization — builds on a simple but powerful modeling idea: the expected value of each entry in the matrix can be written as a product of a row-specific factor and a column-specific factor. The row factor captures the overall scale of a row, for example the length of a document, while the column factors act like cluster centroids, describing the proportions that characterize each group of rows. This formulation, which echoes earlier Poisson clustering methods such as PoissonL and PoissonC, assumes that every row is scaled by its size and that every cluster follows its own characteristic column proportions.</p>
<p>Those assumptions are strict, and real data routinely violate them. The authors illustrate the point with word counts in letters: a greeting phrase appears roughly once regardless of letter length, contradicting the scaling assumption, while common words like &#8216;the&#8217; scale with document length but carry little information about the topic. To handle this, 3CPO divides the columns of the matrix into three disjoint subsets. Columns in the first subset, C1, are genuinely relevant for clustering and follow the full Poisson model with cluster-specific parameters. Columns in C0 scale with the rows but behave similarly across all clusters, so they are modeled with a single shared parameter. Columns in C− are pure noise, modeled by a column-specific average that ignores both row scale and cluster membership.</p>
<p>The algorithm then solves an optimization problem: find the parameters, the clustering of rows, and the column partition that minimize a log-loss derived from the Poisson likelihood. It does so with an Expectation Maximization-style iterative procedure that alternates between updating the expected values, reassigning rows to clusters using a Poisson-based score, and reassigning columns to the three subsets. A penalty term motivated by the Minimum Description Length principle keeps the set of relevant columns small — intuitively, describing K cluster-specific values for a column costs about (K−1) times the logarithm of the column sum in bits, so a column only earns its place in C1 if the improvement in likelihood justifies that cost. Because each step never increases the loss and there are finitely many possible clusterings, the procedure is guaranteed to converge to a local optimum, with a worst-case runtime that grows linearly in the number of rows, columns, clusters, and iterations.</p>
<p>The experiments are extensive. The authors compared 3CPO against Poisson-based baselines, Spherical k-Means, k-Means combined with several normalizations (z-scores, min-max scaling, relative frequencies, and the revealed comparative advantage measure used in economics), and a suite of co-clustering algorithms including CROINFO, CoclustMod, CoclustSpecMod, ELBM, SELBM, and TauCC. They evaluated one synthetic and eleven real-world data sets spanning gene expression, single-cell RNA sequencing, text corpora such as BBCSports, BBCNews, Reuters21578 and 20Newsgroups, handwritten digits, and economic wholesale data, using Unsupervised Clustering Accuracy, Normalized Mutual Information, and the Adjusted Rand Index.</p>
<p>The results are striking. 3CPO was the top performer in eight of twelve comparisons against traditional clustering algorithms, beating all competitors by more than 38 percent on the synthetic data and by more than 6 percent on WebKB, while remaining within one standard deviation of the best method in most cases where it did not win. Against co-clustering methods it ranked among the top three on every data set. On text data, it outperformed k-Means and Spherical k-Means combined with TF-IDF and BM25 weighting in four out of five scenarios, despite working directly on raw word counts. Because it operates on interpretable counts rather than opaque embeddings, domain experts can inspect the selected columns and see exactly which terms drive each cluster — the analysis of BBCSports and BBCNews showed words like &#8216;party&#8217; and &#8216;govern&#8217; characterizing a politics cluster and &#8216;athlete&#8217; and &#8216;olymp&#8217; characterizing an athletics cluster, while stop words were correctly relegated to the uninformative subsets.</p>
<p>The column selection itself proved remarkably effective. On high-dimensional data sets such as BBCSports, Reuters, 20Newsgroups and a gene expression data set, 3CPO kept only a fraction of the original columns — roughly 32 percent for the gene expression data and about 40 percent for BBCSports — while still outperforming methods that used everything. Histogram analyses showed the algorithm was not simply discarding sparse, zero-heavy columns; the distribution of zeros was similar across all three column subsets, indicating that 3CPO responds to genuine structural patterns in the counts. Robustness experiments reinforced the point: when noise columns with uniformly distributed values were added to the synthetic data, 3CPO was the only algorithm whose clustering quality remained perfect, and it tolerated noise values up to around 32 before degrading, far beyond the limits of every competitor.</p>
<p>The authors also built in an optional outlier detection mechanism. Rows are compared against a data set-wide background model, and any row that fits no cluster better than the background is flagged as an outlier rather than forced into one. This improved clustering scores on nearly all data sets, and on the gene expression data set 3CPO achieved a perfect clustering result while identifying only about fourteen outliers. An additional MDL-based penalty even allows the algorithm to estimate the number of clusters itself: it recovered the correct number for the synthetic and BBCSports data sets and was off by only one for the gene expression data, outperforming standard heuristics such as elbow detection, silhouette scores, and BIC-based estimates in the high-dimensional text settings.</p>
<p>The authors are candid about limitations. The Poisson model ties the variance to the mean, so data with strong overdispersion or zero inflation might be better served by a negative binomial formulation; the penalty terms involve heuristic choices; and the number of clusters must either be supplied or estimated with the proposed heuristic, which struggled on some tabular data sets. Still, the overall message is compelling: by taking the statistical nature of counts seriously and letting the data itself reveal which features matter, 3CPO delivers clusters that are both more accurate and far easier to interpret. The code is publicly available on GitHub, and the authors suggest that refining outlier handling and integrating cluster-number estimation more tightly are promising directions for future work. For anyone analyzing contingency tables, text counts, or sequencing data, the study makes a strong case that the essentials are best found by modeling the counts as they are — not by transforming them into something they are not.</p>
<p><strong>Subject of Research:</strong> A Poisson-based subspace clustering algorithm for count data with integrated column selection</p>
<p><strong>Article Title:</strong> Poisson subspace clustering: focusing on the essentials in count data</p>
<p><strong>Article References:</strong> Leiber, C., Puolamäki, K., &amp; Mannila, H. (2026). Poisson subspace clustering: focusing on the essentials in count data. <em>Data Mining and Knowledge Discovery, 40</em>(6), Article 98. <a href="https://doi.org/10.1007/s10618-026-01230-x" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01230-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01230-x" rel="noopener noreferrer">10.1007/s10618-026-01230-x</a></p>
<p><strong>Keywords:</strong> count data, Poisson distribution, clustering, subspace clustering, column selection, expectation maximization, minimum description length, data mining, machine learning, gene expression, text clustering, outlier detection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211694</post-id>	</item>
		<item>
		<title>Nested Dirichlet Mixture Models Bring New Precision to Clustering Count Data</title>
		<link>https://scienmag.com/nested-dirichlet-mixture-models-bring-new-precision-to-clustering-count-data/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 01:00:29 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced mixture models for data analysis]]></category>
		<category><![CDATA[clustering]]></category>
		<category><![CDATA[clustering of frequency-based data]]></category>
		<category><![CDATA[count data]]></category>
		<category><![CDATA[count data clustering]]></category>
		<category><![CDATA[Dirichlet distribution applications]]></category>
		<category><![CDATA[energy disaggregation]]></category>
		<category><![CDATA[expectation maximization]]></category>
		<category><![CDATA[finite mixture models]]></category>
		<category><![CDATA[Fisher information matrix]]></category>
		<category><![CDATA[high-dimensional count vector analysis]]></category>
		<category><![CDATA[image clustering]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[minimum message length]]></category>
		<category><![CDATA[multinomial nested Dirichlet mixture model]]></category>
		<category><![CDATA[nested Dirichlet distribution]]></category>
		<category><![CDATA[Nested Dirichlet mixture models]]></category>
		<category><![CDATA[overdispersion in count data]]></category>
		<category><![CDATA[probabilistic modeling of sparse data]]></category>
		<category><![CDATA[statistical methods for energy consumption data]]></category>
		<category><![CDATA[text and image feature clustering]]></category>
		<category><![CDATA[text clustering]]></category>
		<category><![CDATA[unsupervised machine learning for count data]]></category>
		<category><![CDATA[video categorization]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204828</guid>

					<description><![CDATA[Researchers at Concordia University have introduced a nested Dirichlet finite mixture model with an exact Fisher information matrix and an exponential approximation that improves clustering of sparse, overdispersed count data across text, image, video, and energy applications.]]></description>
										<content:encoded><![CDATA[<p>Count data are everywhere in modern science and industry. Every time a document is represented by how many times each word appears, an image is described by the frequency of visual features, or a household&#8217;s electricity meter logs how much energy each appliance consumes, the result is a vector of nonnegative integers. These high-dimensional count vectors are notoriously awkward to analyze: they are sparse, bursting with sudden spikes, and plagued by a statistical phenomenon known as overdispersion, in which the observed variance far exceeds what standard models such as the multinomial distribution would predict. Clustering such data—grouping similar observations without any labels—has therefore remained one of the more stubborn challenges in unsupervised machine learning.</p>
<p>A new study published in Data Mining and Knowledge Discovery by Fares Alkhawaja, Manar Amayri, and Nizar Bouguila of the Concordia Institute for Information Systems Engineering at Concordia University in Montreal tackles this problem head-on. The researchers introduce a family of finite mixture models built on the nested Dirichlet distribution, a flexible probability distribution that serves as the kernel of their new multinomial nested Dirichlet mixture model, abbreviated MNDM. The work generalizes the widely used Dirichlet compound multinomial approach, which has long been a workhorse for clustering count data but suffers from structural constraints that limit how faithfully it can represent real datasets.</p>
<p>To understand why the nested Dirichlet matters, it helps to recall the lineage of Dirichlet-based models. The classic Dirichlet distribution, a generalization of the beta distribution to multiple categories, dates back decades and has been a cornerstone of Bayesian statistics. In 1969, Connor and Mosimann proposed a generalized Dirichlet distribution that relaxes some of the strong independence assumptions of the original. Later, the nested Dirichlet distribution emerged as a further extension, allowing components of a proportion vector to be grouped into nested subsets, which mirrors the natural hierarchical structure of many real datasets—for example, documents organized into topics and subtopics, or a home&#8217;s electrical load organized into circuits and appliances. By adopting this nested structure as the kernel of a finite mixture, the Montreal team created a model that can capture richer covariance patterns among count variables than its predecessors.</p>
<p>Finite mixture models work by assuming that the data are drawn from a weighted combination of several component distributions, each representing a cluster. Fitting such models is typically done with the expectation maximization algorithm, a decades-old iterative procedure that alternates between estimating cluster memberships and updating the model parameters. The authors go a step further by employing the deterministic annealing variant of expectation maximization, known as DAEM. Deterministic annealing gradually sharpens the assignment of data points to clusters, starting soft and becoming increasingly confident, which helps the algorithm avoid poor local optima—a common failure mode in mixture estimation. The parameters are initialized using the method of moments, a classical estimation technique that matches theoretical moments of the distribution to their empirical counterparts.</p>
<p>One of the paper&#8217;s most technically significant contributions concerns the Fisher information matrix, a fundamental object in statistics that measures how much information the data carry about the unknown parameters. Model selection criteria such as the minimum message length, or MML, principle—rooted in the idea that the best model is the one that compresses the data most efficiently—require evaluating this matrix. In previous work on Dirichlet-type mixtures, researchers often relied on approximations of the Fisher information, which could introduce inaccuracies into the model selection process. Alkhawaja and colleagues derive the exact Fisher information matrix for their nested mixture model, eliminating the errors introduced by approximation. The exact computation, however, comes at a price: convergence takes considerably longer.</p>
<p>To recover computational efficiency without sacrificing too much accuracy, the authors also develop an exponential approximation of the MNDM. By casting the model into the exponential family of distributions, they unlock a suite of mathematical conveniences that the exponential family affords, including simplified parameter estimation and closed-form manipulations. The trade-off is elegant and practical: the exact Fisher information version delivers the best clustering performance but converges slowly, while the exponentially approximated version achieves relatively good performance at a fraction of the computational cost. This gives practitioners a dial to turn depending on whether accuracy or speed is the priority for their application.</p>
<p>The team validated their framework on four demanding real-world applications: text clustering, image clustering, video categorization, and energy disaggregation. The text experiments draw on well-known benchmarks including large-scale sentiment datasets, the Reuters news corpus, and web page collections, where documents are encoded as word-count vectors. Image clustering experiments use standard datasets of handwritten digits, scene images, textures, and faces, with features derived from established computer vision descriptors. Video categorization spans action recognition benchmarks, where spatio-temporal interest points extracted from footage are counted to characterize human motion. Finally, energy disaggregation—the task of inferring how much electricity individual appliances use from a single whole-house meter, central to nonintrusive load monitoring—was tested on public datasets of residential and commercial power consumption.</p>
<p>Across these applications, the pattern of results was consistent: the exact Fisher information variant of the nested mixture model outperformed both the approximate version and established baselines built on the Dirichlet compound multinomial and generalized Dirichlet mixtures, but at a higher convergence cost. The exponential approximation closed much of the performance gap while remaining practical for large datasets. The MML-based model selection criterion reliably identified sensible numbers of clusters, sidestepping the need for arbitrary choices about how many groups the data contain. The findings suggest that the nested Dirichlet family, with its hierarchical flexibility, offers a genuinely better statistical language for describing sparse, bursty count data than the distributions that dominated the previous two decades.</p>
<p>The implications extend well beyond the benchmark datasets tested in the paper. Count-valued data arise in genomics, where microbial communities are profiled by counting taxonomic sequences; in linguistics, where word frequencies follow famously bursty patterns; in economics, where purchase counts and insurance claims exhibit overdispersion; and in the growing field of smart-grid analytics, where appliance-level energy inference underpins demand-response programs. A clustering framework that models these phenomena natively, rather than forcing count data into Gaussian-shaped assumptions, could improve everything from recommendation systems to medical diagnosis pipelines. The research was supported by the Natural Sciences and Engineering Research Council of Canada, the Fonds de Recherche du Québec, and a start-up grant from Concordia University, and the authors report no conflicts of interest.</p>
<p>Looking ahead, the Concordia group&#8217;s program of work—spanning earlier unsupervised nested Dirichlet mixtures, simultaneous clustering and feature selection, and hierarchical count data clustering—points toward increasingly sophisticated generative models for discrete data. The exact Fisher information derivation technique developed here may also transfer to other members of the Dirichlet family, giving statisticians sharper tools for principled model selection across the unsupervised learning landscape. As machine learning systems are asked to make sense of ever-larger torrents of discrete events, from clickstreams to sensor readings, models that respect the true statistical texture of counts may prove essential. This study demonstrates that with the right distributional foundation, even the messiest count data can reveal coherent, actionable structure.</p>
<p><strong>Subject of Research:</strong> Clustering of high-dimensional count data using nested multinomial Dirichlet finite mixture models with exact Fisher information and exponential approximation</p>
<p><strong>Article Title:</strong> Clustering of count data using nested multinomial Dirichlet finite mixture model and its extensions</p>
<p><strong>Article References:</strong> Alkhawaja, F., Amayri, M., &amp; Bouguila, N. (2026). Clustering of count data using nested multinomial Dirichlet finite mixture model and its extensions. <em>Data Mining and Knowledge Discovery, 40</em>(6), Article 93. <a href="https://doi.org/10.1007/s10618-026-01250-7" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01250-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01250-7" rel="noopener noreferrer">10.1007/s10618-026-01250-7</a></p>
<p><strong>Keywords:</strong> count data, clustering, nested Dirichlet distribution, finite mixture models, Fisher information matrix, minimum message length, expectation maximization, text clustering, image clustering, video categorization, energy disaggregation, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204828</post-id>	</item>
	</channel>
</rss>
