<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>censoring &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/censoring/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 12:23:54 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>censoring &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Hidden Microbial Zeros: New Statistical Model Sharpens the Hunt for Disease-Linked Gut Bacteria</title>
		<link>https://scienmag.com/hidden-microbial-zeros-new-statistical-model-sharpens-the-hunt-for-disease-linked-gut-bacteria/</link>
		
		<dc:creator><![CDATA[Morgan Morrow]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 12:23:54 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adenoma]]></category>
		<category><![CDATA[advanced biostatistics for microbiome studies]]></category>
		<category><![CDATA[biostatistics]]></category>
		<category><![CDATA[censoring]]></category>
		<category><![CDATA[Colorectal cancer]]></category>
		<category><![CDATA[differential abundance analysis]]></category>
		<category><![CDATA[differentiating structural zeros in microbiome datasets]]></category>
		<category><![CDATA[disease-linked gut bacteria detection]]></category>
		<category><![CDATA[gut microbiome and disease research]]></category>
		<category><![CDATA[gut microbiome statistical methods]]></category>
		<category><![CDATA[joint mixture Tobit model for microbiome analysis]]></category>
		<category><![CDATA[latent abundance]]></category>
		<category><![CDATA[microbial abundance data analysis]]></category>
		<category><![CDATA[microbiome]]></category>
		<category><![CDATA[microbiome sequencing]]></category>
		<category><![CDATA[microbiome zero interpretation]]></category>
		<category><![CDATA[PLOS Computational Biology]]></category>
		<category><![CDATA[sequencing data]]></category>
		<category><![CDATA[sequencing technology challenges in microbiome research]]></category>
		<category><![CDATA[statistical modeling]]></category>
		<category><![CDATA[Tobit model]]></category>
		<category><![CDATA[zero inflation]]></category>
		<category><![CDATA[zero inflation in microbiome data]]></category>
		<category><![CDATA[zero-inflated models in microbiology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=247578</guid>

					<description><![CDATA[A new joint mixture Tobit statistical model distinguishes true microbial absence from undetected low abundance, improving the detection of microbiome-disease links in zero-inflated sequencing data.]]></description>
										<content:encoded><![CDATA[<p>Every time scientists sequence the microbes living in a human gut, they confront a deceptively simple question: what does a zero actually mean? When a bacterial species shows up as absent in a stool sample, is it truly gone from that person&#8217;s body, or was it simply too rare for the sequencer to catch? This distinction, long glossed over in the rush to link microbiomes to disease, sits at the heart of a new statistical framework published in PLOS Computational Biology. A team of biostatisticians at Fudan University, led by Junrong Deng and Yue-Qing Hu, has developed a method called the joint mixture Tobit model, or joint mTobit, which promises to make the search for disease-associated microbes both more honest and more powerful.</p>
<p>The problem the method tackles is known as zero inflation, and it plagues virtually every microbiome study conducted to date. Sequencing technologies such as 16S rRNA gene sequencing and whole-metagenome shotgun sequencing produce abundance tables in which a large fraction of entries are zero. But those zeros are not all born equal. Some are structural zeros, meaning the microbe is genuinely absent from the sample, perhaps because the host was never exposed to that organism at all. Others are sampling zeros, where the microbe is present at low abundance but falls below the detection limit of the sequencing process, a phenomenon statisticians describe as left-censoring, borrowing a concept from engineering and economics where measurements below an instrument&#8217;s threshold are recorded as zero even though the true value is positive.</p>
<p>Most existing approaches to differential abundance analysis, the workhorse technique for identifying microbes that differ between healthy and diseased groups, handle these zeros crudely. One common strategy simply replaces zeros with small pseudo-counts, an artificial fix that can introduce bias and leaves researchers guessing about the right magnitude of the substitution. Another family of methods uses two-part or zero-inflated models that acknowledge two sources of zeros but stop short of modeling the true abundance lurking beneath the observation. Hurdle models go further in the wrong direction for some datasets, assuming every zero is structural. The consequence of these compromises is tangible: inflated false positive rates, lost statistical power, and unstable estimates that can make rare but genuinely disease-relevant taxa invisible.</p>
<p>The Fudan team&#8217;s insight was to treat sampling zeros the way economists have long treated censored income data, through the Tobit model, originally devised by James Tobin in 1958. In the classical Tobit framework, a continuous outcome is observed only when it exceeds a threshold; below it, the value is censored to zero. The researchers adapted this idea to microbial relative abundances, but with a crucial twist. Because microbiome data contain both structural and sampling zeros, they built a mixture Tobit model: a point-mass component captures the probability that a subject is not exposed to the microbe at all, while a censored lognormal component models the latent true abundance of exposed subjects whose microbes fall below the detection limit. The detection limit itself is estimated from the data as the normalized minimum of all positive counts across samples.</p>
<p>What makes the method genuinely novel, however, is its joint architecture. Rather than first estimating microbial abundance and then separately testing its relationship with disease, the joint mTobit framework links three things simultaneously within a single probabilistic model: the observed abundance, the latent true abundance, and the binary disease status. The disease outcome is connected to the latent true abundance through a logistic regression, meaning that associations are inferred at the level of the underlying biological signal rather than the imperfect measurement. When a positive abundance is observed, the latent value coincides with the observation. When a zero is observed, the model weighs two possibilities: a structural zero from non-exposure, or a censored positive abundance hidden below the detection threshold. This explicit decomposition allows confounding factors such as age, body mass index, and lipid levels to be adjusted for across all components of the model.</p>
<p>Fitting such a model is not trivial. The likelihood contains an integral of a non-linear function that has no closed-form solution, so the researchers resort to numerical integration and iterative optimization using the Nelder-Mead algorithm, with a carefully designed three-step initialization procedure to ensure convergence. Standard errors are computed with a sandwich covariance estimator, and inference on the key abundance-disease association parameter proceeds via a Wald test. The extra computational effort, the authors argue, is a small price for the inferential gains the framework delivers.</p>
<p>Those gains were demonstrated through extensive simulations. In parametric simulations spanning sample sizes of 100 and 200, non-exposure rates from 10 to 30 percent, and censoring rates from 10 to 30 percent, the joint mTobit method kept false positive rates closest to the nominal 0.05 level while achieving the highest and most stable statistical power. The contrast with a simpler joint Tobit model, which assumes all zeros arise from censoring, was striking: that model&#8217;s false positive rate frequently soared above 0.3, revealing how badly misclassifying structural zeros distorts inference. A second round of simulations built on real colorectal cancer and inflammatory bowel disease datasets using the SIMBA benchmarking framework confirmed the pattern, with joint mTobit achieving mean power of roughly 0.76 to 0.82 at a sample size of 200 while competitors lagged well behind, and the simple Tobit variant again showing false positive rates up to 0.183.</p>
<p>The real-world test came from a dataset of 46 colorectal cancer patients, 47 adenoma patients, and 61 healthy controls, with 249 microbial species retained after filtering. In the cancer-versus-control comparison, joint mTobit identified 19 differentially abundant taxa, including Peptostreptococcus stomatis, a well-established colorectal cancer-associated bacterium, and Veillonella dispar, previously reported to promote the disease, while health-associated species such as Propionibacterium freudenreichii and Eubacterium hallii appeared depleted. In the adenoma comparison, the method found 21 significant species, among them Bacteroides fragilis and Bacteroides faecis, both previously linked to premalignant lesions, and Enterococcus faecalis, whose inverse association aligns with experimental evidence that certain strains suppress polyp development in mice. Quantile-quantile plots of the p-values showed the method&#8217;s hallmark behavior: tight adherence to the null distribution where no signal exists, and a clean upward deviation where true associations emerge, in contrast to competitors that either produced excess small p-values or missed signals entirely.</p>
<p>The implications reach beyond the gut. The authors note that the combination of latent-variable modeling, explicit separation of structural and sampling zeros, and joint inference could be adapted to other sparse, detection-limited biological data, including single-cell and spatial transcriptomics and proteomics, where observed zeros likewise mix true absence with measurement failure. Limitations remain: the method depends on a lognormal assumption that holds for most but not all taxa, estimation can struggle in small samples with sparse non-exposed groups, and the numerical optimization is computationally heavier than simpler alternatives. The team suggests future extensions using negative binomial or gamma distributions, improved computational efficiency, and support for longitudinal microbiome data. For now, the message is clear: a zero in microbiome data is a question, not an answer, and treating it as one may finally reveal the rare microbes that matter most in human disease.</p>
<p><strong>Subject of Research:</strong> A censoring-based joint mixture Tobit statistical method for detecting microbiome-disease associations in zero-inflated sequencing data</p>
<p><strong>Article Title:</strong> A joint mixture Tobit method with latent microbial abundance improves detection of microbiome–disease associations in zero-inflated data</p>
<p><strong>Article References:</strong> Deng, J., Chen, D., Shen, S., Zhou, Y., Cui, H., Qiu, Y., Li, Q., &amp; Hu, Y.-Q. (2026). A joint mixture Tobit method with latent microbial abundance improves detection of microbiome–disease associations in zero-inflated data. <em>PLOS Computational Biology, 22</em>(10), e1014844. <a href="https://doi.org/10.1371/journal.pcbi.1014844" rel="noopener noreferrer">https://doi.org/10.1371/journal.pcbi.1014844</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pcbi.1014844" rel="noopener noreferrer">10.1371/journal.pcbi.1014844</a></p>
<p><strong>Keywords:</strong> microbiome, zero inflation, Tobit model, differential abundance analysis, latent abundance, colorectal cancer, adenoma, statistical modeling, censoring, sequencing data, biostatistics, PLOS Computational Biology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">247578</post-id>	</item>
	</channel>
</rss>
