<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>sex inference &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/sex-inference/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 03:37:17 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>sex inference &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Tool Reads Hidden Host DNA in Microbiome Data to Predict Biological Sex</title>
		<link>https://scienmag.com/new-tool-reads-hidden-host-dna-in-microbiome-data-to-predict-biological-sex/</link>
		
		<dc:creator><![CDATA[Morgan Morrow]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 03:37:17 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[animal species]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[chromosomal sex]]></category>
		<category><![CDATA[host DNA]]></category>
		<category><![CDATA[host sex influence on microbiome composition]]></category>
		<category><![CDATA[human microbiome project]]></category>
		<category><![CDATA[limiting the depth of biological insights.]]></category>
		<category><![CDATA[low host biomass]]></category>
		<category><![CDATA[metagenomics]]></category>
		<category><![CDATA[microbiome]]></category>
		<category><![CDATA[multinomial likelihood]]></category>
		<category><![CDATA[open-source software]]></category>
		<category><![CDATA[quality control]]></category>
		<category><![CDATA[sex inference]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=236686</guid>

					<description><![CDATA[A new bioinformatic tool called SCiMS predicts host chromosomal sex from the small amount of host DNA present in shotgun metagenomic data, even in low host-biomass samples and across multiple animal species.]]></description>
										<content:encoded><![CDATA[<p>Every shotgun metagenomic sequencing run carries a hidden passenger. When researchers sequence the microbes living in or on a host, a small fraction of the reads that come off the sequencer do not belong to bacteria, viruses, or fungi at all. They are fragments of the host&#8217;s own genome, shed into the sample alongside the microbial community. For years, most microbiome studies have treated this host DNA as contamination to be filtered out and discarded. A new open-source tool called SCiMS, short for Sex Calling in Metagenomic Sequences, turns that discarded signal into something genuinely useful: a reliable prediction of the host&#8217;s chromosomal sex, even when almost no host DNA is present in the sample.</p>
<p>The work, published in the journal Microbiome by Hanh N. Tran, Kobie J. Kirven, and Emily R. Davenport of Pennsylvania State University, addresses a surprisingly persistent gap in microbiome research. Host sex is one of the most important determinants of microbial community structure across many species, shaped by hormonal profiles, physiology, and sex-stratified behaviors. Yet sex metadata is frequently missing from microbiome datasets, including studies of animal-associated samples. Without knowing whether a sample came from a male or a female, researchers cannot properly account for a variable that may be quietly driving much of the variation they observe in microbial composition.</p>
<p>The technical challenge is that existing genomic sex prediction tools were built for a different problem. Tools designed to infer chromosomal sex from sequencing data typically rely on fixed coverage thresholds calibrated specifically for human X and Y chromosomes. They compare the depth of sequencing coverage on the X chromosome against the depth on autosomes, using the logic that an XX individual should show roughly twice the X-chromosome coverage of an XY individual relative to the rest of the genome. The problem is that these approaches demand a relatively high number of host reads to make the coverage ratios statistically meaningful. In low host-biomass samples, such as stool, where host DNA is scarce to begin with, the coverage estimates become too noisy for the thresholds to work. The tools also assume a human-style XY sex-determination system, which makes them useless for organisms with different chromosomal architectures.</p>
<p>SCiMS takes a fundamentally different statistical approach. Rather than relying on rigid coverage thresholds, the tool computes a multinomial likelihood from the observed read counts mapped to each chromosome under each candidate sex karyotype. In plain terms, it asks a probabilistic question: given the number of sequencing reads observed mapping to each chromosome, what is the likelihood that those reads were drawn from a genome with one copy of the X chromosome and one copy of the Y chromosome, versus two copies of the X chromosome? The expected read distribution under each hypothesis is derived directly from the chromosome lengths and the ploidy of each chromosome under the candidate karyotype. Because the expectation is computed from first principles rather than from a fixed human-calibrated threshold, the method generalizes naturally to any organism with a heterogametic sex-determination system, meaning any species in which males and females carry different sex chromosome complements.</p>
<p>This generality is what sets SCiMS apart. A multinomial likelihood framework scales gracefully with data: when host reads are plentiful, the likelihood strongly favors the correct karyotype, and when host reads are scarce, the tool still extracts whatever signal exists in the ratio of reads mapping to sex chromosomes versus autosomes. The authors report chromosomal sex calls with an explicit statistical basis, allowing downstream users to understand not just what the call is but how confident the underlying model is in that call. The tool is designed to slot into standard metagenomic processing pipelines, running on the host-derived reads that most workflows already identify and remove during the decontamination step.</p>
<p>To validate the approach, the team benchmarked SCiMS against existing sex prediction tools across three distinct testing regimes. First, they used simulated metagenomic data, where the true host sex is known by construction and the number of host reads can be controlled precisely. This allowed them to map out exactly how performance degrades as host coverage drops, and to quantify the false positive and false negative rates of each method at every depth. Second, they applied the tools to real human metagenomic samples spanning multiple body sites, drawn from resources such as the Human Microbiome Project, where documented sex metadata exists to serve as ground truth. Third, and most tellingly for the tool&#8217;s generality, they tested it on metagenomic samples from seven different animal species, organisms whose sex chromosomes and genome structures differ substantially from the human reference.</p>
<p>The results were consistent across all three regimes: SCiMS matched or outperformed the existing tools, with its most noticeable advantage appearing precisely where current methods fail, in the low host read conditions that characterize stool samples and other low host-biomass specimen types. In simulations, the multinomial approach continued to make accurate calls at host read depths where threshold-based methods collapsed into coin flips. On real human samples from multiple body sites, SCiMS recovered known sex labels with high accuracy, and on the seven animal species it demonstrated the cross-species portability that human-calibrated tools structurally cannot achieve. The benchmarking also tracked true positives, false positives, and false negatives systematically, giving the field a transparent picture of where each method can be trusted.</p>
<p>The implications for microbiome science reach beyond simple convenience. Missing metadata is a quiet epidemic in publicly available sequencing data. Samples deposited in repositories years ago, sometimes by labs that have since moved on, often lack critical host variables, and sex is among the most commonly absent. Because sex shapes microbial communities through hormones, immune physiology, and behavior, its absence can confound meta-analyses that pool thousands of samples from many studies. A tool that can recover chromosomal sex directly from the sequencing data itself, retroactively and at scale, functions as a form of quality control for the entire discipline. Researchers can reanalyze archived metagenomic datasets, fill in the missing variable, and either confirm that their findings hold after sex-stratified analysis or discover that a previously unexplained pattern was a sex effect all along.</p>
<p>The cross-species capability opens its own set of doors. Wildlife microbiome studies, livestock research, and comparative genomics all generate shotgun metagenomic data from animals whose sex is unrecorded or ambiguously documented, particularly in field-collected samples. Because SCiMS derives its expectations from chromosome lengths and ploidy under each candidate karyotype, it can be applied to any heterogametic species with a reference genome, without needing species-specific calibration. That means a single tool can now serve a laboratory working on human gut samples, mouse models, and wild-caught animal specimens alike, a level of generality that fixed-threshold approaches were never designed to offer.</p>
<p>SCiMS is freely available on GitHub through the Davenport lab, and the underlying article is open access, so the barrier to adoption is essentially zero. The work was supported by National Institutes of Health grants, including R35GM146980 to Davenport and a training fellowship supporting Tran. As microbiome datasets continue to grow in size and ecological scope, the ability to recover missing host metadata from the data itself represents a small but consequential shift: the contaminating reads that pipelines once threw away are now a resource, and one of the most important variables in host-associated microbiology can be reconstructed from samples whose donors may never have been asked a single question about it.</p>
<p><strong>Subject of Research:</strong> A computational method for inferring host chromosomal sex from host-derived reads in shotgun metagenomic microbiome data</p>
<p><strong>Article Title:</strong> SCiMS: Sex Calling in Metagenomic Sequences</p>
<p><strong>Article References:</strong> Tran, H. N., Kirven, K. J., &amp; Davenport, E. R. (2026). SCiMS: Sex Calling in Metagenomic Sequences. <em>Microbiome</em>. <a href="https://doi.org/10.1186/s40168-026-02547-x" rel="noopener noreferrer">https://doi.org/10.1186/s40168-026-02547-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40168-026-02547-x" rel="noopener noreferrer">10.1186/s40168-026-02547-x</a></p>
<p><strong>Keywords:</strong> metagenomics, microbiome, sex inference, bioinformatics, chromosomal sex, host DNA, multinomial likelihood, quality control, Human Microbiome Project, open-source software, animal species, low host biomass</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">236686</post-id>	</item>
	</channel>
</rss>
