<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>mucin gene diversity &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/mucin-gene-diversity/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 12:03:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>mucin gene diversity &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes</title>
		<link>https://scienmag.com/pangenome-graphs-and-a-cosine-trick-let-shallow-sequencing-crack-complex-genes/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 12:03:00 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[ancient DNA]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[complex gene region genotyping]]></category>
		<category><![CDATA[copy number variation]]></category>
		<category><![CDATA[COSIGT genotyping tool]]></category>
		<category><![CDATA[cosine similarity]]></category>
		<category><![CDATA[cosine similarity in genomics]]></category>
		<category><![CDATA[CYP2D6]]></category>
		<category><![CDATA[Genome Biology]]></category>
		<category><![CDATA[genome graph-based variant calling]]></category>
		<category><![CDATA[genomic haplotype reconstruction]]></category>
		<category><![CDATA[genotyping]]></category>
		<category><![CDATA[HLA locus variation]]></category>
		<category><![CDATA[HLA typing]]></category>
		<category><![CDATA[low-coverage sequencing]]></category>
		<category><![CDATA[mucin gene diversity]]></category>
		<category><![CDATA[multi-kilobase insertions and deletions]]></category>
		<category><![CDATA[pangenome graph analysis]]></category>
		<category><![CDATA[pangenome graphs]]></category>
		<category><![CDATA[pangenome reference frameworks]]></category>
		<category><![CDATA[population genomics]]></category>
		<category><![CDATA[shallow sequencing techniques]]></category>
		<category><![CDATA[structural variant detection in human genome]]></category>
		<category><![CDATA[structural variation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222518</guid>

					<description><![CDATA[A new pangenome graph tool called COSIGT uses cosine similarity to accurately genotype complex genomic loci from sequencing data as shallow as 1X coverage, including degraded ancient DNA.]]></description>
										<content:encoded><![CDATA[<p>Some of the most medically important stretches of the human genome are also the hardest to read. Genes such as the cytochrome P450 drug-metabolism family, the human leukocyte antigen (HLA) loci, and the mucin genes exist in a shifting landscape of multi-kilobase insertions, deletions, duplications, and copy-number variants. A single linear reference genome simply cannot represent the full range of allelic diversity at these regions, and standard variant-calling pipelines routinely fail to recover the true structural haplotypes hiding within them. Now, a team of researchers led by Davide Bolognini and Andrea Guarracino, working across Human Technopole in Milan, the University of Tennessee Health Science Center, and partner institutions, has introduced a tool designed to change that calculus. Called COSIGT, short for COsine SImilarity-based GenoTyper, the method is described in a brief report published in Genome Biology and is built to genotype complex loci accurately even from sequencing data so sparse that existing approaches collapse.</p>
<p>The core problem COSIGT addresses is one of depth. Recent methods such as Locityper have already shown that pangenome references, which store many haplotypes as paths through a graph rather than a single consensus sequence, can dramatically improve targeted genotyping at difficult loci when read data are abundant. Locityper aligns reads to locus-specific haplotypes and selects the best-fitting diploid pair by optimizing alignment accuracy, insert-size concordance, and coverage balance. But that alignment-based likelihood machinery has an Achilles heel: as sequencing coverage falls, the signal-to-noise ratio degrades and genotyping accuracy drops with it. That matters enormously for population-scale studies and biobank cohorts sequenced at variable and often shallow depth, and it matters even more for ancient DNA, where postmortem degradation and microbial contamination routinely push effective coverage below 2X.</p>
<p>COSIGT takes a fundamentally different mathematical route. The pipeline first constructs a local variation graph for the target locus from haplotype-resolved assemblies, rather than querying an entire genome-wide pangenome. This locus-specific design allows parameter tuning, rapid incorporation of new assemblies, and faster genotyping. Reads pre-aligned to the region are extracted, supplemented with unmapped reads rescued by a k-mer filtering tool called kfilt, and then mapped to the local graph. From the graph nodes that reads traverse, the pipeline builds a coverage vector for the sample, in which each element reflects length-normalized, multi-mapping-aware coverage at a node. Every haplotype in the graph is likewise represented as a coverage vector of node traversal counts. COSIGT then enumerates all possible diploid haplotype pairs, sums their vectors to create synthetic genotype profiles, and computes the cosine similarity between each synthetic profile and the observed sample vector. The pair with the highest similarity wins.</p>
<p>The elegance of the approach lies in what cosine similarity measures. Because the metric captures the orientation of a vector rather than its magnitude, it is inherently invariant to overall sequencing depth. A sample sequenced at 1X coverage produces the same relative coverage profile shape as one sequenced at 30X, merely scaled down, and cosine similarity is blind to that scaling. Likelihood-based methods, by contrast, depend on absolute read counts and lose power as counts dwindle. In benchmarks across 326 challenging medically relevant genes and 265 structurally variable regions using short-read data from the 1000 Genomes Project, with pangenome graphs built from assemblies of the Human Pangenome Reference Consortium and the Human Genome Structural Variation Consortium, both COSIGT and Locityper performed well at 5X and 30X coverage, with more than 93 percent high-quality calls. Locityper retained an edge at 30X, reaching 98.2 percent versus 93.9 percent for COSIGT at the medically relevant genes.</p>
<p>At low coverage, the picture reversed dramatically. At 1X, COSIGT delivered 93.4 percent of calls at mid-or-higher quality compared with 84.5 percent for Locityper, and at 2X the gap persisted at 95.8 percent versus 93.4 percent. The advantage became even starker on simulated ancient DNA. Across 48 medically relevant genes enriched for pharmacogenetic and immunogenetic content, including HLAs, CYPs, and mucins, COSIGT maintained roughly 94 percent mid-or-higher quality calls at 1X while Locityper fell to about 46 percent; at 2X the figures were roughly 95 percent versus 60 percent. COSIGT outperformed Locityper at all 48 genes and showed only marginal degradation relative to matched modern DNA at the same coverage. The simulations, generated with a purpose-built simulator called ancestralsim, incorporated realistic ancient DNA damage patterns, short fragment lengths, and contamination levels of 0 or 10 percent, making the result a demanding test.</p>
<p>Robustness to missing references was assessed through leave-all-out benchmarks, in which the true haplotypes of each sample were excluded from the pangenome graph. Even then, COSIGT remained near-optimal, achieving at least 87 percent of the best possible quality at most loci, with 87.6 percent of medically relevant genes and 90.8 percent of structural variant regions falling in the top quintile of achievable quality. An ancestry-mismatched test, in which Peruvian and Colombian samples were genotyped against a graph containing no admixed American assemblies, yielded 87.2 percent high-quality calls at the medically relevant genes, suggesting the method degrades gracefully when the reference panel does not perfectly match the population under study.</p>
<p>The team also demonstrated the tool at genuine population scale. Applying COSIGT to 1,085 whole-genome samples from the Italian Moli-sani cohort, sequenced at roughly 20X coverage, the researchers performed HLA typing across seven classical HLA genes and achieved a mean haplotype-level accuracy of 89.5 percent against types imputed with HLA*IMP:02 from microarray data, comparable to the 90.8 percent achieved by the dedicated HLA genotyping tool T1K. More broadly, the authors report having applied COSIGT to more than 6,000 modern and ancient human genomes, demonstrating population-scalable analysis of complex repeats and multi-copy genes. The pipeline is implemented in Snakemake with containerized deployment, parallelizes across regions and samples, and its per-sample genotyping steps are computationally light, with a median runtime of about 0.02 minutes and median memory use of 52 megabytes for the core genotyping step.</p>
<p>The authors are candid about limitations. Genotyping accuracy ultimately depends on the quality and completeness of the input pangenome: haplotypes absent from the reference panel cannot be called exactly, and novel structural variants will be assigned to the most similar available haplotype. The leave-all-out benchmarks quantify how much accuracy survives this constraint, but they also make clear that expanding pangenome references remains essential. COSIGT currently requires pre-aligned BAM or CRAM files, creating a dependency on reference-based alignment, though future work aims to accept raw FASTQ input directly and to implement haplotype subsampling to keep the method scalable as pangenomes grow. The framework also currently supports diploid genotyping, although the underlying mathematics generalizes to arbitrary ploidy by enumerating k-tuples of haplotypes.</p>
<p>The implications reach well beyond a single software release. Pangenome references are expanding rapidly in size, quality, and taxonomic breadth, and as they do, the bottleneck shifts from reference completeness to scalable genotyping. By making complex-locus genotyping reliable at 1 to 2X coverage, COSIGT opens the door to systematically mining the vast archives of low-coverage sequencing data already sitting in biobanks and legacy datasets, and to bringing archaeological and ancient genomes into population-level biomedical analyses from which their shallow depth previously excluded them. Larger cohorts become affordable at reduced sequencing cost, samples of heterogeneous quality can be combined in a single analysis, and the allelic diversity captured in pangenome references becomes practically accessible across the full spectrum of sequencing depths. The pipeline is compatible with long-read input as well, which could extend its utility to low-coverage long-read datasets too sparse for assembly. For a field that has long treated structurally complex loci as terra incognita in shallow data, the message is that the map, and the compass to read it, may finally have arrived together.</p>
<p><strong>Subject of Research:</strong> Pangenome graph-based genotyping of structurally complex genomic loci from low-coverage sequencing data</p>
<p><strong>Article Title:</strong> COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs</p>
<p><strong>Article References:</strong> Bolognini, D., Guarracino, A., Paleni, C., Dudley, T. S., Iacoviello, L., Raveane, A., Sudmant, P. H., Garrison, E., &amp; Soranzo, N. (2026). COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs. <em>Genome Biology, 27</em>(1), Article 286. <a href="https://doi.org/10.1186/s13059-026-04242-4" rel="noopener noreferrer">https://doi.org/10.1186/s13059-026-04242-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13059-026-04242-4" rel="noopener noreferrer">10.1186/s13059-026-04242-4</a></p>
<p><strong>Keywords:</strong> pangenome graphs, genotyping, low-coverage sequencing, ancient DNA, structural variation, copy-number variation, HLA typing, CYP2D6, cosine similarity, population genomics, Genome Biology, bioinformatics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222518</post-id>	</item>
	</channel>
</rss>
