Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes

October 1, 2026
in Biology
Juliet Wilcox
By Juliet Wilcox Scienmag Editorial Profile - Human Genetics
Reading Time: 5 mins read
0
Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes

Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes

Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Some of the most medically important stretches of the human genome are also the hardest to read. Genes such as the cytochrome P450 drug-metabolism family, the human leukocyte antigen (HLA) loci, and the mucin genes exist in a shifting landscape of multi-kilobase insertions, deletions, duplications, and copy-number variants. A single linear reference genome simply cannot represent the full range of allelic diversity at these regions, and standard variant-calling pipelines routinely fail to recover the true structural haplotypes hiding within them. Now, a team of researchers led by Davide Bolognini and Andrea Guarracino, working across Human Technopole in Milan, the University of Tennessee Health Science Center, and partner institutions, has introduced a tool designed to change that calculus. Called COSIGT, short for COsine SImilarity-based GenoTyper, the method is described in a brief report published in Genome Biology and is built to genotype complex loci accurately even from sequencing data so sparse that existing approaches collapse.

The core problem COSIGT addresses is one of depth. Recent methods such as Locityper have already shown that pangenome references, which store many haplotypes as paths through a graph rather than a single consensus sequence, can dramatically improve targeted genotyping at difficult loci when read data are abundant. Locityper aligns reads to locus-specific haplotypes and selects the best-fitting diploid pair by optimizing alignment accuracy, insert-size concordance, and coverage balance. But that alignment-based likelihood machinery has an Achilles heel: as sequencing coverage falls, the signal-to-noise ratio degrades and genotyping accuracy drops with it. That matters enormously for population-scale studies and biobank cohorts sequenced at variable and often shallow depth, and it matters even more for ancient DNA, where postmortem degradation and microbial contamination routinely push effective coverage below 2X.

COSIGT takes a fundamentally different mathematical route. The pipeline first constructs a local variation graph for the target locus from haplotype-resolved assemblies, rather than querying an entire genome-wide pangenome. This locus-specific design allows parameter tuning, rapid incorporation of new assemblies, and faster genotyping. Reads pre-aligned to the region are extracted, supplemented with unmapped reads rescued by a k-mer filtering tool called kfilt, and then mapped to the local graph. From the graph nodes that reads traverse, the pipeline builds a coverage vector for the sample, in which each element reflects length-normalized, multi-mapping-aware coverage at a node. Every haplotype in the graph is likewise represented as a coverage vector of node traversal counts. COSIGT then enumerates all possible diploid haplotype pairs, sums their vectors to create synthetic genotype profiles, and computes the cosine similarity between each synthetic profile and the observed sample vector. The pair with the highest similarity wins.

The elegance of the approach lies in what cosine similarity measures. Because the metric captures the orientation of a vector rather than its magnitude, it is inherently invariant to overall sequencing depth. A sample sequenced at 1X coverage produces the same relative coverage profile shape as one sequenced at 30X, merely scaled down, and cosine similarity is blind to that scaling. Likelihood-based methods, by contrast, depend on absolute read counts and lose power as counts dwindle. In benchmarks across 326 challenging medically relevant genes and 265 structurally variable regions using short-read data from the 1000 Genomes Project, with pangenome graphs built from assemblies of the Human Pangenome Reference Consortium and the Human Genome Structural Variation Consortium, both COSIGT and Locityper performed well at 5X and 30X coverage, with more than 93 percent high-quality calls. Locityper retained an edge at 30X, reaching 98.2 percent versus 93.9 percent for COSIGT at the medically relevant genes.

At low coverage, the picture reversed dramatically. At 1X, COSIGT delivered 93.4 percent of calls at mid-or-higher quality compared with 84.5 percent for Locityper, and at 2X the gap persisted at 95.8 percent versus 93.4 percent. The advantage became even starker on simulated ancient DNA. Across 48 medically relevant genes enriched for pharmacogenetic and immunogenetic content, including HLAs, CYPs, and mucins, COSIGT maintained roughly 94 percent mid-or-higher quality calls at 1X while Locityper fell to about 46 percent; at 2X the figures were roughly 95 percent versus 60 percent. COSIGT outperformed Locityper at all 48 genes and showed only marginal degradation relative to matched modern DNA at the same coverage. The simulations, generated with a purpose-built simulator called ancestralsim, incorporated realistic ancient DNA damage patterns, short fragment lengths, and contamination levels of 0 or 10 percent, making the result a demanding test.

Robustness to missing references was assessed through leave-all-out benchmarks, in which the true haplotypes of each sample were excluded from the pangenome graph. Even then, COSIGT remained near-optimal, achieving at least 87 percent of the best possible quality at most loci, with 87.6 percent of medically relevant genes and 90.8 percent of structural variant regions falling in the top quintile of achievable quality. An ancestry-mismatched test, in which Peruvian and Colombian samples were genotyped against a graph containing no admixed American assemblies, yielded 87.2 percent high-quality calls at the medically relevant genes, suggesting the method degrades gracefully when the reference panel does not perfectly match the population under study.

The team also demonstrated the tool at genuine population scale. Applying COSIGT to 1,085 whole-genome samples from the Italian Moli-sani cohort, sequenced at roughly 20X coverage, the researchers performed HLA typing across seven classical HLA genes and achieved a mean haplotype-level accuracy of 89.5 percent against types imputed with HLA*IMP:02 from microarray data, comparable to the 90.8 percent achieved by the dedicated HLA genotyping tool T1K. More broadly, the authors report having applied COSIGT to more than 6,000 modern and ancient human genomes, demonstrating population-scalable analysis of complex repeats and multi-copy genes. The pipeline is implemented in Snakemake with containerized deployment, parallelizes across regions and samples, and its per-sample genotyping steps are computationally light, with a median runtime of about 0.02 minutes and median memory use of 52 megabytes for the core genotyping step.

The authors are candid about limitations. Genotyping accuracy ultimately depends on the quality and completeness of the input pangenome: haplotypes absent from the reference panel cannot be called exactly, and novel structural variants will be assigned to the most similar available haplotype. The leave-all-out benchmarks quantify how much accuracy survives this constraint, but they also make clear that expanding pangenome references remains essential. COSIGT currently requires pre-aligned BAM or CRAM files, creating a dependency on reference-based alignment, though future work aims to accept raw FASTQ input directly and to implement haplotype subsampling to keep the method scalable as pangenomes grow. The framework also currently supports diploid genotyping, although the underlying mathematics generalizes to arbitrary ploidy by enumerating k-tuples of haplotypes.

The implications reach well beyond a single software release. Pangenome references are expanding rapidly in size, quality, and taxonomic breadth, and as they do, the bottleneck shifts from reference completeness to scalable genotyping. By making complex-locus genotyping reliable at 1 to 2X coverage, COSIGT opens the door to systematically mining the vast archives of low-coverage sequencing data already sitting in biobanks and legacy datasets, and to bringing archaeological and ancient genomes into population-level biomedical analyses from which their shallow depth previously excluded them. Larger cohorts become affordable at reduced sequencing cost, samples of heterogeneous quality can be combined in a single analysis, and the allelic diversity captured in pangenome references becomes practically accessible across the full spectrum of sequencing depths. The pipeline is compatible with long-read input as well, which could extend its utility to low-coverage long-read datasets too sparse for assembly. For a field that has long treated structurally complex loci as terra incognita in shallow data, the message is that the map, and the compass to read it, may finally have arrived together.

Subject of Research: Pangenome graph-based genotyping of structurally complex genomic loci from low-coverage sequencing data

Article Title: COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs

Article References: Bolognini, D., Guarracino, A., Paleni, C., Dudley, T. S., Iacoviello, L., Raveane, A., Sudmant, P. H., Garrison, E., & Soranzo, N. (2026). COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs. Genome Biology, 27(1), Article 286. https://doi.org/10.1186/s13059-026-04242-4

Image Credits: AI Generated

DOI: 10.1186/s13059-026-04242-4

Keywords: pangenome graphs, genotyping, low-coverage sequencing, ancient DNA, structural variation, copy-number variation, HLA typing, CYP2D6, cosine similarity, population genomics, Genome Biology, bioinformatics

Cite Scienmag News

Juliet Wilcox. (October 1, 2026). Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes. Scienmag. https://scienmag.com/pangenome-graphs-and-a-cosine-trick-let-shallow-sequencing-crack-complex-genes/

Juliet Wilcox. "Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes." Scienmag, 1 October 2026, https://scienmag.com/pangenome-graphs-and-a-cosine-trick-let-shallow-sequencing-crack-complex-genes/. Accessed 1 October 2026.

Juliet Wilcox. "Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes." Scienmag. October 1, 2026. https://scienmag.com/pangenome-graphs-and-a-cosine-trick-let-shallow-sequencing-crack-complex-genes/

Tags: ancient DNAbioinformaticscomplex gene region genotypingcopy number variationCOSIGT genotyping toolcosine similaritycosine similarity in genomicsCYP2D6Genome Biologygenome graph-based variant callinggenomic haplotype reconstructiongenotypingHLA locus variationHLA typinglow-coverage sequencingmucin gene diversitymulti-kilobase insertions and deletionspangenome graph analysispangenome graphspangenome reference frameworkspopulation genomicsshallow sequencing techniquesstructural variant detection in human genomestructural variation
Share26Tweet16
Previous Post

Skin Side Effects of Breast Cancer Drug Sacituzumab Govitecan Are Mostly Mild, Global Study Finds

Next Post

AI-Powered Review Maps 50 Years of Research on Milk’s Hidden Bioactives Carnitine and Betaine

Related Posts

Ancient Chinese Pigs Carry Genetic Secrets Behind Their Record-Breaking Teat Counts
Biology

Ancient Chinese Pigs Carry Genetic Secrets Behind Their Record-Breaking Teat Counts

October 1, 2026
Mothers Know Best: Plant Seed Coats Use Epigenetic Switches to Sculpt the Embryo
Biology

Mothers Know Best: Plant Seed Coats Use Epigenetic Switches to Sculpt the Embryo

October 1, 2026
Genetic Sleuthing Reveals Hidden Wildcat Lineages in Italy’s Apennine Mountains
Biology

Genetic Sleuthing Reveals Hidden Wildcat Lineages in Italy’s Apennine Mountains

October 1, 2026
BK Virus Kidney Infection Shows Broad Immune Disruption in Transplant Patients
Biology

BK Virus Kidney Infection Shows Broad Immune Disruption in Transplant Patients

October 1, 2026
Manure-Grown Bloodworms Could Replace Costly Imported Fish Feed for Catfish Fry
Biology

Manure-Grown Bloodworms Could Replace Costly Imported Fish Feed for Catfish Fry

October 1, 2026
Gut Bacteria Team Up to Destroy the Nerves That Keep the Bowel Moving
Biology

Gut Bacteria Team Up to Destroy the Nerves That Keep the Bowel Moving

October 1, 2026
Next Post
AI-Powered Review Maps 50 Years of Research on Milk’s Hidden Bioactives Carnitine and Betaine

AI-Powered Review Maps 50 Years of Research on Milk's Hidden Bioactives Carnitine and Betaine

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Ancient Chinese Pigs Carry Genetic Secrets Behind Their Record-Breaking Teat Counts
  • AI-Powered Review Maps 50 Years of Research on Milk’s Hidden Bioactives Carnitine and Betaine
  • Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes
  • Skin Side Effects of Breast Cancer Drug Sacituzumab Govitecan Are Mostly Mild, Global Study Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading