<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>genotyping &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/genotyping/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 12:03:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>genotyping &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Pangenome Graphs and a Cosine Trick Let Shallow Sequencing Crack Complex Genes</title>
		<link>https://scienmag.com/pangenome-graphs-and-a-cosine-trick-let-shallow-sequencing-crack-complex-genes/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 12:03:00 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[ancient DNA]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[complex gene region genotyping]]></category>
		<category><![CDATA[copy number variation]]></category>
		<category><![CDATA[COSIGT genotyping tool]]></category>
		<category><![CDATA[cosine similarity]]></category>
		<category><![CDATA[cosine similarity in genomics]]></category>
		<category><![CDATA[CYP2D6]]></category>
		<category><![CDATA[Genome Biology]]></category>
		<category><![CDATA[genome graph-based variant calling]]></category>
		<category><![CDATA[genomic haplotype reconstruction]]></category>
		<category><![CDATA[genotyping]]></category>
		<category><![CDATA[HLA locus variation]]></category>
		<category><![CDATA[HLA typing]]></category>
		<category><![CDATA[low-coverage sequencing]]></category>
		<category><![CDATA[mucin gene diversity]]></category>
		<category><![CDATA[multi-kilobase insertions and deletions]]></category>
		<category><![CDATA[pangenome graph analysis]]></category>
		<category><![CDATA[pangenome graphs]]></category>
		<category><![CDATA[pangenome reference frameworks]]></category>
		<category><![CDATA[population genomics]]></category>
		<category><![CDATA[shallow sequencing techniques]]></category>
		<category><![CDATA[structural variant detection in human genome]]></category>
		<category><![CDATA[structural variation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222518</guid>

					<description><![CDATA[A new pangenome graph tool called COSIGT uses cosine similarity to accurately genotype complex genomic loci from sequencing data as shallow as 1X coverage, including degraded ancient DNA.]]></description>
										<content:encoded><![CDATA[<p>Some of the most medically important stretches of the human genome are also the hardest to read. Genes such as the cytochrome P450 drug-metabolism family, the human leukocyte antigen (HLA) loci, and the mucin genes exist in a shifting landscape of multi-kilobase insertions, deletions, duplications, and copy-number variants. A single linear reference genome simply cannot represent the full range of allelic diversity at these regions, and standard variant-calling pipelines routinely fail to recover the true structural haplotypes hiding within them. Now, a team of researchers led by Davide Bolognini and Andrea Guarracino, working across Human Technopole in Milan, the University of Tennessee Health Science Center, and partner institutions, has introduced a tool designed to change that calculus. Called COSIGT, short for COsine SImilarity-based GenoTyper, the method is described in a brief report published in Genome Biology and is built to genotype complex loci accurately even from sequencing data so sparse that existing approaches collapse.</p>
<p>The core problem COSIGT addresses is one of depth. Recent methods such as Locityper have already shown that pangenome references, which store many haplotypes as paths through a graph rather than a single consensus sequence, can dramatically improve targeted genotyping at difficult loci when read data are abundant. Locityper aligns reads to locus-specific haplotypes and selects the best-fitting diploid pair by optimizing alignment accuracy, insert-size concordance, and coverage balance. But that alignment-based likelihood machinery has an Achilles heel: as sequencing coverage falls, the signal-to-noise ratio degrades and genotyping accuracy drops with it. That matters enormously for population-scale studies and biobank cohorts sequenced at variable and often shallow depth, and it matters even more for ancient DNA, where postmortem degradation and microbial contamination routinely push effective coverage below 2X.</p>
<p>COSIGT takes a fundamentally different mathematical route. The pipeline first constructs a local variation graph for the target locus from haplotype-resolved assemblies, rather than querying an entire genome-wide pangenome. This locus-specific design allows parameter tuning, rapid incorporation of new assemblies, and faster genotyping. Reads pre-aligned to the region are extracted, supplemented with unmapped reads rescued by a k-mer filtering tool called kfilt, and then mapped to the local graph. From the graph nodes that reads traverse, the pipeline builds a coverage vector for the sample, in which each element reflects length-normalized, multi-mapping-aware coverage at a node. Every haplotype in the graph is likewise represented as a coverage vector of node traversal counts. COSIGT then enumerates all possible diploid haplotype pairs, sums their vectors to create synthetic genotype profiles, and computes the cosine similarity between each synthetic profile and the observed sample vector. The pair with the highest similarity wins.</p>
<p>The elegance of the approach lies in what cosine similarity measures. Because the metric captures the orientation of a vector rather than its magnitude, it is inherently invariant to overall sequencing depth. A sample sequenced at 1X coverage produces the same relative coverage profile shape as one sequenced at 30X, merely scaled down, and cosine similarity is blind to that scaling. Likelihood-based methods, by contrast, depend on absolute read counts and lose power as counts dwindle. In benchmarks across 326 challenging medically relevant genes and 265 structurally variable regions using short-read data from the 1000 Genomes Project, with pangenome graphs built from assemblies of the Human Pangenome Reference Consortium and the Human Genome Structural Variation Consortium, both COSIGT and Locityper performed well at 5X and 30X coverage, with more than 93 percent high-quality calls. Locityper retained an edge at 30X, reaching 98.2 percent versus 93.9 percent for COSIGT at the medically relevant genes.</p>
<p>At low coverage, the picture reversed dramatically. At 1X, COSIGT delivered 93.4 percent of calls at mid-or-higher quality compared with 84.5 percent for Locityper, and at 2X the gap persisted at 95.8 percent versus 93.4 percent. The advantage became even starker on simulated ancient DNA. Across 48 medically relevant genes enriched for pharmacogenetic and immunogenetic content, including HLAs, CYPs, and mucins, COSIGT maintained roughly 94 percent mid-or-higher quality calls at 1X while Locityper fell to about 46 percent; at 2X the figures were roughly 95 percent versus 60 percent. COSIGT outperformed Locityper at all 48 genes and showed only marginal degradation relative to matched modern DNA at the same coverage. The simulations, generated with a purpose-built simulator called ancestralsim, incorporated realistic ancient DNA damage patterns, short fragment lengths, and contamination levels of 0 or 10 percent, making the result a demanding test.</p>
<p>Robustness to missing references was assessed through leave-all-out benchmarks, in which the true haplotypes of each sample were excluded from the pangenome graph. Even then, COSIGT remained near-optimal, achieving at least 87 percent of the best possible quality at most loci, with 87.6 percent of medically relevant genes and 90.8 percent of structural variant regions falling in the top quintile of achievable quality. An ancestry-mismatched test, in which Peruvian and Colombian samples were genotyped against a graph containing no admixed American assemblies, yielded 87.2 percent high-quality calls at the medically relevant genes, suggesting the method degrades gracefully when the reference panel does not perfectly match the population under study.</p>
<p>The team also demonstrated the tool at genuine population scale. Applying COSIGT to 1,085 whole-genome samples from the Italian Moli-sani cohort, sequenced at roughly 20X coverage, the researchers performed HLA typing across seven classical HLA genes and achieved a mean haplotype-level accuracy of 89.5 percent against types imputed with HLA*IMP:02 from microarray data, comparable to the 90.8 percent achieved by the dedicated HLA genotyping tool T1K. More broadly, the authors report having applied COSIGT to more than 6,000 modern and ancient human genomes, demonstrating population-scalable analysis of complex repeats and multi-copy genes. The pipeline is implemented in Snakemake with containerized deployment, parallelizes across regions and samples, and its per-sample genotyping steps are computationally light, with a median runtime of about 0.02 minutes and median memory use of 52 megabytes for the core genotyping step.</p>
<p>The authors are candid about limitations. Genotyping accuracy ultimately depends on the quality and completeness of the input pangenome: haplotypes absent from the reference panel cannot be called exactly, and novel structural variants will be assigned to the most similar available haplotype. The leave-all-out benchmarks quantify how much accuracy survives this constraint, but they also make clear that expanding pangenome references remains essential. COSIGT currently requires pre-aligned BAM or CRAM files, creating a dependency on reference-based alignment, though future work aims to accept raw FASTQ input directly and to implement haplotype subsampling to keep the method scalable as pangenomes grow. The framework also currently supports diploid genotyping, although the underlying mathematics generalizes to arbitrary ploidy by enumerating k-tuples of haplotypes.</p>
<p>The implications reach well beyond a single software release. Pangenome references are expanding rapidly in size, quality, and taxonomic breadth, and as they do, the bottleneck shifts from reference completeness to scalable genotyping. By making complex-locus genotyping reliable at 1 to 2X coverage, COSIGT opens the door to systematically mining the vast archives of low-coverage sequencing data already sitting in biobanks and legacy datasets, and to bringing archaeological and ancient genomes into population-level biomedical analyses from which their shallow depth previously excluded them. Larger cohorts become affordable at reduced sequencing cost, samples of heterogeneous quality can be combined in a single analysis, and the allelic diversity captured in pangenome references becomes practically accessible across the full spectrum of sequencing depths. The pipeline is compatible with long-read input as well, which could extend its utility to low-coverage long-read datasets too sparse for assembly. For a field that has long treated structurally complex loci as terra incognita in shallow data, the message is that the map, and the compass to read it, may finally have arrived together.</p>
<p><strong>Subject of Research:</strong> Pangenome graph-based genotyping of structurally complex genomic loci from low-coverage sequencing data</p>
<p><strong>Article Title:</strong> COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs</p>
<p><strong>Article References:</strong> Bolognini, D., Guarracino, A., Paleni, C., Dudley, T. S., Iacoviello, L., Raveane, A., Sudmant, P. H., Garrison, E., &amp; Soranzo, N. (2026). COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs. <em>Genome Biology, 27</em>(1), Article 286. <a href="https://doi.org/10.1186/s13059-026-04242-4" rel="noopener noreferrer">https://doi.org/10.1186/s13059-026-04242-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13059-026-04242-4" rel="noopener noreferrer">10.1186/s13059-026-04242-4</a></p>
<p><strong>Keywords:</strong> pangenome graphs, genotyping, low-coverage sequencing, ancient DNA, structural variation, copy-number variation, HLA typing, CYP2D6, cosine similarity, population genomics, Genome Biology, bioinformatics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222518</post-id>	</item>
		<item>
		<title>Deep Learning Tool Sharpens SNP Detection in Nanopore Transcriptome Sequencing</title>
		<link>https://scienmag.com/deep-learning-tool-sharpens-snp-detection-in-nanopore-transcriptome-sequencing/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 14:51:08 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[allele-specific expression]]></category>
		<category><![CDATA[computational genomics]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning SNP detection]]></category>
		<category><![CDATA[direct RNA sequencing]]></category>
		<category><![CDATA[direct RNA sequencing improvements]]></category>
		<category><![CDATA[genotyping]]></category>
		<category><![CDATA[genotyping in nanopore data]]></category>
		<category><![CDATA[long-read transcriptome analysis]]></category>
		<category><![CDATA[long-read transcriptomics]]></category>
		<category><![CDATA[Nanopore long-read sequencing accuracy]]></category>
		<category><![CDATA[nanopore sequencing]]></category>
		<category><![CDATA[nanopore sequencing noise reduction]]></category>
		<category><![CDATA[nanopore transcriptome sequencing]]></category>
		<category><![CDATA[NanoTS]]></category>
		<category><![CDATA[Nature Methods]]></category>
		<category><![CDATA[neural network SNP identification]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[RNA modification preservation in sequencing]]></category>
		<category><![CDATA[RNA sequencing error correction]]></category>
		<category><![CDATA[SNP calling]]></category>
		<category><![CDATA[third-generation sequencing variant analysis]]></category>
		<category><![CDATA[transcript isoform sequencing]]></category>
		<category><![CDATA[variant detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205999</guid>

					<description><![CDATA[NanoTS, a new deep learning tool published in Nature Methods, enables accurate SNP detection and genotype calling from nanopore long-read transcriptome sequencing data.]]></description>
										<content:encoded><![CDATA[<p>Researchers have unveiled NanoTS, a deep learning tool designed to bring new accuracy to the detection of single nucleotide polymorphisms, or SNPs, in nanopore long-read transcriptome sequencing data. The tool, described in a study published in Nature Methods, addresses one of the most persistent challenges in third-generation sequencing: extracting reliable genetic variants from RNA reads that are notoriously prone to errors. By harnessing neural networks trained to distinguish true biological variation from sequencing noise, NanoTS promises to make direct RNA sequencing a far more powerful instrument for genotyping and transcriptomic analysis.</p>
<p>Nanopore sequencing has transformed genomics by allowing DNA and RNA molecules to be read in their entirety, producing long reads that span entire transcripts and reveal complex splicing patterns that short-read technologies often miss. Oxford Nanopore platforms thread individual molecules through protein nanopores and infer their sequence from disruptions in an electrical current. This approach delivers reads that can stretch tens of thousands of bases, capturing full-length transcript isoforms and preserving modifications that are lost when RNA is fragmented and amplified. Yet the technology has long been hampered by a higher raw error rate compared with Illumina short-read sequencing, and this limitation has been particularly consequential for SNP calling, where a single misread base can be mistaken for a genuine variant.</p>
<p>Conventional approaches to correcting these errors typically rely on consensus building, aligning many reads to a reference genome and counting the frequency of each base at every position. Such methods work reasonably well for DNA sequencing at high coverage, but transcriptome data presents unique complications. Expression levels vary enormously across genes, meaning some transcripts are covered by thousands of reads while others yield only a handful. Low-coverage transcripts offer too few observations for statistical consensus methods to separate signal from noise, and highly expressed genes can harbor natural RNA editing events that complicate variant interpretation further. The result has been a landscape in which SNP detection from nanopore transcriptome data has demanded heavy computational filtering and often produced inconsistent results across tools and datasets.</p>
<p>NanoTS takes a fundamentally different approach by framing SNP calling as a pattern recognition problem that deep learning is exceptionally well suited to solve. Rather than relying solely on base-level alignments, the tool learns from the rich features embedded in nanopore signal data, the raw electrical current measurements that underlie every base call. Neural networks can detect subtle signatures that distinguish systematic sequencing errors, which tend to follow repeatable patterns tied to sequence context and pore dynamics, from genuine genetic variation, which manifests consistently across independent molecules. This capacity to model complex, high-dimensional relationships in the data allows NanoTS to make confident genotype calls even in regions where traditional statistical methods falter.</p>
<p>The implications of accurate SNP calling from direct RNA sequencing extend across a wide range of biological and biomedical applications. Variant detection at the transcript level reveals which alleles are actively expressed, a dimension invisible to DNA-based genotyping. Allele-specific expression, in which one copy of a gene is transcribed more abundantly than the other, plays a critical role in imprinted genes, autoimmune disease, cancer biology and X-chromosome inactivation. Researchers studying these phenomena have historically needed to combine DNA genotyping with RNA expression analysis from separate experiments, introducing uncertainty about whether observed expression patterns truly reflect the same cells and conditions. A tool that can call genotypes directly from transcriptome data collapses these two measurements into a single, internally consistent dataset.</p>
<p>The deep learning architecture at the heart of NanoTS represents part of a broader movement in genomics toward learned models that outperform hand-crafted algorithms. Basecalling, the process of converting raw nanopore current signals into nucleotide sequences, was revolutionized years ago by neural networks that replaced earlier hidden Markov model approaches, dramatically improving read accuracy. Variant calling is now following a similar trajectory. DeepVariant, developed by Google, demonstrated that convolutional neural networks could outperform classical statistical variant callers on Illumina data by treating genomic alignment piles as images. NanoTS extends this paradigm to the specific and demanding context of long-read transcriptome sequencing, where error profiles, coverage distributions and biological variation patterns differ substantially from genomic DNA sequencing.</p>
<p>Training such a model requires data of exceptional quality, since the network must learn to associate signal and alignment patterns with known ground-truth genotypes. The developers of NanoTS tackled this challenge by curating training examples in which true variants could be established with high confidence, allowing the network to internalize the distinctions between authentic polymorphism and sequencing artifact. Once trained, the model generalizes across genes, transcripts and coverage regimes, producing variant calls that remain reliable even for transcripts sequenced at low depth, a scenario that has historically been the Achilles heel of long-read transcriptome analysis.</p>
<p>The clinical dimension of this work is particularly compelling. Personalized medicine increasingly depends on knowing both which variants a patient carries and how those variants are expressed in relevant tissues. RNA sequencing of patient samples, whether from tumor biopsies or blood cells, is already routine in many clinical and research settings. If those same RNA reads can yield accurate genotype information without a parallel DNA sequencing run, the cost and complexity of comprehensive molecular profiling drops considerably. Cancer genomics stands to benefit substantially, as tumor transcripts carry both expressed mutations and the allele-specific expression patterns that reveal loss of heterozygosity and other genome alterations relevant to treatment decisions.</p>
<p>Beyond clinical applications, NanoTS opens doors for basic research into population genetics and evolutionary biology at the transcript level. Understanding how genetic variation shapes transcript diversity across individuals and species requires tools that can reliably measure both. Long-read transcriptome sequencing captures the full complexity of isoform expression, and with accurate SNP calling layered on top, researchers can begin to map how specific variants influence splicing, expression levels and transcript structure in ways that short-read studies have only approximated. The combination of full-length isoform resolution and direct genotype calling provides a uniquely integrated view of the relationship between genome and transcriptome.</p>
<p>The publication of NanoTS in Nature Methods underscores the growing recognition that computational innovation is as essential to unlocking the potential of new sequencing technologies as the hardware itself. As nanopore sequencing continues its rapid improvements in accuracy and throughput, tools like NanoTS ensure that the analytical pipeline keeps pace, converting noisy raw signals into biologically meaningful and clinically actionable insights. For laboratories worldwide working with direct RNA sequencing, the arrival of a robust deep learning solution for SNP calling marks a significant step toward making long-read transcriptomics a complete, standalone platform for both discovery and diagnosis.</p>
<p><strong>Subject of Research:</strong> A deep learning tool for accurate SNP calling and genotype detection in nanopore long-read transcriptome sequencing data.</p>
<p><strong>Article Title:</strong> NanoTS: a deep learning tool for accurate SNP calling in nanopore long-read transcriptome data</p>
<p><strong>Article References:</strong> NanoTS: a deep learning tool for accurate SNP calling in nanopore long-read transcriptome data. (n.d.). <a href="https://doi.org/10.1038/s41592-026-03225-4" rel="noopener noreferrer">https://doi.org/10.1038/s41592-026-03225-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41592-026-03225-4" rel="noopener noreferrer">10.1038/s41592-026-03225-4</a></p>
<p><strong>Keywords:</strong> NanoTS, deep learning, SNP calling, nanopore sequencing, long-read transcriptomics, genotyping, direct RNA sequencing, variant detection, neural networks, Nature Methods, allele-specific expression, computational genomics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205999</post-id>	</item>
		<item>
		<title>New Software pSTRminer Uncovers Thousands of Forensic DNA Markers in Cattle</title>
		<link>https://scienmag.com/new-software-pstrminer-uncovers-thousands-of-forensic-dna-markers-in-cattle/</link>
		
		<dc:creator><![CDATA[William Thompson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 12:25:12 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[animal forensic genetics]]></category>
		<category><![CDATA[animal forensics]]></category>
		<category><![CDATA[automated forensic DNA evaluation]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[bioinformatics tools for DNA discovery]]></category>
		<category><![CDATA[cattle]]></category>
		<category><![CDATA[DNA fingerprinting in criminal investigations]]></category>
		<category><![CDATA[DNA profiling of livestock]]></category>
		<category><![CDATA[forensic analysis of poached animals]]></category>
		<category><![CDATA[forensic DNA markers in cattle]]></category>
		<category><![CDATA[forensic genetics]]></category>
		<category><![CDATA[forensic investigation of wildlife crimes]]></category>
		<category><![CDATA[genome-wide DNA marker identification]]></category>
		<category><![CDATA[genotyping]]></category>
		<category><![CDATA[next-generation sequencing]]></category>
		<category><![CDATA[polymorphic short tandem repeats in animals]]></category>
		<category><![CDATA[polymorphism]]></category>
		<category><![CDATA[population genetics]]></category>
		<category><![CDATA[pSTRminer]]></category>
		<category><![CDATA[pSTRminer software for DNA analysis]]></category>
		<category><![CDATA[short tandem repeats]]></category>
		<category><![CDATA[standardization in animal forensics]]></category>
		<category><![CDATA[STR database]]></category>
		<category><![CDATA[tetranucleotide STRs]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194111</guid>

					<description><![CDATA[Researchers have developed pSTRminer, an integrated bioinformatic tool that mines cattle genomes for polymorphic short tandem repeats and builds a standardized forensic marker database.]]></description>
										<content:encoded><![CDATA[<p>When investigators arrive at a crime scene, they do not always find human DNA. Hair from a dog, blood from a cat, or traces of livestock can link a suspect to a location, identify poached wildlife, or resolve disputes over stolen animals. Animal forensic genetics has quietly become an essential pillar of modern criminal investigation, yet it has long operated with far less standardization than its human counterpart. Now, a team of forensic scientists at Sun Yat-sen University in Guangzhou, China, has unveiled a tool that could change that. In a study published in the International Journal of Legal Medicine, the researchers introduce pSTRminer, an integrated bioinformatic software package designed to automate the discovery and evaluation of polymorphic short tandem repeats, or STRs, across entire genomes and entire populations.</p>
<p>Short tandem repeats are stretches of DNA in which a short sequence of two to six base pairs is repeated over and over, such as ATATATAT. Because the number of repeats varies widely between individuals, STRs form the backbone of DNA profiling in human forensics. Standardized human STR genotyping systems, built on carefully validated panels of markers, allow laboratories around the world to produce comparable, court-admissible profiles. Animal forensics has never enjoyed that level of coordination. Validated STR markers for most domestic and wild species are scarce, and the markers that do exist are often dinucleotide STRs, repeats of just two base pairs, which are notoriously prone to genotyping artifacts such as stutter, the generation of spurious off-by-one peaks that complicate interpretation. Population data, which allow forensic scientists to calculate the statistical weight of a match, are frequently missing altogether.</p>
<p>The team behind pSTRminer, led by Jiajun Liu, Zhentang Liu, and senior authors Hongyu Sun and Riga Wu of the Faculty of Forensic Medicine at Zhongshan School of Medicine, set out to close these gaps with a single, scalable computational framework. The software automates what has traditionally been a fragmented, largely manual workflow: scanning a reference genome for STR loci, genotyping those loci in large collections of whole-genome sequencing data, and then scoring each locus for the properties that matter in forensic practice, including genotyping success rate and polymorphism information content, a standard measure of how informative a genetic marker is for distinguishing individuals.</p>
<p>To demonstrate the power of the approach, the researchers turned to domestic cattle, Bos taurus, one of the most economically and forensically significant livestock species in the world. Applying pSTRminer to the cattle reference genome, they identified 775,444 STRs de novo, a catalog of repeat loci far exceeding anything previously assembled for the species. They then genotyped this catalog using whole-genome sequencing data from 60 Chinese cattle and 111 African cattle, two populations chosen to represent sharply divergent genetic backgrounds. The logic is straightforward but important: a marker that appears highly variable in only one breed or region may be nearly useless elsewhere, so evaluating polymorphism across diverse lineages is essential before any locus can be recommended for global forensic use.</p>
<p>From this population-scale analysis, the team constructed the cattle STR database, or CSDB, a curated resource containing only those loci that met stringent quality criteria: a genotyping success rate of at least 40 percent and a polymorphism information content of at least 0.5. These thresholds ensure that the database holds markers that both amplify reliably in the laboratory and carry enough variation to discriminate between individuals. The sensitivity of the database to the genotyping success rate threshold was examined in supplementary analyses, giving future users a transparent view of how the marker set changes as criteria are tightened or relaxed.</p>
<p>Computational screening alone, however, is not enough for forensic work. Markers destined for casework must survive contact with real samples. The researchers therefore experimentally validated a panel of loci in a local Chinese cattle population of 145 animals using next-generation sequencing. Thirty tetranucleotide STRs, repeats of four base pairs, and 33 dinucleotide STRs were randomly selected from the database and tested. The validation confirmed that the markers were reliable, and it produced a nuanced picture of the trade-offs between repeat types. Tetranucleotide STRs showed lower average polymorphism than their dinucleotide counterparts, meaning they tend to be somewhat less variable across individuals. But they carried a decisive advantage: significantly lower stutter ratios, a difference the authors report as statistically significant at p less than 0.05. In practical terms, four-base-pair repeats generate fewer genotyping artifacts, producing cleaner, easier-to-interpret profiles.</p>
<p>That finding matters because stutter is one of the most persistent headaches in STR analysis. When a polymerase copies a repeat tract, it occasionally slips, adding or dropping a repeat unit and creating a minor artifact peak one repeat shorter or longer than the true allele. Dinucleotide repeats, with their short two-base motif, are especially vulnerable to this slippage. A marker with high stutter can obscure genuine alleles, particularly in degraded or low-template samples common in forensic contexts. By demonstrating that certain tetranucleotide STRs can actually surpass dinucleotide STRs in polymorphism while producing far fewer artifacts, the study lays out a viable path toward building animal STR panels that are both highly discriminative and technically robust. Systematic screening across the CSDB revealed that such high-performing tetranucleotide loci are not rare exceptions but a discoverable resource waiting to be tapped.</p>
<p>The broader significance of pSTRminer extends well beyond cattle. The software integrates established components of the modern genomics pipeline, drawing on widely used tools for read preprocessing, alignment, and STR genotyping, and wraps them into a reproducible workflow that reduces the manual operations required to move from raw sequencing data to a validated marker panel. The authors provide detailed documentation of the commands needed to reproduce their analyses, and supplementary tables include the formulas used to calculate forensic parameters, the overlap between STRs currently in use and those newly identified in cattle, and recommended analytical thresholds for heterozygote balance at validated loci. In effect, the study offers not just a database but a blueprint that other laboratories can follow to develop standardized STR systems for dogs, cats, horses, yaks, wildlife species, or any organism with a reference genome and population sequencing data.</p>
<p>The need for such tools is well documented in the forensic literature. Individual identification systems based on STR panels have been developed for domestic cats, dogs, and horses, and microsatellite marker sets have been proposed for parentage testing in cattle, yaks, and Chinese Holstein bulls. Yet each of these efforts has relied on comparatively small collections of markers, often selected without genome-wide polymorphism data or population-scale validation. Human forensics, by contrast, has moved decisively toward expanded multiplex systems and sequencing-based genotyping, supported by open population databases built from large-scale sequencing projects such as the 1000 Genomes Project. pSTRminer aims to bring animal forensics closer to that standard, enabling marker discovery at the scale the human field now takes for granted.</p>
<p>The work, supported by the National Natural Science Foundation of China, was approved by the Institutional Animal Care and Use Committee of Sun Yat-sen University, and the authors acknowledge the publicly available whole-genome sequencing data from the NCBI Sequence Read Archive that made the population analyses possible. For forensic scientists, the arrival of pSTRminer and the CSDB marks a shift from ad hoc marker selection to systematic, data-driven panel design. A single hair from a stolen calf, a bloodstain on a suspect&#8217;s boot, or a trace of tissue from a poached animal could soon be profiled with the same rigor and statistical confidence that human DNA evidence enjoys. As genome sequencing becomes cheaper and reference genomes accumulate for more species, the framework promises to make animal forensic genetics faster, cleaner, and more defensible, one well-validated repeat at a time.</p>
<p><strong>Subject of Research:</strong> Genome-wide identification and population-scale evaluation of polymorphic short tandem repeats for animal forensic genetics</p>
<p><strong>Article Title:</strong> pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats</p>
<p><strong>Article References:</strong> pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats. (n.d.). <a href="https://doi.org/10.1007/s00414-026-04002-w" rel="noopener noreferrer">https://doi.org/10.1007/s00414-026-04002-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00414-026-04002-w" rel="noopener noreferrer">10.1007/s00414-026-04002-w</a></p>
<p><strong>Keywords:</strong> pSTRminer, short tandem repeats, forensic genetics, cattle, bioinformatics, genotyping, polymorphism, next-generation sequencing, STR database, animal forensics, population genetics, tetranucleotide STRs</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194111</post-id>	</item>
	</channel>
</rss>
