Structural variants—large-scale rearrangements of the genome that include deletions, duplications, inversions and insertions of thousands to millions of DNA letters—are among the most common sources of genetic disease in humans. Yet they remain stubbornly difficult to find with the technology that dominates clinical sequencing today: short-read sequencing, which chops the genome into small fragments and reassembles the sequence computationally. A new method called dicast, described in the journal Genome Biology, applies machine learning to this problem and reports substantially more accurate detection of structural variants from short-read data than existing tools, along with evidence that it can recover pathogenic variants that consensus approaches leave behind.
The method was developed by a team led by Nico Alavi, M-Hossein Moeinzadeh and Jakob Hertzberg, who contributed equally, working across the Max Planck Institute for Molecular Genetics in Berlin, Lucid Genomics, and collaborating institutions in Italy and Finland, with Martin Vingron of the Max Planck Institute as corresponding author. Their central insight is that deciding whether a candidate structural variant call is real is a problem well suited to machine learning: candidate calls carry a wealth of evidence in their alignment patterns and genomic context, and a trained classifier can weigh that evidence far more consistently than the hard-coded heuristics that traditional detection tools rely on.
How dicast works, in technical terms, is a two-stage design. Upstream algorithms first generate candidate structural variant calls from the alignment of short sequencing reads to a reference genome. dicast then takes each candidate and scores it using a set of features extracted from the read alignments—such as the evidence supplied by split reads that map to two different genomic locations, and by read pairs whose insert sizes deviate from expectation—combined with features describing the genomic context around the candidate breakpoint. A machine-learning model, trained on examples of calls known to be true and false, converts these features into a calibrated score that discriminates genuine variants from sequencing artifacts and mapping errors.
The quality of that training depends entirely on the ground truth, and this is where the team invested heavily. Rather than borrowing a previously published benchmark, the researchers built a new multi-technology ground truth from nine samples, integrating evidence from multiple sequencing technologies and applying extensive manual curation. Structural variants sit at the awkward intersection of being too large for standard short-read analysis to resolve cleanly and too numerous for manual review of every call, so combining independent sources of evidence—and verifying them by hand—produces a training resource with unusually high confidence. This matters because previous machine-learning attempts in this area have often been limited by noisy or incompletely validated labels, which cap the accuracy any model can achieve.
On benchmarks built from this curated ground truth, dicast outperformed existing short-read structural variant callers and also outperformed consensus approaches, which pool the outputs of multiple callers and keep only the calls on which they agree. The most striking result is a gain in sensitivity at high precision: dicast recovered substantially more true positives while maintaining accuracy levels comparable to or better than the tools it was compared against. This is the regime that matters most in practice, because consensus-based pipelines trade sensitivity for precision—filtering out artifacts by requiring agreement, but in the process discarding genuine variants that only one caller detects. By learning which single-caller evidence is trustworthy, dicast aims to keep those true variants without importing the noise.
The clinical implications were tested directly. The team applied dicast to multiple disease cohorts and found that it identified all pathogenic structural variants in those cohorts—every variant that clinical analysis had already established as causative for the patients’ disease. Demonstrating that no known pathogenic call is lost is the baseline requirement for any diagnostic tool, since a missed pathogenic variant in a clinical setting means a missed or delayed diagnosis. dicast passed that test across the cohorts examined, which included patient groups for whom structural variant detection is a central diagnostic question.
Beyond simply matching existing clinical findings, the method found variants that current practice misses. In the disease cohorts analyzed, dicast identified 20 percent more candidate pathogenic deletions than consensus approaches did. Deletions—stretches of DNA removed from the genome—are a particularly common class of pathogenic structural variant, implicated in developmental disorders, congenital malformations and neuromuscular disease. A 20 percent increase in candidate pathogenic deletions recovered from the same short-read data, without any new sequencing, represents diagnostic information that was already present in the patients’ files but invisible to standard analysis pipelines. For patients whose previous genetic tests came back negative, reanalysis with a more sensitive detector can turn an unexplained condition into a molecularly diagnosed one.
The cohorts used for the clinical evaluation reflect the diversity of settings in which structural variant detection matters. The study analyzed a congenital limb malformation cohort approved by the ethics committee of Charité – Universitätsmedizin Berlin, a neuromuscular disease cohort studied in collaboration with researchers at the University of Helsinki and Helsinki University Hospital, and an atrial fibrillation cohort drawn from a previously published study. Across these groups, written informed consent was obtained from all individuals or their legal guardians, and the analyses were conducted under the relevant ethical approvals. The breadth of the evaluation suggests the method’s benefits are not confined to one disease area or one type of variant signature, but hold across the heterogeneous landscape of clinically relevant rearrangements.
Why has this problem resisted solution for so long? Short-read sequencing dominates clinical workflows for reasons of cost, throughput and established regulatory practice, even though longer reads resolve structural variants far more directly. A single short read cannot span a deletion of tens of thousands of bases; instead, detectors must infer the variant from indirect signals—reads that straddle the breakpoint, or read pairs that map suspiciously far apart. Each of these signals is noisy, and different tools weight them differently, which is why their call sets disagree so often. Machine learning offers a way to integrate the signals rather than choosing between them, and to learn the weights from data rather than from developer intuition. dicast’s contribution is to show that, given a sufficiently rigorous training set, this integration can exceed both individual tools and their consensus.
The work, published open access in Genome Biology on 16 September 2026 and carried out with support from the German Federal Ministry of Education and Research under the iGenVar program, arrives at a moment when clinical genetics is pushing hard toward more complete variant detection at lower cost. If its performance holds in independent clinical settings, the practical consequence would be straightforward: laboratories running ordinary short-read sequencing would extract more diagnostic value from the same data, patients would receive fewer unresolved test results, and the long-standing gap between what short-read instruments can technically capture and what analysis pipelines actually recover would narrow. The study also underscores a broader lesson for genomics, that the ground truth on which computational methods are trained is not a detail but the foundation—dicast’s accuracy is inseparable from the manually curated, multi-technology benchmark its creators built to teach it.
Subject of Research: Machine learning for structural variant detection from short-read genome sequencing
Article Title: dicast: a machine learning method for accurate structural variant detection from short-read sequencing data
Article References: Alavi, N., Moeinzadeh, M.-H., Hertzberg, J., Melo, U. S., Al Raei, L. W., Infantino, P., Ghareghani, M., Savarese, M., Mundlos, S., & Vingron, M. (2026). dicast: a machine learning method for accurate structural variant detection from short-read sequencing data. Genome Biology, 27(1), Article 292. https://doi.org/10.1186/s13059-026-04280-y
Image Credits: AI Generated
DOI: 10.1186/s13059-026-04280-y
Keywords: structural variants, machine learning, short-read sequencing, genomics, dicast, Genome Biology, diagnostics, deletions, variant calling, clinical genetics, ground truth, precision medicine
Cite Scienmag News
Juliet Wilcox. (September 23, 2026). Machine learning method dicast spots disease-causing DNA deletions short-read tools miss. Scienmag. https://scienmag.com/machine-learning-method-dicast-spots-disease-causing-dna-deletions-short-read-tools-miss/
Juliet Wilcox. "Machine learning method dicast spots disease-causing DNA deletions short-read tools miss." Scienmag, 23 September 2026, https://scienmag.com/machine-learning-method-dicast-spots-disease-causing-dna-deletions-short-read-tools-miss/. Accessed 23 September 2026.
Juliet Wilcox. "Machine learning method dicast spots disease-causing DNA deletions short-read tools miss." Scienmag. September 23, 2026. https://scienmag.com/machine-learning-method-dicast-spots-disease-causing-dna-deletions-short-read-tools-miss/








