Scientists have reconstructed the most complete diploid human genome yet produced, creating a high-resolution sequence that contains both copies of every chromosome inherited from an individual’s parents. The achievement, led by researchers from Johns Hopkins University, the National Human Genome Research Institute and the National Institute of Standards and Technology, marks a major advance beyond the conventional human reference genome. Rather than identifying a person’s genetic differences by comparing them with an incomplete standard, the new method reconstructs the individual genome itself, including regions that have historically been too repetitive or complex to decode.
The work was carried out by the Telomere-to-Telomere, or T2T, Consortium using HG002, a human genome sample obtained from a living donor and widely used as a reference material by sequencing and diagnostic laboratories. The researchers assembled each chromosome from one telomere, the protective structure at one end, to the other, producing two separate chromosome sets that represent the maternal and paternal genomes. This is technically more difficult than assembling a single genome because the two copies are highly similar but not identical. Computational systems must determine which DNA fragments belong to which parental chromosome while preserving every small difference between them.
The new sequence adds more than 900 million DNA letters that were absent from previous benchmarks and reveals roughly 15% more of the genome than the earlier reference standard. These newly accessible regions include parts of both sex chromosomes, highly repetitive stretches, and sequences containing genes and regulatory elements that may influence disease risk. Such regions have often been excluded from clinical sequencing because standard technologies struggled to read them accurately or because researchers could not determine their correct position within the genome. By resolving these difficult segments, the T2T approach could expose genetic variants that have remained invisible in routine testing.
The advance builds on the consortium’s landmark 2022 completion of the first truly complete human genome sequence. That project filled in approximately the final 8% of a single reference genome, including many repetitive regions and the previously incomplete Y chromosome. The new effort goes further by applying improved sequencing platforms, assembly algorithms and validation methods to a diploid genome. Long-read sequencing technologies were central to the work because they generate DNA fragments thousands or even millions of letters long, allowing researchers to span repetitive sequences that would be broken into ambiguous pieces by older short-read methods.
Accurately assigning genes to each chromosome copy was another essential part of the project. Scientists at Johns Hopkins led by computational biologist Steven Salzberg analyzed the two chromosome sets to identify and annotate their genes, while Michael Schatz’s laboratory contributed to extensive validation of the assembly. Independent checks were used to test whether the reconstructed sequence contained errors, missing segments or incorrectly joined fragments. The result is intended not only as a biological reference but also as a measurement standard for companies developing DNA sequencing instruments, analysis software and clinical diagnostics.
Researchers say the development could change the logic of medical genomics. Current clinical analyses generally search for variants that differ from a standard reference genome. This strategy can perform well when a patient’s DNA resembles the reference, but it becomes less reliable in genomic regions where the reference is incomplete or structurally different. A complete genome assembled for each patient would instead provide an individualized baseline. Genetic analysis could then examine substitutions, insertions, deletions, duplications and larger rearrangements across the entire sequence without automatically discarding regions that do not align well with the traditional reference.
The immediate medical benefit could be improved diagnosis for children and adults with rare genetic disorders. Genome sequencing is already used in such cases, but more than half of patients may still leave testing without a clear molecular explanation. Missing or misread regions can conceal the mutation responsible for disease, particularly when it lies in a repetitive sequence or involves a complex structural change. A complete diploid assembly could help clinicians identify these causes more accurately, potentially ending years of uncertainty for families and guiding treatment, monitoring and reproductive decisions.
The same approach may eventually strengthen predictions for common diseases. Variants in the BRCA1 and BRCA2 genes are already used to estimate breast cancer risk, but researchers believe that many additional risk-associated changes remain undiscovered in difficult-to-sequence portions of the genome. More complete reference data could improve studies of cancer, cardiovascular disease, immune disorders and neuropsychiatric conditions. When combined with genomes from large and diverse populations, these sequences could also support artificial intelligence models trained to recognize disease-related patterns while reducing the bias created by relying on a single, historically limited reference genome.
The consortium estimates that a complete and highly accurate human genome can now be generated for about $5,000, compared with the roughly $5 billion, in current dollars, spent on the Human Genome Project, which concluded in 2003. Although routine whole-genome sequencing still raises questions about privacy, data storage, consent and the interpretation of uncertain findings, the technical barrier is rapidly falling. The researchers envision a future in which a person’s complete genome is sequenced early in life, securely linked to medical records and revisited as scientific knowledge improves. The study is part of a broader package of work in Cell and Cell Genomics that also presents complete or near-complete genomes for macaques, marmosets, zebra finches, rats, voles, horses, donkeys and giraffes, extending the same genomic precision to research on evolution, biodiversity, agriculture and animal health.
Subject of Research: Complete diploid human genome sequencing and personalized genomics
Article Title: Complete, high-quality diploid human genome reconstructed from telomere to telomere
Web References:
https://engineering.jhu.edu/faculty/adam-phillippy/
https://hub.jhu.edu/2022/03/31/johns-hopkins-scientists-first-complete-sequence-human-genome/
https://www.bme.jhu.edu/people/faculty/steven-l-salzberg/
https://engineering.jhu.edu/faculty/michael-schatz/
References:
Cell, DOI: 10.1016/j.cell.2026.06.016
Keywords
Human genome sequencing, diploid genome, Telomere-to-Telomere Consortium, personalized genomics, genetic disease diagnosis, long-read sequencing, genomic medicine, structural variants, precision medicine, genome assembly

