<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>telomere-to-telomere genome assembly &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/telomere-to-telomere-genome-assembly/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 28 Aug 2026 11:51:35 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>telomere-to-telomere genome assembly &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Complete Guazuma ulmifolia Genome Reveals Evolution, Drought Adaptation, and Flavonoid Biosynthesis</title>
		<link>https://scienmag.com/complete-guazuma-ulmifolia-genome-reveals-evolution-drought-adaptation-and-flavonoid-biosynthesis/</link>
		
		<dc:creator><![CDATA[Gavin Prescott]]></dc:creator>
		<pubDate>Fri, 28 Aug 2026 11:51:32 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[chromosome evolution in plants]]></category>
		<category><![CDATA[chromosome structure and stability in plants]]></category>
		<category><![CDATA[chromosome structure in Malvaceae]]></category>
		<category><![CDATA[climate resilience in cacao relatives]]></category>
		<category><![CDATA[drought tolerance in tropical trees]]></category>
		<category><![CDATA[flavonoid biosynthesis pathways]]></category>
		<category><![CDATA[genetic basis of drought resistance]]></category>
		<category><![CDATA[genome sequencing of Malvaceae species]]></category>
		<category><![CDATA[Guazuma ulmifolia genome]]></category>
		<category><![CDATA[medicinal plant compounds]]></category>
		<category><![CDATA[plant evolutionary genomics]]></category>
		<category><![CDATA[plant genome sequencing techniques]]></category>
		<category><![CDATA[plant stress adaptation genetics]]></category>
		<category><![CDATA[telomere-to-telomere genome assembly]]></category>
		<guid isPermaLink="false">https://scienmag.com/complete-guazuma-ulmifolia-genome-reveals-evolution-drought-adaptation-and-flavonoid-biosynthesis/</guid>

					<description><![CDATA[A wild relative of cacao has yielded a remarkably complete genetic blueprint that could help scientists understand how tropical trees withstand drought, evolve new chromosome structures and produce medically interesting plant compounds. In a study published in Plant Cell Reports, researchers report the first telomere-to-telomere, chromosome-level genome assembly of Guazuma ulmifolia, a Malvaceae species known [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>A wild relative of cacao has yielded a remarkably complete genetic blueprint that could help scientists understand how tropical trees withstand drought, evolve new chromosome structures and produce medically interesting plant compounds. In a study published in <em>Plant Cell Reports</em>, researchers report the first telomere-to-telomere, chromosome-level genome assembly of <em>Guazuma ulmifolia</em>, a Malvaceae species known by names including West Indian elm and guácima. The 311.31-million-base-pair genome offers an unusually detailed view of a plant lineage that is ecologically valuable, used in traditional medicine and closely related to <em>Theobroma cacao</em>, the tree responsible for chocolate. As climate change intensifies drought risk in cacao-growing regions, the genome could become a foundation for identifying genetic features associated with stress tolerance in cacao and other crops.</p>
<p>The scale of the achievement lies in the completeness of the assembly. Rather than stitching together a genome that still contains numerous unresolved gaps, the researchers produced a telomere-to-telomere, or T2T, reference genome designed to represent chromosome sequences from one end to the other. Telomeres are repetitive DNA structures that protect chromosome ends, while centromeres are specialized regions involved in chromosome movement during cell division. Both regions are difficult to sequence and assemble because they often contain long, repetitive stretches of DNA. The new assembly reaches a contig N50 of 35.19 million base pairs, a measure indicating that relatively long continuous DNA segments make up the assembly, and achieves 98.70 percent BUSCO completeness. BUSCO assesses whether a genome contains a standardized collection of genes expected to be conserved in a particular lineage, making the result a strong indication that most of the organism’s core genetic content has been captured.</p>
<p>The genome also reveals that repetitive DNA occupies 27.43 percent of <em>G. ulmifolia</em>. The largest contribution comes from long terminal repeat retrotransposons, a class of mobile genetic elements that can copy themselves through an RNA intermediate and insert the copy elsewhere in the genome. These elements are sometimes described as genomic parasites, but they can also influence genome structure, gene regulation and evolutionary change. Their accumulation can expand genome size, alter the spacing between genes and contribute to chromosome rearrangements. By comparing <em>G. ulmifolia</em> with other members of the Malvaceae, the researchers found that differences in genome size are associated with two evolutionary forces: the history of polyploidization and the activity of transposable elements. Polyploidization occurs when an organism acquires additional complete sets of chromosomes, while transposable-element dynamics can add or remove large quantities of DNA over time.</p>
<p>That comparison places the cacao relatives within a broader history of genomic expansion and contraction. In plants, polyploid genomes may later undergo diploidization, a long process in which duplicated genes are lost, silenced or reorganized until the genome behaves more like a diploid one. Repeated cycles of duplication and restructuring can leave behind duplicated genes and altered chromosome relationships. Transposable elements add another layer of change by moving through the genome and generating mutations or large-scale rearrangements. A high-quality assembly makes it possible to distinguish these processes more accurately than a fragmented draft genome would. Instead of seeing isolated sequences, scientists can examine how genes and repeats are positioned along entire chromosomes and compare those arrangements across related species.</p>
<p>One of the most striking findings concerns chromosome evolution. Using comparative genomic analyses and ancestral karyotype reconstruction, the team identified five lineage-specific chromosome fusion events that distinguish <em>G. ulmifolia</em> from <em>T. cacao</em>. A chromosome fusion occurs when two ancestral chromosomes become joined into one, changing the number and organization of chromosomes without necessarily destroying the genes they carry. Such events can affect meiotic pairing, gene linkage and the inheritance of traits. Reconstructing them is similar to comparing the layouts of related genomes and tracing which segments were joined, separated or rearranged during evolution. The result offers a clearer explanation of how the chromosomes of this wild cacao relative came to differ from those of cultivated cacao and provides a framework for interpreting structural variation within the group.</p>
<p>The study’s practical importance centers on drought adaptation. Climate change is already placing pressure on tropical agriculture, and cacao is particularly vulnerable because its production depends on stable moisture and temperature conditions. The researchers identified tandem duplication-associated expansions in two stress-related gene families: late embryogenesis abundant, or LEA, genes and glutathione S-transferase, or GST, genes. Tandem duplication occurs when a DNA segment is copied and the resulting gene copies remain adjacent on the same chromosome. Over evolutionary time, duplicated copies can retain the original function, divide the original function between them or acquire new roles. LEA proteins are commonly associated with protection against cellular dehydration, while GST enzymes participate in detoxification and help plants manage reactive molecules generated during environmental stress.</p>
<p>The presence of expanded LEA and GST families does not by itself prove that these genes make <em>G. ulmifolia</em> drought tolerant. Establishing that connection will require experiments in which plants are exposed to controlled water limitation and the activity of individual genes is measured alongside physiological traits such as water use, photosynthesis, membrane stability and recovery after rewatering. Nevertheless, the genomic pattern gives researchers a shortlist of candidates for such tests. Because the species is a wild relative of cacao, its stress-associated genes could eventually inform comparative breeding or genetic engineering strategies. The immediate value is as a discovery resource: scientists can now investigate whether particular versions or arrangements of LEA and GST genes are associated with survival in dry environments.</p>
<p>The genome also sheds light on flavonoid biosynthesis, the metabolic pathway that produces a diverse family of plant compounds involved in pigmentation, defense and protection from ultraviolet radiation. Flavonoids include molecules with antioxidant properties, although activity observed in chemical or cell-based assays should not automatically be interpreted as a clinical benefit in humans. In <em>G. ulmifolia</em>, genes involved in flavonoid production were largely conserved in copy number rather than dramatically expanded. Their activity, however, varied among tissues, indicating that the plant regulates the pathway according to biological context. A gene expressed strongly in leaves may contribute to protection from sunlight or herbivores, whereas activity in bark or other tissues could reflect different defensive or developmental functions.</p>
<p>This distinction between gene copy number and gene expression is central to understanding plant chemistry. A pathway can produce different quantities or combinations of compounds without acquiring many additional genes, simply by switching existing genes on or off in different organs or at different stages of development. The researchers’ expression results therefore identify candidate genes for investigating the secondary metabolism of <em>G. ulmifolia</em>, a species already associated with tannins, proanthocyanidins and other phenolic compounds. Future work could connect tissue-specific expression with measured metabolite profiles, environmental conditions and biological activity. Such studies may clarify which genomic features control the plant’s chemical diversity, but the new genome itself is a starting point rather than evidence that extracts from the tree are safe or effective treatments.</p>
<p>The assembly was generated through a combination of modern genome-analysis approaches, including long-read sequencing, chromosome-scale organization and comparative computational analysis. Long reads are valuable because they can span repetitive regions that defeat short-read methods, while chromosome conformation data can reveal which DNA fragments physically interact inside the nucleus and therefore belong near one another. Gene prediction and functional annotation then match genomic sequences to likely coding regions and known protein families. The authors have deposited the raw sequencing data in the GenBank Sequence Read Archive under project PRJNA1279801, allowing other researchers to examine the underlying data. The work was funded by the National Natural Science Foundation of China and the Key Laboratory of Mass Spectrometry Imaging and Metabolomics at Minzu University of China.</p>
<p>For cacao researchers, the new reference genome could serve as a bridge between evolutionary biology and crop improvement. Wild relatives often contain genetic variation lost during domestication, including traits that help plants cope with pathogens, heat or water scarcity. A reference genome does not immediately create a drought-resistant cacao variety, but it enables more precise comparisons between species and populations. Researchers can search for conserved genes, detect structural differences, map candidate regions associated with stress responses and design molecular markers for breeding. The chromosome fusion history is equally important because large rearrangements can influence how easily genes are inherited together. With a complete genomic map in hand, scientists have a sharper tool for exploring the evolutionary innovations that allowed a tropical tree related to cacao to persist across changing environments—and for asking whether some of those innovations can help protect the future of chocolate.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> The telomere-to-telomere genome, chromosome evolution, drought adaptation and flavonoid biosynthesis of <i>Guazuma ulmifolia</i>, a wild relative of cacao</p>
<p><strong>Article Title:</strong> Complete telomere-to-telomere genome assembly of <i>Guazuma ulmifolia</i> uncovers evolutionary mechanisms, drought adaptation, and flavonoid biosynthesis</p>
<p><strong>Article References:</strong> Dorjee, T., Cui, Y., Liu, B., Richardson, J. E., &amp; Gao, F. (2026). Complete telomere-to-telomere genome assembly of Guazuma ulmifolia uncovers evolutionary mechanisms, drought adaptation, and flavonoid biosynthesis. <em>Plant Cell Reports, 45</em>(9), Article 273. <a href="https://doi.org/10.1007/s00299-026-03948-w" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s00299-026-03948-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00299-026-03948-w" target="_blank" rel="noopener noreferrer">10.1007/s00299-026-03948-w</a></p>
<p><strong>Keywords:</strong> Guazuma ulmifolia, cacao wild relatives, telomere-to-telomere genome assembly, drought adaptation, comparative genomics, chromosome fusion, transposable elements, LEA genes, GST genes, flavonoid biosynthesis</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">183571</post-id>	</item>
		<item>
		<title>Scientists produce the most complete brown rat DNA profile yet</title>
		<link>https://scienmag.com/scientists-produce-the-most-complete-brown-rat-dna-profile-yet/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Fri, 07 Aug 2026 07:26:39 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advancements in genomics technology]]></category>
		<category><![CDATA[Brown rat genome sequencing]]></category>
		<category><![CDATA[complete DNA profile]]></category>
		<category><![CDATA[genetic basis of disease in rats]]></category>
		<category><![CDATA[genetic variation in rats]]></category>
		<category><![CDATA[genome complexity and gene discovery]]></category>
		<category><![CDATA[impact on preclinical experiment interpretation]]></category>
		<category><![CDATA[implications for biomedical research]]></category>
		<category><![CDATA[long-read DNA sequencing]]></category>
		<category><![CDATA[rat models in disease studies]]></category>
		<category><![CDATA[sex-chromosome organization in rodents]]></category>
		<category><![CDATA[telomere-to-telomere genome assembly]]></category>
		<guid isPermaLink="false">https://scienmag.com/scientists-produce-the-most-complete-brown-rat-dna-profile-yet/</guid>

					<description><![CDATA[Scientists have produced the most complete genetic map yet of the brown rat, revealing previously hidden genes, extensive DNA variation, and an unexpected system of sex-chromosome organization. The new genome assembly, led by researchers at UTHealth Houston, offers a powerful reference for studying how genes contribute to heart disease, kidney disease, high blood pressure, stroke, [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Scientists have produced the most complete genetic map yet of the brown rat, revealing previously hidden genes, extensive DNA variation, and an unexpected system of sex-chromosome organization. The new genome assembly, led by researchers at UTHealth Houston, offers a powerful reference for studying how genes contribute to heart disease, kidney disease, high blood pressure, stroke, immune disorders, and other conditions. Because rats are among the most widely used animals in biomedical research, the findings could reshape the way scientists interpret results from preclinical experiments.</p>
<p>Published in <em>Cell Genomics</em>, the study was led by Peter Doris, PhD, director of the Center for Human Genetics at The Brown Foundation Institute of Molecular Medicine within McGovern Medical School at UTHealth Houston. The researchers used advanced long-read DNA sequencing to construct a telomere-to-telomere assembly of the brown rat genome. Unlike earlier genome drafts, which contained gaps and unresolved repetitive regions, the new assembly provides continuous sequences extending from one telomere—the protective DNA structure at a chromosome’s end—to the other.</p>
<p>The completed genome revealed that the rat’s genetic architecture is considerably more complex than previously recognized. The team identified more than 60 genes that had not been accurately captured in earlier reference genomes. Many of these genes lie in regions that are difficult to sequence because they contain repeated or nearly identical DNA segments. Some appear to be involved in immunity and other biological processes, raising the possibility that missing genetic information has contributed to incomplete or misleading interpretations of rat-based disease research.</p>
<p>One of the most surprising discoveries involved the rat’s X and Y chromosomes. In humans and most other mammals, these chromosomes contain a shared segment called the pseudoautosomal region, or PAR. The PAR contains genes present on both the X and Y chromosomes, allowing the two chromosomes to pair during the formation of reproductive cells and to replicate correctly. Although the X and Y chromosomes differ substantially, this shared region acts as a genetic bridge between them.</p>
<p>The researchers found that the brown rat has lost these PAR genes from its sex chromosomes. Instead, the genes have moved to ordinary, non-sex chromosomes. The team also identified newly organized DNA sequences that appear to allow the rat’s X and Y chromosomes to pair in a head-to-tail configuration, rather than the head-to-head arrangement seen in most other mammals. This finding suggests that the mechanics of rat reproduction have evolved along a distinct genetic pathway, despite the animal’s close relevance to human biology.</p>
<p>“Sexual reproduction in the rat can take place, but it’s not taking place in exactly the same way that it is in humans,” Doris said. The unusual chromosome structure would have been difficult to detect without a highly accurate genome assembly, because incomplete reference sequences can obscure rearrangements and make genes appear to be missing, misplaced, or incorrectly duplicated.</p>
<p>The new work also addresses a long-standing problem in genetic disease research. Researchers often compare the genomes of laboratory rats with those of other strains or with disease-associated genetic regions, but missing segments can make it difficult to determine which DNA differences are biologically meaningful. Gene duplications are especially challenging: when two copies are nearly identical, conventional sequencing methods may collapse them into a single sequence. Yet duplicated genes can acquire different functions, allowing one copy to retain an original role while the other becomes specialized.</p>
<p>To capture this diversity, the team assembled eight reference-quality genomes from different brown rat strains. These assemblies were combined into a pangenome—a comprehensive genetic resource that represents variation across multiple individuals rather than treating one genome as the definitive standard. The rat pangenome contains approximately 7% more sequence than the previously available reference genome, revealing genetic regions that would otherwise remain invisible. Scientists can now examine a gene across several rat strains and determine whether its sequence, copy number, or biological function varies between animals.</p>
<p>Such variation may have direct implications for laboratory studies. A gene involved in digestion, for example, may have been duplicated in some rats, with one copy retaining a digestive function while the other becomes involved in immune activity. If researchers use different strains without accounting for these differences, they may obtain conflicting results or fail to reproduce findings. The pangenome provides a framework for identifying these differences before they influence an experiment, potentially improving the reliability of studies that use rats to investigate human disease.</p>
<p>The rat genome consists of 22 chromosome pairs, and the new assembly describes each chromosome in an unbroken sequence. By filling the gaps in the genetic “map,” the study gives researchers a more precise way to locate disease-associated variants, study chromosome evolution, and compare rat biology with human biology. The resource is expected to support future research into cardiovascular and metabolic disease, kidney function, inflammation, immunity, and neurological disorders. Alongside Doris, the study included Yaming Zhu of UTHealth Houston and collaborators from the University of Kentucky, the National Institutes of Health, the University of Louisville, and The Jackson Laboratory.</p>
<p><strong>Subject of Research</strong>: Genetic sequencing, genome assembly, pangenomics, chromosome biology, and biomedical rat research</p>
<p><strong>Article Title</strong>: Telomere-to-telomere genome assembly and a pangenome for the rat</p>
<p><strong>News Publication Date</strong>: 6-Aug-2026</p>
<p><strong>Web References</strong>: <a href="https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00143-6">https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00143-6</a></p>
<p><strong>References</strong>: <em>Cell Genomics</em>, “Telomere-to-telomere genome assembly and a pangenome for the rat”</p>
<p><strong>Image Credits</strong>: Photo by UTHealth Houston; Peter Doris, PhD, director of the Center for Human Genetics at The Brown Foundation Institute of Molecular Medicine within McGovern Medical School at UTHealth Houston.</p>
<p><strong>Keywords</strong>: Brown rat, rat genome, telomere-to-telomere assembly, pangenome, genomics, genetics, sex chromosomes, pseudoautosomal region, gene duplication, disease research, biomedical research, long-read sequencing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">177628</post-id>	</item>
		<item>
		<title>Human Genome Breakthrough Paves Way for Personalized Genomics</title>
		<link>https://scienmag.com/human-genome-breakthrough-paves-way-for-personalized-genomics/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Fri, 07 Aug 2026 04:58:19 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[chromosome-level genome assembly]]></category>
		<category><![CDATA[complex repetitive DNA decoding]]></category>
		<category><![CDATA[diploid human genome sequencing]]></category>
		<category><![CDATA[genome sequencing technology]]></category>
		<category><![CDATA[genomic differences and variations]]></category>
		<category><![CDATA[high-resolution genome sequencing]]></category>
		<category><![CDATA[human genome reconstruction]]></category>
		<category><![CDATA[human genome reference improvements]]></category>
		<category><![CDATA[implications for personalized medicine]]></category>
		<category><![CDATA[maternal and paternal genome differentiation]]></category>
		<category><![CDATA[personalized genomics advancements]]></category>
		<category><![CDATA[telomere-to-telomere genome assembly]]></category>
		<guid isPermaLink="false">https://scienmag.com/human-genome-breakthrough-paves-way-for-personalized-genomics/</guid>

					<description><![CDATA[Scientists have reconstructed the most complete diploid human genome yet produced, creating a high-resolution sequence that contains both copies of every chromosome inherited from an individual’s parents. The achievement, led by researchers from Johns Hopkins University, the National Human Genome Research Institute and the National Institute of Standards and Technology, marks a major advance beyond [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Scientists have reconstructed the most complete diploid human genome yet produced, creating a high-resolution sequence that contains both copies of every chromosome inherited from an individual’s parents. The achievement, led by researchers from Johns Hopkins University, the National Human Genome Research Institute and the National Institute of Standards and Technology, marks a major advance beyond the conventional human reference genome. Rather than identifying a person’s genetic differences by comparing them with an incomplete standard, the new method reconstructs the individual genome itself, including regions that have historically been too repetitive or complex to decode.</p>
<p>The work was carried out by the Telomere-to-Telomere, or T2T, Consortium using HG002, a human genome sample obtained from a living donor and widely used as a reference material by sequencing and diagnostic laboratories. The researchers assembled each chromosome from one telomere, the protective structure at one end, to the other, producing two separate chromosome sets that represent the maternal and paternal genomes. This is technically more difficult than assembling a single genome because the two copies are highly similar but not identical. Computational systems must determine which DNA fragments belong to which parental chromosome while preserving every small difference between them.</p>
<p>The new sequence adds more than 900 million DNA letters that were absent from previous benchmarks and reveals roughly 15% more of the genome than the earlier reference standard. These newly accessible regions include parts of both sex chromosomes, highly repetitive stretches, and sequences containing genes and regulatory elements that may influence disease risk. Such regions have often been excluded from clinical sequencing because standard technologies struggled to read them accurately or because researchers could not determine their correct position within the genome. By resolving these difficult segments, the T2T approach could expose genetic variants that have remained invisible in routine testing.</p>
<p>The advance builds on the consortium’s landmark 2022 completion of the first truly complete human genome sequence. That project filled in approximately the final 8% of a single reference genome, including many repetitive regions and the previously incomplete Y chromosome. The new effort goes further by applying improved sequencing platforms, assembly algorithms and validation methods to a diploid genome. Long-read sequencing technologies were central to the work because they generate DNA fragments thousands or even millions of letters long, allowing researchers to span repetitive sequences that would be broken into ambiguous pieces by older short-read methods.</p>
<p>Accurately assigning genes to each chromosome copy was another essential part of the project. Scientists at Johns Hopkins led by computational biologist Steven Salzberg analyzed the two chromosome sets to identify and annotate their genes, while Michael Schatz’s laboratory contributed to extensive validation of the assembly. Independent checks were used to test whether the reconstructed sequence contained errors, missing segments or incorrectly joined fragments. The result is intended not only as a biological reference but also as a measurement standard for companies developing DNA sequencing instruments, analysis software and clinical diagnostics.</p>
<p>Researchers say the development could change the logic of medical genomics. Current clinical analyses generally search for variants that differ from a standard reference genome. This strategy can perform well when a patient’s DNA resembles the reference, but it becomes less reliable in genomic regions where the reference is incomplete or structurally different. A complete genome assembled for each patient would instead provide an individualized baseline. Genetic analysis could then examine substitutions, insertions, deletions, duplications and larger rearrangements across the entire sequence without automatically discarding regions that do not align well with the traditional reference.</p>
<p>The immediate medical benefit could be improved diagnosis for children and adults with rare genetic disorders. Genome sequencing is already used in such cases, but more than half of patients may still leave testing without a clear molecular explanation. Missing or misread regions can conceal the mutation responsible for disease, particularly when it lies in a repetitive sequence or involves a complex structural change. A complete diploid assembly could help clinicians identify these causes more accurately, potentially ending years of uncertainty for families and guiding treatment, monitoring and reproductive decisions.</p>
<p>The same approach may eventually strengthen predictions for common diseases. Variants in the BRCA1 and BRCA2 genes are already used to estimate breast cancer risk, but researchers believe that many additional risk-associated changes remain undiscovered in difficult-to-sequence portions of the genome. More complete reference data could improve studies of cancer, cardiovascular disease, immune disorders and neuropsychiatric conditions. When combined with genomes from large and diverse populations, these sequences could also support artificial intelligence models trained to recognize disease-related patterns while reducing the bias created by relying on a single, historically limited reference genome.</p>
<p>The consortium estimates that a complete and highly accurate human genome can now be generated for about $5,000, compared with the roughly $5 billion, in current dollars, spent on the Human Genome Project, which concluded in 2003. Although routine whole-genome sequencing still raises questions about privacy, data storage, consent and the interpretation of uncertain findings, the technical barrier is rapidly falling. The researchers envision a future in which a person’s complete genome is sequenced early in life, securely linked to medical records and revisited as scientific knowledge improves. The study is part of a broader package of work in <em>Cell</em> and <em>Cell Genomics</em> that also presents complete or near-complete genomes for macaques, marmosets, zebra finches, rats, voles, horses, donkeys and giraffes, extending the same genomic precision to research on evolution, biodiversity, agriculture and animal health.</p>
<p><strong>Subject of Research</strong>: Complete diploid human genome sequencing and personalized genomics</p>
<p><strong>Article Title</strong>: Complete, high-quality diploid human genome reconstructed from telomere to telomere</p>
<p><strong>Web References</strong>:<br />
<a href="https://engineering.jhu.edu/faculty/adam-phillippy/">https://engineering.jhu.edu/faculty/adam-phillippy/</a><br />
<a href="https://hub.jhu.edu/2022/03/31/johns-hopkins-scientists-first-complete-sequence-human-genome/">https://hub.jhu.edu/2022/03/31/johns-hopkins-scientists-first-complete-sequence-human-genome/</a><br />
<a href="https://www.bme.jhu.edu/people/faculty/steven-l-salzberg/">https://www.bme.jhu.edu/people/faculty/steven-l-salzberg/</a><br />
<a href="https://engineering.jhu.edu/faculty/michael-schatz/">https://engineering.jhu.edu/faculty/michael-schatz/</a></p>
<p><strong>References</strong>:<br />
Cell, DOI: 10.1016/j.cell.2026.06.016</p>
<h4><strong>Keywords</strong></h4>
<p>Human genome sequencing, diploid genome, Telomere-to-Telomere Consortium, personalized genomics, genetic disease diagnosis, long-read sequencing, genomic medicine, structural variants, precision medicine, genome assembly</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">177598</post-id>	</item>
		<item>
		<title>Telomere-to-Telomere Assembly with HERRO Nanopore</title>
		<link>https://scienmag.com/telomere-to-telomere-assembly-with-herro-nanopore/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Mon, 27 Apr 2026 16:52:38 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced haplotype phasing algorithms]]></category>
		<category><![CDATA[complete chromosome-length contiguity]]></category>
		<category><![CDATA[cost-effective genome assembly methods]]></category>
		<category><![CDATA[diploid and polyploid genome sequencing]]></category>
		<category><![CDATA[genomic repetitive region resolution]]></category>
		<category><![CDATA[haplotype-aware genome phasing]]></category>
		<category><![CDATA[HERRO deep learning error correction]]></category>
		<category><![CDATA[high-fidelity long-read sequencing alternatives]]></category>
		<category><![CDATA[ONT Simplex read enhancement]]></category>
		<category><![CDATA[scalable T2T assembly techniques]]></category>
		<category><![CDATA[telomere-to-telomere genome assembly]]></category>
		<category><![CDATA[ultra-long Oxford Nanopore sequencing]]></category>
		<guid isPermaLink="false">https://scienmag.com/telomere-to-telomere-assembly-with-herro-nanopore/</guid>

					<description><![CDATA[In the ever-evolving landscape of genome sequencing, achieving reference-quality, telomere-to-telomere (T2T) phased assemblies has long stood as an aspirational benchmark. These assemblies offer the most complete and accurate genomic maps, capturing the entire length of chromosomes from one telomere to the other. Traditionally, such endeavors have been both technically challenging and financially prohibitive, especially when [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the ever-evolving landscape of genome sequencing, achieving reference-quality, telomere-to-telomere (T2T) phased assemblies has long stood as an aspirational benchmark. These assemblies offer the most complete and accurate genomic maps, capturing the entire length of chromosomes from one telomere to the other. Traditionally, such endeavors have been both technically challenging and financially prohibitive, especially when scaled across diploid and polyploid genomes. The complexity arises chiefly because assembling complete genomes demands integrating multiple types of sequencing reads, each with their own strengths and limitations.</p>
<p>Conventionally, generating T2T assemblies requires the amalgamation of highly accurate long-read technologies, exemplified by PacBio HiFi sequencing or the now-obsolete ONT Duplex reads. These high-fidelity reads provide the precision necessary to correctly resolve genomic sequences. However, to span complex repetitive regions and achieve chromosome-length contiguity, they must be paired with ultra-long Oxford Nanopore Technologies (ONT) Simplex reads. The combination, while effective, drastically escalates costs and the required input of high-quality genomic DNA, thus limiting broad adoption especially in studies involving numerous or large genomes.</p>
<p>Addressing these obstacles, an innovative approach has recently emerged that leverages deep learning to dramatically enhance the utility of ultra-long ONT Simplex reads. Known as HERRO, an acronym for Haplotype-aware ERRor cOrrection, this novel framework harnesses artificial intelligence algorithms to correct sequencing errors while specifically preserving the subtle differences vital for distinguishing haplotypes or repeated genomic regions. HERRO’s design uniquely acknowledges and utilizes informative polymorphic sites, enabling unparalleled error correction fidelity across complex diploid human genomes.</p>
<p>The underlying principle of HERRO lies in its ability to model the intricate sequence variation inherent in diploid genomes. By incorporating a haplotype-aware correction strategy, HERRO maintains true biological differences rather than erroneously smoothing them out, a common pitfall in conventional error correction algorithms. This precision facilitates a read accuracy increase up to 100-fold, setting a new standard for the quality of ONT Simplex reads. Importantly, this leap in accuracy comparably matches—and in some cases rivals—that of more established yet costlier approaches dependent on multiple sequencing platforms.</p>
<p>Integrating HERRO-corrected ONT reads with Verkko, a cutting-edge de novo genome assembler optimized for long reads, researchers have reconstructed up to 32 chromosomes telomere-to-telomere, including both sex chromosomes X and Y. The assemblies generated demonstrate impressive contiguity, consistently achieving NGA50 values exceeding 100 megabases across multiple human genomes. Such milestones signal a profound shift, as they indicate that ultra-long ONT reads, once considered too error-prone for standalone assembly, can now independently underpin near-complete diploid genome assemblies.</p>
<p>Further underscoring its broad applicability, HERRO supports both major ONT Simplex chemistries—R9.4.1 and the newer R10.4.1—ensuring immediate relevance across current sequencing platforms. Additionally, the method generalizes effectively to non-human species, suggesting avenues for revolutionizing genome assembly in a diversity of biological contexts. This cross-species adaptability opens up possibilities for scaling high-quality genomic studies without prohibitive increases in input material or sequencing complexity.</p>
<p>The implications of HERRO’s development extend well beyond cost reduction. By streamlining the sequencing workflow into a single-platform solution reliant on error-corrected ultra-long ONT reads, the process circumvents the logistical and technical hurdles associated with integrating multiple sequencing technologies. This paradigm not only reduces financial barriers but also minimizes necessary sample handling, which is crucial for studies involving rare or precious biological specimens.</p>
<p>On a technical front, HERRO’s architecture employs deep neural network models trained on extensive genomic datasets, enabling nuanced error recognition and correction while preserving true polymorphisms. The innovation lies in balancing error suppression with biological authenticity, a feat achieved by leveraging sequence context and haplotype structure information during correction. This contrasts with earlier error correction strategies that often indiscriminately corrected mismatches, inadvertently erasing meaningful haplotype differences.</p>
<p>The deep learning framework also adapts dynamically to sequencing chemistry properties, accommodating systematic errors characteristic of different ONT flow cell versions. This adaptability enhances robustness and positions HERRO as a future-proof solution able to keep pace with ONT technology’s rapid evolution.</p>
<p>In practice, implementation of HERRO followed by Verkko assembly demonstrated substantial improvements in contiguity and accuracy benchmarks when applied to diverse human samples. The reconstructions unveiled nearly complete chromosomal assemblies with telomere-to-telomere continuity, setting a new standard for diploid genome resolution. Moreover, the approach effectively uncovered and preserved structural variations and repeat expansions, which are often challenging to resolve but critical for comprehensive genomic analyses.</p>
<p>HERRO thus exemplifies how computational innovation, applied judiciously, can unlock latent potential in existing sequencing chemistries. By transforming traditionally high-error ultra-long ONT reads into highly reliable substrates for assembly, the technology paves the way for expanding high-fidelity, full-chromosome assemblies in both research and clinical domains.</p>
<p>Looking forward, the implications of this technique are vast. Equipping labs worldwide with cost-effective means to produce reference-grade genome assemblies democratizes access to genomic information, accelerating discovery in fields ranging from evolutionary biology to personalized medicine. Further, as sequencing continues its rapid trajectory of advancement, tools like HERRO ensure that data quality improvements keep pace, helping to realize the full promise of genomics.</p>
<p>In sum, HERRO’s haplotype-aware error correction combined with ultra-long ONT read assembly heralds a new era in genome sequencing where completeness and accuracy no longer require prohibitive budgets or complex multi-platform strategies. By harnessing deep learning and the unique strengths of nanopore sequencing, this innovation charts a promising path toward comprehensive, accessible, and affordable chromosome-scale genomics.</p>
<hr />
<p>Subject of Research: Telomere-to-telomere genome assembly and error correction of Oxford Nanopore ultra-long Simplex reads using deep learning.</p>
<p>Article Title: Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads.</p>
<p>Article References:<br />
Stanojević, D., Lin, D., Nurk, S. et al. Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads. Nature (2026). https://doi.org/10.1038/s41586-026-10563-y</p>
<p>Image Credits: AI Generated</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">154792</post-id>	</item>
		<item>
		<title>Complex Genetic Variation in Nearly Complete Genomes</title>
		<link>https://scienmag.com/complex-genetic-variation-in-nearly-complete-genomes/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Wed, 23 Jul 2025 18:23:43 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced genomic research techniques]]></category>
		<category><![CDATA[complex structural variants]]></category>
		<category><![CDATA[disease susceptibility and genetics]]></category>
		<category><![CDATA[genomic complexity in human evolution]]></category>
		<category><![CDATA[genomic instability]]></category>
		<category><![CDATA[human genomic variation]]></category>
		<category><![CDATA[long-read genome assemblies]]></category>
		<category><![CDATA[mobile element insertions]]></category>
		<category><![CDATA[repetitive DNA regions in genomes]]></category>
		<category><![CDATA[segmental duplications]]></category>
		<category><![CDATA[structural variant detection methods]]></category>
		<category><![CDATA[telomere-to-telomere genome assembly]]></category>
		<guid isPermaLink="false">https://scienmag.com/complex-genetic-variation-in-nearly-complete-genomes/</guid>

					<description><![CDATA[In a groundbreaking advance, researchers have leveraged the power of long-read genome assemblies to dissect the most complex forms of structural variation within the human genome. These intricate alterations, coined complex structural variants (CSVs), represent singular genomic events composed of simpler structural variants that extend across multiple repair junctions, often embedded within highly repetitive DNA [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking advance, researchers have leveraged the power of long-read genome assemblies to dissect the most complex forms of structural variation within the human genome. These intricate alterations, coined complex structural variants (CSVs), represent singular genomic events composed of simpler structural variants that extend across multiple repair junctions, often embedded within highly repetitive DNA regions. This novel approach opens an unprecedented window into the hidden landscape of genomic complexity, with profound implications for understanding human evolution, disease susceptibility, and genomic instability.</p>
<p>The challenge in identifying CSVs has long been their tendency to arise within genomic regions laden with segmental duplications (SDs) and mobile element insertions (MEIs). These repetitive sequences notoriously confound traditional sequencing and mapping efforts, masking the true architecture of these variants. By upgrading the existing PAV tool, the research team enabled a heightened sensitivity in capturing CSVs embedded in these large, complex repeats. Applying this enhanced method against the telomere-to-telomere assembled CHM13 reference genome, an average of 72 CSVs were detected per genome, revealing a rich spectrum of 1,247 distinct CSV events with 128 unique complex reference signatures across human populations.</p>
<p>Further interrogation revealed that a substantial fraction of these CSVs embodies local sequence duplications and inversions — approximately 27% exhibited duplications while 38% contained inversions. Intriguingly, many of these variants are orchestrated through mechanisms involving SDs which mediate elaborate architectures, such as INVDUP-INV-DEL, DEL-INV-DEL, and INVDUP-INV-INVDUP. These configurations combine deletions, inversions, and duplications in a complex interplay of genomic rearrangements. One remarkable example highlights CSVs involving the NOTCH2NL and NBPF gene families, loci intrinsically tied to the expansion of the human brain and its evolutionary trajectory.</p>
<p>Previously intractable to resolution through conventional methods such as optical mapping, these CSVs now reveal at least three distinct haplotypes: a reference haplotype with a 13.7% allele frequency, a 930-kilobase inversion-deletion variant affecting NBPF8 and deleting NOTCH2NLR and NBPF26 found at 35.9% frequency, and a 513-kilobase variant involving a distal template switch that replaces NBPF8 with NBPF9, seen in over half of sampled genomes. These findings not only underscore the variability of human haplotypes but also provide precise molecular characterization of loci previously obscured by genomic complexity.</p>
<p>The study&#8217;s scope extends beyond CSVs to structurally challenging gene regions implicated in disease. One prime example is the SMN locus, comprising SMN1 and SMN2 gene copies, central to spinal muscular atrophy pathogenesis and therapeutic targeting. These genes reside within an approximately 1.5 megabase segmental duplication hotspot, historically resistant to full sequence resolution. By successfully assembling and validating 101 complete haplotypes, the team achieved comprehensive characterization of SMN1/2 copy number and structure, alongside related genes such as SERF1A/B, NAIP, and GTF2H2/C.</p>
<p>Intriguingly, nearly half of these haplotypes maintain exactly two copies of SMN1/2 and its associated gene cluster members, reflecting a conserved genomic architecture. However, deviations in this copy number highlight the landscape of structural genomic diversity with potential clinical ramifications. Comparative analysis with short-read genotyping tools Parascopy and SMNCopyNumberCaller affirmed the accuracy of the long-read assembly-derived copy number calls, ensuring reliability. Moreover, findings revealed rare haplotypes lacking SMN1, potentially representing genomic configurations predisposing individuals to disease risk via mechanisms such as interlocus gene conversion.</p>
<p>Expanding the inquiry to other complex, multi-copy genes, the team delved into the amylase gene locus on chromosome 1, which features genes AMY1A, AMY1B, AMY1C, AMY2A, and AMY2B. This locus spans over 200 kilobases and exhibits high structural variability critical to dietary adaptation and metabolic phenotypes. Analysis of 65 fully resolved genome assemblies yielded 39 distinct amylase haplotypes, covering a significant majority of the population&#8217;s haplotype diversity. A remarkable breadth of haplotype lengths—from ~111 kb to over 580 kb—reflects evolutionary expansion and contraction events shaping this locus.</p>
<p>Among these diverse haplotypes, four dominate in prevalence, collectively making up over half of all observed haplotypes, affirming a blend of common and rare structural configurations in modern humans. The study is particularly notable for fully resolving the largest known amylase haplotype, containing eleven tandem AMY1 gene copies, a locus previously only partially characterized via optical genome mapping. The resolution of such intricate haplotypes illuminates the evolutionary complexity and functional implications embedded within seemingly inscrutable genomic regions.</p>
<p>The collective impact of this research lies in bridging the gap between reference-quality genome assemblies and the nuanced, individualized structural variation defining human genomic diversity. By applying refined computational tools to ultra-long read data, the study surfaces a wealth of complex variation concealed within difficult genomic terrain. The resulting catalogs not only expand our understanding of structural genomic diversity but also provide invaluable resources for future studies dissecting genotype-to-phenotype relationships, disease mechanisms, and evolutionary history.</p>
<p>Furthermore, the delineation of distinct haplotype structures across populations paves the way for population-scale assessments of genetic risk, improving the resolution of genetic diagnostics and personalized medicine. For disorders like spinal muscular atrophy, where gene copy number and arrangement are crucial for prognosis and therapy, such precision genomics can revolutionize patient care. Moreover, the ability to phase and characterize complex loci across thousands of genomes opens novel vistas in evolutionary biology, functional genomics, and bioinformatics.</p>
<p>By overcoming the limitations of short-read sequencing and the ambiguity of optical mapping, this integrative approach represents a paradigm shift in human genomics. It underscores the critical need for high-fidelity, contiguous genome assemblies capturing the full spectrum of structural variation. These insights reveal the inherent plasticity of the human genome and its capacity for generating complexity that shapes both health and disease.</p>
<p>As the field marches toward comprehensive population-scale long-read sequencing efforts, tools like the enhanced PAV and tailored assembly pipelines will be indispensable. They empower researchers to not only detect but precisely delineate the architecture of complex variants. Ultimately, this progress enriches the foundational knowledge of human genetic variation and drives forward the promise of truly personalized genomics.</p>
<hr />
<p><strong>Subject of Research</strong>: Complex structural variants and genomic diversity in near-complete human genome assemblies</p>
<p><strong>Article Title</strong>: Complex genetic variation in nearly complete human genomes</p>
<p><strong>Article References</strong>:<br />
Logsdon, G.A., Ebert, P., Audano, P.A. <em>et al.</em> Complex genetic variation in nearly complete human genomes. <em>Nature</em> (2025). <a href="https://doi.org/10.1038/s41586-025-09140-6">https://doi.org/10.1038/s41586-025-09140-6</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">58924</post-id>	</item>
	</channel>
</rss>
