<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>protein evolution &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/protein-evolution/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 13 Sep 2026 01:50:55 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>protein evolution &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI model maps the entire protein universe in a single view</title>
		<link>https://scienmag.com/new-ai-model-maps-the-entire-protein-universe-in-a-single-view/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 01:50:55 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI-driven understanding of cellular functions]]></category>
		<category><![CDATA[amino acid sequence]]></category>
		<category><![CDATA[amino acid sequence and 3D structure integration]]></category>
		<category><![CDATA[artificial intelligence in biochemistry]]></category>
		<category><![CDATA[bioinformatics tools for protein research]]></category>
		<category><![CDATA[CATH]]></category>
		<category><![CDATA[CLSS]]></category>
		<category><![CDATA[CLSS model for protein analysis]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[deep learning for protein analysis]]></category>
		<category><![CDATA[ECOD]]></category>
		<category><![CDATA[evolution of protein families]]></category>
		<category><![CDATA[evolutionary biochemistry]]></category>
		<category><![CDATA[Institute of Science Tokyo]]></category>
		<category><![CDATA[interdisciplinary approaches in molecular biology]]></category>
		<category><![CDATA[mapping biological diversity]]></category>
		<category><![CDATA[protein classification]]></category>
		<category><![CDATA[protein embeddings]]></category>
		<category><![CDATA[protein evolution]]></category>
		<category><![CDATA[protein folding and molecular tasks]]></category>
		<category><![CDATA[protein language model]]></category>
		<category><![CDATA[protein structure]]></category>
		<category><![CDATA[protein structure prediction]]></category>
		<category><![CDATA[protein universe mapping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200584</guid>

					<description><![CDATA[An international research team has developed CLSS, a protein language model that unites amino acid sequence and structural information into a single map of protein space, revealing evolutionary relationships across billions of years.]]></description>
										<content:encoded><![CDATA[<p>Every living cell depends on thousands of distinct protein families, each folding into precise three-dimensional shapes to carry out the molecular tasks that sustain life. Where all of this diversity came from, and how the different families relate to one another across billions of years of evolution, remains one of the deepest open questions in biochemistry. An international team of researchers, including the Earth-Life Science Institute (ELSI) at Institute of Science Tokyo, has now unveiled a new artificial intelligence tool that brings scientists closer to an answer by fusing the two fundamental languages of proteins—amino acid sequence and three-dimensional structure—into a single, unified representation. The work, published in Proceedings of the National Academy of Sciences, promises to transform how researchers explore the vast and largely unmapped protein universe.</p>
<p>The study was led by Professor Rachel Kolodny and PhD candidate Guy Yanai of the University of Haifa, together with Professor Nir Ben-Tal and graduate student Gabriel Axel of Tel Aviv University, and Specially Appointed Associate Professor Liam M. Longo of ELSI. Kolodny also spent five months as a visiting researcher at ELSI, developing methods to analyze the new model. Their creation, dubbed CLSS for Contrastive Learning Sequence-Structure, is a protein language model designed to overcome a stubborn problem that has limited previous computational approaches: the awkward relationship between what a protein&#8217;s sequence says and what its structure actually does.</p>
<p>Scientists have long organized proteins into hierarchical groups based on relatedness, much like the genus and species categories biologists use to classify organisms. These curated systems, such as the widely used ECOD and CATH databases, distill decades of expert knowledge. But with artificial intelligence now capable of generating &#8217;embeddings&#8217;—numerical representations in which proteins with similar properties receive nearby coordinates, like postal codes on a map—researchers can visualize relationships across millions of proteins at once, producing what the team calls a protein world map. The catch is that sequence and structure do not map neatly onto each other. Unrelated sequences can fold into similar shapes, while even identical sequences can sometimes adopt wildly different structures.</p>
<p>Most existing protein language models treat sequence and structure as separate worlds, processing one or the other independently. Even hybrid models that incorporate both kinds of data rarely place the sequence and the structure of the same protein at the same location on a global map, leaving researchers with two conflicting atlases of protein space. CLSS was engineered specifically to resolve this discordance. Using a machine learning strategy known as contrastive learning, the model is trained on pairs of protein sequences and their corresponding structures, learning to pull matching sequence-structure pairs together in the embedding space while pushing unrelated pairs apart.</p>
<p>The result is a single shared map in which a protein occupies essentially the same location whether the model is given its sequence or its structure. When benchmarked against other state-of-the-art protein language models, CLSS succeeded in producing a cohesive unified representation, something its predecessors could not achieve. Remarkably, the model&#8217;s maps closely reproduced the relationships recorded in the expert-curated ECOD and CATH classification systems, even though those classifications were never shown to the model during training. In direct classification tests, CLSS also performed strongly, demonstrating that merging sequence and structure information yields genuinely more informative protein representations.</p>
<p>Perhaps the most exciting feature of CLSS is its ability to handle fragments. Most protein language models require a complete sequence or structure to generate a meaningful embedding, but CLSS showed that short sequence fragments can in many cases be positioned meaningfully alongside full-length proteins and structures. This capability matters enormously for evolutionary studies, because small pieces of proteins have been repeatedly reused and rearranged throughout the history of life. Some fragments may even have served as the primordial building blocks from which the earliest protein domains were assembled, meaning that similar fragments appearing in otherwise unrelated proteins can hint at ancient evolutionary connections.</p>
<p>The maps produced by CLSS also revealed sweeping patterns across protein space that were previously difficult to see. When the researchers overlaid biological properties onto the maps, proteins associated with organic cofactors turned out to cluster in particular regions, while metal-binding proteins were scattered more broadly. Such patterns illustrate how global protein maps can serve not only as classification tools but as instruments for exploring the interplay between sequence, structure, function, and deep evolutionary history, potentially exposing large-scale patterns invisible to conventional pairwise comparison methods.</p>
<p>&#8216;This gives us a way to look at the protein universe through sequence and structure at the same time, rather than treating them as separate worlds,&#8217; said Longo. &#8216;What is particularly exciting for us is the possibility of using these maps to uncover large-scale evolutionary patterns that are difficult to recognise using conventional approaches.&#8217; The team ultimately envisions unified sequence-structure representations opening new frontiers in database searches, protein engineering, and the reconstruction of evolutionary trajectories—offering a fresh window onto how the staggering diversity of proteins found in life today emerged over nearly four billion years of evolution.</p>
<p><strong>Subject of Research:</strong> A contrastive-learning protein language model that unifies protein sequence and structure representations to map the protein universe</p>
<p><strong>Article Title:</strong> Uniting sequence and structure to map the protein universe</p>
<p><strong>Article References:</strong> Uniting sequence and structure to map the protein universe. (n.d.). <a href="https://www.eurekalert.org/news-releases/1142950" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> protein language model, CLSS, protein evolution, contrastive learning, protein structure, amino acid sequence, ECOD, CATH, protein embeddings, evolutionary biochemistry, protein classification, Institute of Science Tokyo</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200584</post-id>	</item>
		<item>
		<title>Hidden Sequence Motif Reveals How Natural Enzymes Harness Unusual Redox Cofactors</title>
		<link>https://scienmag.com/hidden-sequence-motif-reveals-how-natural-enzymes-harness-unusual-redox-cofactors/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 12:48:40 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[biocatalysis]]></category>
		<category><![CDATA[biochemistry of redox-active enzyme cofactors]]></category>
		<category><![CDATA[biosynthetic gene clusters]]></category>
		<category><![CDATA[biotechnological applications of enzyme cofactors]]></category>
		<category><![CDATA[deazaflavin F420]]></category>
		<category><![CDATA[enzyme cofactor discovery and characterization]]></category>
		<category><![CDATA[enzyme diversity beyond canonical cofactors]]></category>
		<category><![CDATA[enzyme engineering]]></category>
		<category><![CDATA[enzyme sequence motif]]></category>
		<category><![CDATA[enzymes]]></category>
		<category><![CDATA[expanding enzymatic chemical repertoire]]></category>
		<category><![CDATA[flavin]]></category>
		<category><![CDATA[genome annotation]]></category>
		<category><![CDATA[hidden enzyme functional motifs]]></category>
		<category><![CDATA[implications for drug discovery and enzyme engineering]]></category>
		<category><![CDATA[natural enzyme electron transfer mechanisms]]></category>
		<category><![CDATA[natural product biosynthesis]]></category>
		<category><![CDATA[Nature Chemical Biology]]></category>
		<category><![CDATA[noncanonical redox cofactors in enzymes]]></category>
		<category><![CDATA[novel enzyme catalysis pathways]]></category>
		<category><![CDATA[protein evolution]]></category>
		<category><![CDATA[redox cofactors]]></category>
		<category><![CDATA[role of cofactors in cellular respiration and biosynthesis]]></category>
		<category><![CDATA[sequence motif]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194447</guid>

					<description><![CDATA[Researchers have identified a conserved sequence motif that enables many natural enzymes to use noncanonical redox cofactors, expanding the known chemical capabilities of biology.]]></description>
										<content:encoded><![CDATA[<p>Enzymes are the workhorses of cellular chemistry, and much of their power comes from small helper molecules known as cofactors. For decades, biochemists have catalogued a relatively short list of canonical redox cofactors—flavins, nicotinamides, hemes and iron–sulfur clusters among them—that carry out the vast majority of electron-transfer reactions in living systems. Yet a growing body of evidence suggests that nature&#8217;s catalytic toolkit is far richer than the textbooks imply. A new study published in Nature Chemical Biology reveals that a previously overlooked sequence motif allows many natural enzymes to employ noncanonical redox cofactors, expanding the known chemical repertoire of biology and opening new avenues for biotechnology and drug discovery.</p>
<p>Redox cofactors are the molecular batteries of the cell. They accept and donate electrons in the tightly choreographed reactions that underpin respiration, photosynthesis, biosynthesis and detoxification. The canonical cofactors—molecules such as flavin adenine dinucleotide (FAD), flavin mononucleotide (FMN), nicotinamide adenine dinucleotide (NAD) and nicotinamide adenine dinucleotide phosphate (NADP)—are so widespread that their presence in an enzyme active site is often assumed rather than demonstrated. But over the past several years, researchers have identified a series of modified and entirely distinct cofactors: prenylated flavins such as flavin adenine dinucleotide modified with a prenyl group, deazaflavins like F420, quinone-derived cofactors such as topaquinone and tryptophan tryptophylquinone, and metal-organic species that defy easy classification. These noncanonical cofactors enable chemistries that standard flavins and nicotinamides cannot easily achieve, including hydride transfers at unusual redox potentials, radical-mediated rearrangements and C–C bond formations that would be difficult with conventional catalysis.</p>
<p>The central puzzle addressed in the new work is one of recognition and assembly. If an enzyme uses a noncanonical cofactor, how does the protein know to bind that cofactor rather than its more abundant canonical cousin? And how can bioinformaticians predict, from sequence alone, which of the millions of uncharacterized proteins in genomic databases depend on these exotic helpers? The answer, according to the study, lies in a short, recurring sequence motif—a conserved stretch of amino acids that acts as a molecular postcode, directing the enzyme&#8217;s cofactor-binding pocket toward noncanonical chemistry.</p>
<p>Sequence motifs have long served as the workhorses of computational biology. Short conserved patterns, such as the P-loop that binds nucleotide phosphates or the zinc-finger motifs that coordinate metal ions in DNA-binding proteins, allow researchers to assign function to proteins that have never been isolated in a laboratory. The newly identified motif performs a similar role for redox cofactor selection. By scanning families of flavin-dependent enzymes and comparing those known to use standard FAD or FMN with the smaller subset confirmed to use modified or alternative cofactors, the researchers identified a conserved pattern of residues that appears with striking regularity in the noncanonical group and is conspicuously absent from the canonical one. Mutational experiments confirmed that altering these residues in a noncanonical enzyme abolished its ability to accommodate the alternative cofactor, while introducing the motif into a canonical scaffold shifted its cofactor preference—a result that establishes the motif as a genuine determinant of cofactor identity rather than a coincidental correlation.</p>
<p>The implications of this finding extend well beyond the specific enzyme families examined in the study. Genomic surveys suggest that proteins carrying the motif are distributed across a remarkable range of organisms, from soil-dwelling actinobacteria—long recognized as prolific producers of bioactive natural products—to human-associated microbes and even some archaeal lineages. In many of these organisms, the motif-bearing enzymes cluster within biosynthetic gene clusters, the compact genomic neighborhoods that encode the assembly lines for antibiotics, antitumor agents and other specialized metabolites. This genomic context hints at a widespread and previously underappreciated role for noncanonical redox chemistry in natural product biosynthesis, suggesting that many of the structurally exotic metabolites isolated from microbes over the past half-century may owe their existence to enzymes quietly using cofactors that standard annotation pipelines would never flag.</p>
<p>One of the most exciting consequences of the work is predictive. Armed with the motif, researchers can now interrogate sequence databases with a simple pattern search and retrieve a curated list of candidate enzymes likely to use noncanonical cofactors. This transforms what has historically been a slow, serendipitous process—discover a strange metabolite, purify the enzyme responsible, and only then realize the cofactor is unusual—into a rational, hypothesis-driven workflow. Biochemistry can then be targeted at the most promising candidates, prioritizing enzymes from gene clusters associated with medicinally relevant compound classes. In an era when the rate of genome sequencing vastly outpaces the rate of experimental characterization, tools that convert sequence information into functional predictions are among the most valuable commodities in the life sciences.</p>
<p>The discovery also carries significant weight for synthetic biology and enzyme engineering. Noncanonical cofactors often possess redox potentials and reactivity profiles that canonical cofactors cannot match. F420, for example, the deazaflavin cofactor best known from methanogenic archaea, mediates hydride transfer reactions at potentials inaccessible to NAD and NADP, and engineered F420-dependent enzymes have already been explored for the degradation of persistent pollutants and the production of pharmaceutical intermediates. Prenylated flavins, meanwhile, catalyze photochemical reactions that ordinary flavins cannot, and their light-driven chemistry is being harnessed in optogenetic tools and photocatalytic cascades. A sequence-level handle on cofactor selection means that protein engineers can now rationally swap cofactor identity in designed enzymes, effectively reprogramming the electrochemical capabilities of a catalytic scaffold without altering its overall fold. This could accelerate the design of biocatalysts for green chemistry, where replacing metal catalysts and harsh reagents with enzyme-based alternatives is a major industrial goal.</p>
<p>From an evolutionary standpoint, the findings raise fascinating questions about how and why biology expanded its redox cofactor repertoire in the first place. The canonical cofactors are ancient, likely predating the last universal common ancestor, and their chemistry is deeply woven into core metabolism. Noncanonical cofactors, by contrast, appear to have arisen as evolutionary innovations in specific ecological and metabolic contexts—perhaps to exploit new redox niches, to escape the thermodynamic constraints of shared metabolic pools, or to protect specialized pathways from cross-talk with housekeeping chemistry. The presence of a dedicated sequence motif suggests that cofactor innovation was accompanied by co-evolution of the protein binding environment, producing a heritable, recognizable signature that could be propagated across enzyme families through duplication and divergence. In this sense, the motif is a fossil record of chemical innovation, preserving in amino acid sequence the memory of evolutionary experiments in electron transfer.</p>
<p>The study also serves as a cautionary tale for genome annotation. Most automated pipelines assign enzyme function by homology, and a protein that resembles a flavin-dependent monooxygenase is typically annotated as such, regardless of which cofactor it actually employs. If a substantial fraction of these enzymes in fact use noncanonical cofactors, then large swaths of existing functional annotations may be subtly or substantially wrong, with consequences for metabolic modeling, pathway reconstruction and the interpretation of gene-expression data. The motif provides a corrective lens, allowing annotators to flag proteins whose cofactor assignments deserve experimental scrutiny. As the authors and commentators in the field note, the lesson is broader: the most abundant cofactors are not necessarily the only ones, and assumptions baked into databases can obscure entire layers of biochemical diversity.</p>
<p>Looking forward, the identification of this sequence motif is likely to catalyze a wave of discovery across several fronts. Experimentalists will purify and characterize motif-bearing enzymes from diverse organisms, likely uncovering new cofactor structures and new reaction types. Computational biologists will refine the motif definition, searching for related patterns that govern the use of other exotic cofactors, and integrating these signals into machine-learning models of enzyme function. Structural biologists will determine how the motif residues reshape the cofactor-binding pocket at atomic resolution, providing design principles for engineered catalysts. And natural products chemists will revisit orphan biosynthetic gene clusters with fresh eyes, suspecting that many of the unexplained transformations encoded within them depend on redox chemistry that no one thought to look for. What began as a search for a short string of amino acids has ended with a map pointing toward a vast, unexplored territory of enzyme chemistry—one that has been hiding in plain sight within the genomes of organisms all around us, waiting only for the right pattern to reveal it.</p>
<p><strong>Subject of Research:</strong> A conserved sequence motif that enables natural enzymes to use noncanonical redox cofactors</p>
<p><strong>Article Title:</strong> A sequence motif enables widespread use of noncanonical redox cofactors in natural enzymes</p>
<p><strong>Article References:</strong> Saleh, S., Hsu, N.-H., Luu, E., Martin, V. C., Ng, H. J. C., Black, W. B., Zhang, S., Kim, J.-K., Sankaran, B., Tran, A. H. T., Hayes, R. L., Siegel, J. B., Qiao, F., &amp; Li, H. (2026). A sequence motif enables widespread use of noncanonical redox cofactors in natural enzymes. <em>Nature Chemical Biology</em>. <a href="https://doi.org/10.1038/s41589-026-02315-w" rel="noopener noreferrer">https://doi.org/10.1038/s41589-026-02315-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41589-026-02315-w" rel="noopener noreferrer">10.1038/s41589-026-02315-w</a></p>
<p><strong>Keywords:</strong> redox cofactors, sequence motif, enzymes, flavin, natural product biosynthesis, genome annotation, enzyme engineering, deazaflavin F420, biocatalysis, protein evolution, biosynthetic gene clusters, Nature Chemical Biology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194447</post-id>	</item>
	</channel>
</rss>
