<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>bioinformatics tools for protein research &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/bioinformatics-tools-for-protein-research/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 13 Sep 2026 01:50:55 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>bioinformatics tools for protein research &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI model maps the entire protein universe in a single view</title>
		<link>https://scienmag.com/new-ai-model-maps-the-entire-protein-universe-in-a-single-view/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 01:50:55 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI-driven understanding of cellular functions]]></category>
		<category><![CDATA[amino acid sequence]]></category>
		<category><![CDATA[amino acid sequence and 3D structure integration]]></category>
		<category><![CDATA[artificial intelligence in biochemistry]]></category>
		<category><![CDATA[bioinformatics tools for protein research]]></category>
		<category><![CDATA[CATH]]></category>
		<category><![CDATA[CLSS]]></category>
		<category><![CDATA[CLSS model for protein analysis]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[deep learning for protein analysis]]></category>
		<category><![CDATA[ECOD]]></category>
		<category><![CDATA[evolution of protein families]]></category>
		<category><![CDATA[evolutionary biochemistry]]></category>
		<category><![CDATA[Institute of Science Tokyo]]></category>
		<category><![CDATA[interdisciplinary approaches in molecular biology]]></category>
		<category><![CDATA[mapping biological diversity]]></category>
		<category><![CDATA[protein classification]]></category>
		<category><![CDATA[protein embeddings]]></category>
		<category><![CDATA[protein evolution]]></category>
		<category><![CDATA[protein folding and molecular tasks]]></category>
		<category><![CDATA[protein language model]]></category>
		<category><![CDATA[protein structure]]></category>
		<category><![CDATA[protein structure prediction]]></category>
		<category><![CDATA[protein universe mapping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200584</guid>

					<description><![CDATA[An international research team has developed CLSS, a protein language model that unites amino acid sequence and structural information into a single map of protein space, revealing evolutionary relationships across billions of years.]]></description>
										<content:encoded><![CDATA[<p>Every living cell depends on thousands of distinct protein families, each folding into precise three-dimensional shapes to carry out the molecular tasks that sustain life. Where all of this diversity came from, and how the different families relate to one another across billions of years of evolution, remains one of the deepest open questions in biochemistry. An international team of researchers, including the Earth-Life Science Institute (ELSI) at Institute of Science Tokyo, has now unveiled a new artificial intelligence tool that brings scientists closer to an answer by fusing the two fundamental languages of proteins—amino acid sequence and three-dimensional structure—into a single, unified representation. The work, published in Proceedings of the National Academy of Sciences, promises to transform how researchers explore the vast and largely unmapped protein universe.</p>
<p>The study was led by Professor Rachel Kolodny and PhD candidate Guy Yanai of the University of Haifa, together with Professor Nir Ben-Tal and graduate student Gabriel Axel of Tel Aviv University, and Specially Appointed Associate Professor Liam M. Longo of ELSI. Kolodny also spent five months as a visiting researcher at ELSI, developing methods to analyze the new model. Their creation, dubbed CLSS for Contrastive Learning Sequence-Structure, is a protein language model designed to overcome a stubborn problem that has limited previous computational approaches: the awkward relationship between what a protein&#8217;s sequence says and what its structure actually does.</p>
<p>Scientists have long organized proteins into hierarchical groups based on relatedness, much like the genus and species categories biologists use to classify organisms. These curated systems, such as the widely used ECOD and CATH databases, distill decades of expert knowledge. But with artificial intelligence now capable of generating &#8217;embeddings&#8217;—numerical representations in which proteins with similar properties receive nearby coordinates, like postal codes on a map—researchers can visualize relationships across millions of proteins at once, producing what the team calls a protein world map. The catch is that sequence and structure do not map neatly onto each other. Unrelated sequences can fold into similar shapes, while even identical sequences can sometimes adopt wildly different structures.</p>
<p>Most existing protein language models treat sequence and structure as separate worlds, processing one or the other independently. Even hybrid models that incorporate both kinds of data rarely place the sequence and the structure of the same protein at the same location on a global map, leaving researchers with two conflicting atlases of protein space. CLSS was engineered specifically to resolve this discordance. Using a machine learning strategy known as contrastive learning, the model is trained on pairs of protein sequences and their corresponding structures, learning to pull matching sequence-structure pairs together in the embedding space while pushing unrelated pairs apart.</p>
<p>The result is a single shared map in which a protein occupies essentially the same location whether the model is given its sequence or its structure. When benchmarked against other state-of-the-art protein language models, CLSS succeeded in producing a cohesive unified representation, something its predecessors could not achieve. Remarkably, the model&#8217;s maps closely reproduced the relationships recorded in the expert-curated ECOD and CATH classification systems, even though those classifications were never shown to the model during training. In direct classification tests, CLSS also performed strongly, demonstrating that merging sequence and structure information yields genuinely more informative protein representations.</p>
<p>Perhaps the most exciting feature of CLSS is its ability to handle fragments. Most protein language models require a complete sequence or structure to generate a meaningful embedding, but CLSS showed that short sequence fragments can in many cases be positioned meaningfully alongside full-length proteins and structures. This capability matters enormously for evolutionary studies, because small pieces of proteins have been repeatedly reused and rearranged throughout the history of life. Some fragments may even have served as the primordial building blocks from which the earliest protein domains were assembled, meaning that similar fragments appearing in otherwise unrelated proteins can hint at ancient evolutionary connections.</p>
<p>The maps produced by CLSS also revealed sweeping patterns across protein space that were previously difficult to see. When the researchers overlaid biological properties onto the maps, proteins associated with organic cofactors turned out to cluster in particular regions, while metal-binding proteins were scattered more broadly. Such patterns illustrate how global protein maps can serve not only as classification tools but as instruments for exploring the interplay between sequence, structure, function, and deep evolutionary history, potentially exposing large-scale patterns invisible to conventional pairwise comparison methods.</p>
<p>&#8216;This gives us a way to look at the protein universe through sequence and structure at the same time, rather than treating them as separate worlds,&#8217; said Longo. &#8216;What is particularly exciting for us is the possibility of using these maps to uncover large-scale evolutionary patterns that are difficult to recognise using conventional approaches.&#8217; The team ultimately envisions unified sequence-structure representations opening new frontiers in database searches, protein engineering, and the reconstruction of evolutionary trajectories—offering a fresh window onto how the staggering diversity of proteins found in life today emerged over nearly four billion years of evolution.</p>
<p><strong>Subject of Research:</strong> A contrastive-learning protein language model that unifies protein sequence and structure representations to map the protein universe</p>
<p><strong>Article Title:</strong> Uniting sequence and structure to map the protein universe</p>
<p><strong>Article References:</strong> Uniting sequence and structure to map the protein universe. (n.d.). <a href="https://www.eurekalert.org/news-releases/1142950" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> protein language model, CLSS, protein evolution, contrastive learning, protein structure, amino acid sequence, ECOD, CATH, protein embeddings, evolutionary biochemistry, protein classification, Institute of Science Tokyo</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200584</post-id>	</item>
	</channel>
</rss>
