<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>protein classification &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/protein-classification/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 26 Sep 2026 21:35:49 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>protein classification &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Tool Fuses Language Models and Evolutionary Profiles to Classify CRISPR Cas Proteins</title>
		<link>https://scienmag.com/ai-tool-fuses-language-models-and-evolutionary-profiles-to-classify-crispr-cas-proteins/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 21:35:49 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[accuracy improvement in CRISPR protein identification]]></category>
		<category><![CDATA[bacterial immune system genomics]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[bioinformatics in microbiology]]></category>
		<category><![CDATA[BMC Bioinformatics]]></category>
		<category><![CDATA[Cas proteins]]></category>
		<category><![CDATA[computational methods for protein classification]]></category>
		<category><![CDATA[CRISPR protein classification]]></category>
		<category><![CDATA[CRISPR-Cas system detection tools]]></category>
		<category><![CDATA[CRISPR-Cas system diversity]]></category>
		<category><![CDATA[CRISPR/Cas]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[development of bioinformatics classifiers]]></category>
		<category><![CDATA[ESM1b]]></category>
		<category><![CDATA[evolutionary profile-based protein analysis]]></category>
		<category><![CDATA[feature fusion]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in microbiology]]></category>
		<category><![CDATA[prokaryotic immune system research]]></category>
		<category><![CDATA[protein classification]]></category>
		<category><![CDATA[protein language models]]></category>
		<category><![CDATA[protein sequence reading techniques]]></category>
		<category><![CDATA[PSSM]]></category>
		<category><![CDATA[Random Forest]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216477</guid>

					<description><![CDATA[A new machine learning method called PrePssmCas fuses protein language model embeddings with evolutionary PSSM features to classify CRISPR-Cas proteins with record accuracy.]]></description>
										<content:encoded><![CDATA[<p>A new computational method promises to sharpen one of the most fiddly tasks in modern microbiology: telling apart the many kinds of proteins that power CRISPR immune systems in bacteria and archaea. In a study published in BMC Bioinformatics, researchers at Northwest A&amp;F University in Yangling, China, describe PrePssmCas, a machine learning classifier that fuses two very different ways of reading a protein sequence and, in doing so, outperforms seven existing tools for classifying CRISPR-Cas systems. On an independent validation set, the method reached an accuracy of 97.98 percent and a Matthews correlation coefficient of 0.962, improving on the previous best method, CRISPRCasStack, by 3.91 percentage points in accuracy and 9.60 percentage points in MCC.</p>
<p>The stakes are higher than they might first appear. CRISPR and the Cas proteins that accompany it form the adaptive immune systems of prokaryotes, the vast group of single-celled organisms that dominate life on Earth. These systems store fragments of invading viruses and plasmids in the famous clustered regularly interspaced short palindromic repeats, then use them as molecular wanted posters to recognize and destroy foreign genetic material. Because different classes of Cas proteins perform different jobs within these systems, identifying which Cas proteins are present in a genome offers a route to classifying the whole system. The trouble is that direct experimental identification of CRISPR-Cas systems remains difficult, so computational classification carries much of the load, and its accuracy matters.</p>
<p>The core insight behind PrePssmCas is that no single numerical representation of a protein captures everything a classifier needs to know. The team therefore extracted two complementary families of features from Cas protein sequences. The first came from pre-trained protein language models, the deep neural networks that have transformed computational biology in recent years by learning statistical patterns from millions of sequences. The second came from position-specific scoring matrices, or PSSMs, a much older technique that encodes the evolutionary conservation of each position in a sequence by comparing it against related proteins in a database.</p>
<p>The language model side of the comparison was ambitious. The researchers systematically evaluated five pre-trained models: Prot_BERT, ALBERT, ProtXLNet, ProtT5 and ESM1b. Each of these models reads a protein sequence like a sentence and produces a high-dimensional vector, known as an embedding, for every amino acid in it. These embeddings encode information about structure, function and evolutionary context that the models absorbed during training on vast protein databases. On the PSSM side, the team compared six different variants, methods that compress the raw scoring matrix into fixed-dimensional descriptors that a classifier can consume. The raw PSSM itself has dimensions that scale with protein length, which makes it unwieldy for machine learning, so the compression step is where much of the engineering happens.</p>
<p>After systematic comparison, a clear winner emerged: embeddings from ESM1b, the protein language model developed by Meta&#8217;s fundamental AI research group, combined with a PSSM variant called RPM-PSSM gave the most comprehensive representation of Cas proteins. The researchers also incorporated an attention-based aggregation strategy, a mechanism that lets the model learn which parts of a sequence deserve the most weight when building its final representation. Attention has become a cornerstone of modern deep learning, and its use here reflects a broader trend of importing techniques from natural language processing into protein analysis.</p>
<p>Fusing the two feature families produced a large pool of candidate descriptors, and not all of them carried useful signal. To distill the mixture, the team applied random forest-based feature selection, a technique that uses an ensemble of decision trees to score each feature by how much it contributes to accurate classification. The result was a compact 143-dimensional feature vector, comprising 87 features drawn from the pre-trained language model embeddings and 56 features from the RPM-PSSM representation. That reduction, from potentially thousands of dimensions down to 143, is what allows the final classifier to generalize rather than memorize, focusing on the descriptors that genuinely distinguish one type of Cas protein from another.</p>
<p>The performance gains were substantial. On the independent validation set, the selected features achieved 97.98 percent accuracy and an MCC of 0.962. The Matthews correlation coefficient is widely regarded as a more honest metric than raw accuracy because it accounts for all four cells of the confusion matrix and remains informative even when classes are imbalanced, as they typically are in biological datasets. An MCC above 0.95 indicates near-perfect agreement between predictions and ground truth. Crucially, the improvements were measured against CRISPRCasStack, itself a recent and strong baseline, meaning the gains are not merely the result of comparing against weak predecessors.</p>
<p>The benchmarking went further. PrePssmCas was evaluated against seven existing methods: HMMCAS, CASPredict, CRISPRone, CRISPRCasFinder, CRISPRCasTyper, CRISPRloci and CRISPRCasStack. These tools represent a range of strategies, from hidden Markov models that capture profile signatures of protein families to more recent machine learning pipelines. On the independent validation set, PrePssmCas outperformed all of them, a result the authors attribute to the complementary nature of the fused features. Language model embeddings excel at capturing deep semantic patterns learned across the protein universe, while PSSM-derived features anchor the predictions in the specific evolutionary history of each sequence, and the combination appears to cover blind spots that neither approach addresses alone.</p>
<p>The work also contributes to a lively debate in computational biology about whether large pre-trained language models have made classical sequence analysis techniques obsolete. The answer from this study is a qualified no. Although the ESM1b embeddings contributed the larger share of the final feature vector, more than half of the selected dimensions came from the language model, the 56 RPM-PSSM features that survived selection clearly added discriminative power that the embeddings alone did not provide. Evolutionary profiles, computed by aligning a query sequence against a database of homologs with tools like BLAST, encode information about which positions tolerate mutation, and that information is not always fully captured by a language model trained on sequences without explicit alignment data. The fusion framework suggests that the most powerful classifiers of the near future will be hybrids rather than replacements.</p>
<p>The authors are candid about the limits of their achievement. They note that the generalization of PrePssmCas to large-scale datasets and to remotely homologous sequences, proteins whose similarity to known Cas proteins is so distant that standard detection methods struggle, remains to be established. This is a familiar caveat in protein classification, where models trained on curated datasets can falter when confronted with the messy diversity of real genomic data, particularly the rapidly evolving catalog of CRISPR-Cas variants being discovered through metagenomics. Still, the researchers frame their contribution as both a practical tool for classifying CRISPR-Cas systems and a promising feature-fusion framework that could support future Cas protein discovery, potentially helping biologists spot novel defense systems hiding in the growing torrent of microbial genome sequences. The study, led by Ningyi Zhang with corresponding author Qianqian Shi, was conducted without dedicated funding and relied on the High-Performance Computing Center of Northwest A&amp;F University, the publicly available ESM1b model and BLAST software, and the Gold Standard dataset for Cas protein analysis maintained by the research community.</p>
<p><strong>Subject of Research:</strong> Machine learning classification of CRISPR-Cas proteins using fused pre-trained language model and PSSM features</p>
<p><strong>Article Title:</strong> PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification</p>
<p><strong>Article References:</strong> Zhang, N., Zhao, Y., Luo, C., Peng, Z., &amp; Shi, Q. (2026). PrePssmCas: fusing pre-trained features and PSSM features for Cas protein classification. <em>BMC Bioinformatics</em>. <a href="https://doi.org/10.1186/s12859-026-06569-y" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06569-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06569-y" rel="noopener noreferrer">10.1186/s12859-026-06569-y</a></p>
<p><strong>Keywords:</strong> CRISPR-Cas, Cas proteins, protein classification, machine learning, ESM1b, protein language models, PSSM, feature fusion, bioinformatics, BMC Bioinformatics, random forest, deep learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216477</post-id>	</item>
		<item>
		<title>New AI model maps the entire protein universe in a single view</title>
		<link>https://scienmag.com/new-ai-model-maps-the-entire-protein-universe-in-a-single-view/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 01:50:55 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI-driven understanding of cellular functions]]></category>
		<category><![CDATA[amino acid sequence]]></category>
		<category><![CDATA[amino acid sequence and 3D structure integration]]></category>
		<category><![CDATA[artificial intelligence in biochemistry]]></category>
		<category><![CDATA[bioinformatics tools for protein research]]></category>
		<category><![CDATA[CATH]]></category>
		<category><![CDATA[CLSS]]></category>
		<category><![CDATA[CLSS model for protein analysis]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[deep learning for protein analysis]]></category>
		<category><![CDATA[ECOD]]></category>
		<category><![CDATA[evolution of protein families]]></category>
		<category><![CDATA[evolutionary biochemistry]]></category>
		<category><![CDATA[Institute of Science Tokyo]]></category>
		<category><![CDATA[interdisciplinary approaches in molecular biology]]></category>
		<category><![CDATA[mapping biological diversity]]></category>
		<category><![CDATA[protein classification]]></category>
		<category><![CDATA[protein embeddings]]></category>
		<category><![CDATA[protein evolution]]></category>
		<category><![CDATA[protein folding and molecular tasks]]></category>
		<category><![CDATA[protein language model]]></category>
		<category><![CDATA[protein structure]]></category>
		<category><![CDATA[protein structure prediction]]></category>
		<category><![CDATA[protein universe mapping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200584</guid>

					<description><![CDATA[An international research team has developed CLSS, a protein language model that unites amino acid sequence and structural information into a single map of protein space, revealing evolutionary relationships across billions of years.]]></description>
										<content:encoded><![CDATA[<p>Every living cell depends on thousands of distinct protein families, each folding into precise three-dimensional shapes to carry out the molecular tasks that sustain life. Where all of this diversity came from, and how the different families relate to one another across billions of years of evolution, remains one of the deepest open questions in biochemistry. An international team of researchers, including the Earth-Life Science Institute (ELSI) at Institute of Science Tokyo, has now unveiled a new artificial intelligence tool that brings scientists closer to an answer by fusing the two fundamental languages of proteins—amino acid sequence and three-dimensional structure—into a single, unified representation. The work, published in Proceedings of the National Academy of Sciences, promises to transform how researchers explore the vast and largely unmapped protein universe.</p>
<p>The study was led by Professor Rachel Kolodny and PhD candidate Guy Yanai of the University of Haifa, together with Professor Nir Ben-Tal and graduate student Gabriel Axel of Tel Aviv University, and Specially Appointed Associate Professor Liam M. Longo of ELSI. Kolodny also spent five months as a visiting researcher at ELSI, developing methods to analyze the new model. Their creation, dubbed CLSS for Contrastive Learning Sequence-Structure, is a protein language model designed to overcome a stubborn problem that has limited previous computational approaches: the awkward relationship between what a protein&#8217;s sequence says and what its structure actually does.</p>
<p>Scientists have long organized proteins into hierarchical groups based on relatedness, much like the genus and species categories biologists use to classify organisms. These curated systems, such as the widely used ECOD and CATH databases, distill decades of expert knowledge. But with artificial intelligence now capable of generating &#8217;embeddings&#8217;—numerical representations in which proteins with similar properties receive nearby coordinates, like postal codes on a map—researchers can visualize relationships across millions of proteins at once, producing what the team calls a protein world map. The catch is that sequence and structure do not map neatly onto each other. Unrelated sequences can fold into similar shapes, while even identical sequences can sometimes adopt wildly different structures.</p>
<p>Most existing protein language models treat sequence and structure as separate worlds, processing one or the other independently. Even hybrid models that incorporate both kinds of data rarely place the sequence and the structure of the same protein at the same location on a global map, leaving researchers with two conflicting atlases of protein space. CLSS was engineered specifically to resolve this discordance. Using a machine learning strategy known as contrastive learning, the model is trained on pairs of protein sequences and their corresponding structures, learning to pull matching sequence-structure pairs together in the embedding space while pushing unrelated pairs apart.</p>
<p>The result is a single shared map in which a protein occupies essentially the same location whether the model is given its sequence or its structure. When benchmarked against other state-of-the-art protein language models, CLSS succeeded in producing a cohesive unified representation, something its predecessors could not achieve. Remarkably, the model&#8217;s maps closely reproduced the relationships recorded in the expert-curated ECOD and CATH classification systems, even though those classifications were never shown to the model during training. In direct classification tests, CLSS also performed strongly, demonstrating that merging sequence and structure information yields genuinely more informative protein representations.</p>
<p>Perhaps the most exciting feature of CLSS is its ability to handle fragments. Most protein language models require a complete sequence or structure to generate a meaningful embedding, but CLSS showed that short sequence fragments can in many cases be positioned meaningfully alongside full-length proteins and structures. This capability matters enormously for evolutionary studies, because small pieces of proteins have been repeatedly reused and rearranged throughout the history of life. Some fragments may even have served as the primordial building blocks from which the earliest protein domains were assembled, meaning that similar fragments appearing in otherwise unrelated proteins can hint at ancient evolutionary connections.</p>
<p>The maps produced by CLSS also revealed sweeping patterns across protein space that were previously difficult to see. When the researchers overlaid biological properties onto the maps, proteins associated with organic cofactors turned out to cluster in particular regions, while metal-binding proteins were scattered more broadly. Such patterns illustrate how global protein maps can serve not only as classification tools but as instruments for exploring the interplay between sequence, structure, function, and deep evolutionary history, potentially exposing large-scale patterns invisible to conventional pairwise comparison methods.</p>
<p>&#8216;This gives us a way to look at the protein universe through sequence and structure at the same time, rather than treating them as separate worlds,&#8217; said Longo. &#8216;What is particularly exciting for us is the possibility of using these maps to uncover large-scale evolutionary patterns that are difficult to recognise using conventional approaches.&#8217; The team ultimately envisions unified sequence-structure representations opening new frontiers in database searches, protein engineering, and the reconstruction of evolutionary trajectories—offering a fresh window onto how the staggering diversity of proteins found in life today emerged over nearly four billion years of evolution.</p>
<p><strong>Subject of Research:</strong> A contrastive-learning protein language model that unifies protein sequence and structure representations to map the protein universe</p>
<p><strong>Article Title:</strong> Uniting sequence and structure to map the protein universe</p>
<p><strong>Article References:</strong> Uniting sequence and structure to map the protein universe. (n.d.). <a href="https://www.eurekalert.org/news-releases/1142950" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> protein language model, CLSS, protein evolution, contrastive learning, protein structure, amino acid sequence, ECOD, CATH, protein embeddings, evolutionary biochemistry, protein classification, Institute of Science Tokyo</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200584</post-id>	</item>
	</channel>
</rss>
