<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>ProtBERT-BFD &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/protbert-bfd/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 14:02:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>ProtBERT-BFD &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Built for Green Life: DeepGreenGO Reads Plant Proteins Where Other Models Fail</title>
		<link>https://scienmag.com/ai-built-for-green-life-deepgreengo-reads-plant-proteins-where-other-models-fail/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 14:02:34 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[bioinformatics for plant biology]]></category>
		<category><![CDATA[crop gene annotation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for plant genomics]]></category>
		<category><![CDATA[DeepGreenGO]]></category>
		<category><![CDATA[Gene Ontology]]></category>
		<category><![CDATA[Graph Neural Networks]]></category>
		<category><![CDATA[green life computational biology]]></category>
		<category><![CDATA[machine learning in agriculture]]></category>
		<category><![CDATA[plant biology]]></category>
		<category><![CDATA[plant gene function prediction]]></category>
		<category><![CDATA[plant genome annotation tools]]></category>
		<category><![CDATA[plant protein function prediction]]></category>
		<category><![CDATA[plant protein function research]]></category>
		<category><![CDATA[plant proteome analysis]]></category>
		<category><![CDATA[ProtBERT-BFD]]></category>
		<category><![CDATA[protein function prediction]]></category>
		<category><![CDATA[rice]]></category>
		<category><![CDATA[rice seed development genes]]></category>
		<category><![CDATA[seed development]]></category>
		<category><![CDATA[species-specific protein function models]]></category>
		<category><![CDATA[sustainable agriculture]]></category>
		<category><![CDATA[Viridiplantae]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223178</guid>

					<description><![CDATA[A plant-specific deep learning model called DeepGreenGO combines protein language model embeddings with graph neural networks to predict Gene Ontology functions and has identified 372 candidate rice proteins linked to seed development.]]></description>
										<content:encoded><![CDATA[<p>Plants have long been the quiet underdogs of computational biology. While human, mouse, and yeast proteins have accumulated decades of experimentally verified functions, the vast green branch of life known as Viridiplantae remains sparsely annotated, especially in crop species and non-model organisms that matter most for food security. A team of researchers from the University of Colombo and the Sri Lanka Institute of Information Technology has now tackled this gap head-on with DeepGreenGO, a deep learning model described in BMC Bioinformatics that is trained exclusively on plant proteins and designed to predict what plant genes actually do. Rather than borrowing models built on mixed-species datasets where plant proteins are a small minority, the team built a taxonomically focused framework from the ground up, and then demonstrated its power by hunting for genes that govern seed development in rice, one of the world&#8217;s most important staple crops.</p>
<p>The core problem the researchers confronted is a familiar one in genomics: experimental annotation is slow and expensive, while genome sequencing is fast and cheap. When a new plant genome is sequenced, thousands of proteins are catalogued with little more than a name and a sequence. Computational tools can transfer functional labels from well-studied relatives, but this homology-based approach breaks down for proteins that have no close characterized counterparts, which is precisely the situation in many orphan crops. General-purpose deep learning function predictors, trained on multi-taxon datasets, tend to underrepresent experimentally annotated plant proteins, so their celebrated performance on animals and microbes does not necessarily transfer to the green lineage. DeepGreenGO was built to close that transferability gap by learning the patterns of plant protein function in a dataset where plants are not a footnote but the entire curriculum.</p>
<p>Technically, the model combines two complementary views of a protein. The first comes from ProtBERT-BFD, a protein language model trained on the enormous Big Fantastic Database, which converts each amino acid residue into a rich numerical embedding that encodes evolutionary and biochemical context learned from billions of sequences. The second view is structural: the researchers derived contact maps from protein structures, essentially matrices recording which residues lie close together in three-dimensional space. These contact maps define a graph in which residues are nodes and physical proximity defines the edges. This graph representation is then processed by two types of graph neural network layers in sequence: graph convolutional network layers, which aggregate information from neighboring residues, and GATv2 layers, a more expressive form of graph attention that learns to weight which neighbors matter most for each residue. An attention pooling step then compresses the residue-level representations into a single protein-level vector that feeds a multilabel classifier.</p>
<p>The output of that classifier is a set of Gene Ontology terms, the standardized vocabulary that biologists use to describe protein functions across three branches: molecular function, biological process, and cellular component. Because Gene Ontology is hierarchical, predicting a function is inherently a multilabel problem, and DeepGreenGO was trained to predict terms specific to the ontology it was trained on. The training corpus was a carefully curated Viridiplantae dataset of 7,534 experimentally annotated protein structures drawn from the Protein Data Bank, cross-referenced through the SIFTS resource to link structures to sequences and taxonomy. Crucially, the team split these proteins into training, validation, and test sets of 6,026, 754, and 754 PDB chains respectively using sequence-similarity-aware clustering, a methodological safeguard that prevents near-identical proteins from appearing in both training and test data and inflating performance estimates.</p>
<p>The benchmarking results reveal where the plant-specific approach pays off. DeepGreenGO was evaluated against sequence homology-based tools and against other recent sequence- and structure-based deep learning methods, and it showed its strongest performance in the biological process ontology, arguably the hardest and most informative of the three GO branches. There it achieved the highest protein-centric maximum F1 score, known as Fmax, of 0.227, and the lowest Smin score of 19.73, a metric that penalizes both missed and spurious predictions in a semantically aware way. Perhaps most striking was its showing on the information-content weighted area under the precision-recall curve, where it scored highest among all compared methods for biological process. That metric rewards models for correctly predicting rare, informative annotations rather than playing it safe with broad, shallow terms, suggesting that DeepGreenGO is genuinely learning specific plant biology rather than recycling generic predictions.</p>
<p>Ablation experiments, in which components of the model are systematically removed to measure their contribution, added an honest and instructive note to the study. The pretrained ProtBERT-BFD sequence embeddings turned out to provide most of the predictive signal, confirming the now well-established principle that self-supervised protein language models capture a remarkable amount of functional information on their own. The structural contact maps, processed through the graph neural network stack, contributed more modestly, and the authors identify improving the integration of structural information as a clear opportunity for future work. This kind of transparency matters in a field where architectural novelty is often claimed as the source of gains; here the data show that the language model backbone is central, while the graph machinery offers a framework that can be sharpened as structural prediction quality improves across the plant kingdom.</p>
<p>To demonstrate that the model is more than a benchmark winner, the researchers deployed DeepGreenGO across the entire proteome of rice, Oryza sativa, screening all 43,649 rice proteins for functions related to seed development. The scan flagged 372 candidate proteins associated with this agriculturally critical process. The team then validated the biological plausibility of these predictions through two independent lines of evidence. Functional enrichment analysis showed that the predicted proteins were collectively linked to biological processes known to operate during seed formation, and transcriptomic analysis revealed that many of them display elevated expression in ovary, embryo, and endosperm tissues, the anatomical arenas where seed development unfolds. For a purely computational prediction to align with tissue-level expression patterns is a meaningful consistency check, and it suggests the candidate list is a credible starting point for experimental follow-up.</p>
<p>The implications reach well beyond rice. Seed development genes influence yield, grain quality, and stress resilience, traits at the heart of breeding programs aimed at sustainable agriculture in a changing climate. A tool that can prioritize which of tens of thousands of uncharacterized proteins deserve scarce laboratory resources is a practical accelerator for crop science, particularly for orphan crops and underexplored species where homology-based annotation fails. Because DeepGreenGO is taxonomically focused, its predictions are also more trustworthy for proteins with limited detectable similarity to the training set, exactly the proteins that dominate non-model plant genomes. The framework could in principle be retrained or extended as new plant structures are experimentally solved and as AlphaFold-style predicted structures expand the structural universe available to graph-based methods.</p>
<p>There are also broader lessons here for the design of biological AI systems. The study is a case study in domain-specific curation: rather than scraping the largest possible dataset, the team built a plant-only corpus, controlled for sequence similarity during data splitting, and evaluated with metrics that distinguish informative predictions from easy ones. The modest absolute Fmax values in the biological process ontology are a candid reminder that protein function prediction remains genuinely hard, and that headline numbers from multi-taxon benchmarks should not be assumed to hold in specialized domains. The work also highlights the growing role of research groups outside the traditional centers of AI and genomics; the Sri Lankan team, working with donated workstations acknowledged from the Colombo University Faculty of Science Alumni Association in North America and a faculty high-performance computing facility, produced a competitive model without industrial-scale resources.</p>
<p>DeepGreenGO arrives at a moment when the intersection of protein language models, graph learning, and structural biology is reshaping what is computable in the life sciences. Its strongest advantages, in informative biological process annotations and in proteins far from any characterized relative, point toward a future where plant genomes are no longer annotated by proxy from animals and fungi but on their own terms. The rice seed development results offer a template for how such models should be deployed: predict broadly, then triangulate with enrichment and expression data before anyone commits a greenhouse season to a candidate gene. As plant structural datasets grow and structural integration improves, models of this kind could become standard instruments in the crop biologist&#8217;s toolkit, quietly converting sequence catalogs into testable biology, one green protein at a time.</p>
<p><strong>Subject of Research:</strong> Graph neural network-based deep learning for plant-specific protein function prediction</p>
<p><strong>Article Title:</strong> DeepGreenGO: a graph neural network-based deep learning model for plant-specific protein function prediction</p>
<p><strong>Article References:</strong> Sridharan, G., Weththasinghe, S. A., Sridharan, A., Upeka, W. M. M., Abeywardhana, D. L., &amp; Fernando, P. C. (2026). DeepGreenGO: a graph neural network-based deep learning model for plant-specific protein function prediction. <em>BMC Bioinformatics</em>. <a href="https://doi.org/10.1186/s12859-026-06677-9" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06677-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06677-9" rel="noopener noreferrer">10.1186/s12859-026-06677-9</a></p>
<p><strong>Keywords:</strong> DeepGreenGO, protein function prediction, graph neural networks, plant biology, Gene Ontology, ProtBERT-BFD, rice, seed development, bioinformatics, deep learning, Viridiplantae, sustainable agriculture</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223178</post-id>	</item>
	</channel>
</rss>
