<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>structured biological knowledge databases &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/structured-biological-knowledge-databases/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 03:23:16 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>structured biological knowledge databases &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On</title>
		<link>https://scienmag.com/ai-framework-reads-the-literature-to-map-protein-modifications-it-was-never-trained-on/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 03:23:16 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[advancements in proteomics research]]></category>
		<category><![CDATA[AI-based literature mining]]></category>
		<category><![CDATA[automated annotation of protein modifications]]></category>
		<category><![CDATA[bioinformatics tools for PTMs]]></category>
		<category><![CDATA[BiomedBERT]]></category>
		<category><![CDATA[BMC Bioinformatics]]></category>
		<category><![CDATA[discovery of novel PTMs]]></category>
		<category><![CDATA[GenPTM software for PTM mapping]]></category>
		<category><![CDATA[handling large-scale scientific literature]]></category>
		<category><![CDATA[information extraction]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in biological research]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[phosphorylation]]></category>
		<category><![CDATA[post-translational modification]]></category>
		<category><![CDATA[protein post-translational modifications]]></category>
		<category><![CDATA[protein sites]]></category>
		<category><![CDATA[Proteomics]]></category>
		<category><![CDATA[proteomics data curation]]></category>
		<category><![CDATA[PTM knowledge extraction]]></category>
		<category><![CDATA[PubMed]]></category>
		<category><![CDATA[structured biological knowledge databases]]></category>
		<category><![CDATA[text mining]]></category>
		<category><![CDATA[ubiquitination]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225386</guid>

					<description><![CDATA[Researchers have developed GenPTM, a text-mining framework that extracts protein post-translational modification information from scientific abstracts and generalizes to modification types it was never trained on.]]></description>
										<content:encoded><![CDATA[<p>Proteins are not static machines. After they are assembled inside a cell, most of them are chemically decorated with small molecular tags that switch them on, shut them down, send them to different compartments, or mark them for destruction. These decorations, known as post-translational modifications, or PTMs, are among the most important regulatory mechanisms in biology, and they are described constantly in the scientific literature. Yet capturing that knowledge in a usable, structured form has remained a stubborn bottleneck for the field of proteomics. A team of researchers at the University of Delaware and Georgetown University Medical Center now reports a solution that could change how PTM knowledge is harvested from millions of published papers. Their tool, called GenPTM, is described in an open-access paper published in BMC Bioinformatics on 24 September 2026.</p>
<p>The scale of the problem is easy to underestimate. Curated databases do exist that catalogue which proteins carry which modifications and at which amino acid positions, and these resources are indispensable for experimentalists planning experiments or interpreting mass-spectrometry data. But most of those databases cover only a limited number of PTM types, and they are not updated regularly. Meanwhile, the volume of PTM-related findings buried in PubMed abstracts keeps growing year after year. The result is a widening information gap: knowledge that has already been discovered and published sits locked inside unstructured text, inaccessible to computational analysis, while the manual curation effort needed to unlock it falls further behind. Building a separate, dedicated text-mining system for every one of the hundreds of known PTM types is simply not feasible.</p>
<p>GenPTM attacks this problem with a deliberately generalizable design. The central insight behind the framework is that, despite the bewildering chemical diversity of PTMs, the sentences scientists use to describe them are remarkably similar in structure. A paper might report that a protein is ubiquitinated at a lysine residue, phosphorylated on a serine, or methylated at an arginine, but the underlying linguistic pattern, in which a protein, a modification event, and a site are linked together, repeats across modification types. Rather than teaching a model the vocabulary of each modification separately, the Delaware team replaces PTM-specific modification names and chemical group mentions with generic placeholders. This unified text representation strategy forces the model to focus on the shared textual patterns that express modification events, regardless of which particular modification is being discussed.</p>
<p>On top of this representation, the researchers fine-tuned a classifier built on BiomedBERT, a language model pre-trained on biomedical text. The classifier&#8217;s job is deceptively simple: given a candidate protein or a candidate amino acid site in a sentence, decide whether that protein or site is genuinely modified in the context described. This is a harder judgment than it sounds, because scientific prose is full of negations, speculations, and references to other studies. A sentence saying that a site was not found to be phosphorylated, or that phosphorylation at a position is suspected but unconfirmed, must be distinguished from a clean report of an observed modification. The fine-tuned BiomedBERT classifier makes that determination for each candidate, and a post-processing module then assembles the individual decisions into final predictions: modified proteins, modified sites, or complete protein-site pairs.</p>
<p>The training and evaluation strategy is what makes the claim of generalizability convincing. The model was trained on data covering five major PTM types, including well-studied modifications such as ubiquitination and phosphorylation, which dominate the literature and the curated databases. The real test, however, came from evaluation on eight additional PTM types that the model had never seen during training. These included PTMs that are mentioned only rarely in scientific articles, such as citrullination and AMPylation, precisely the kinds of modifications for which dedicated curation resources are weakest and the need for automated extraction is greatest. A system that only works on abundant, well-documented modifications would offer little advantage over existing tools; GenPTM was designed to work on the long tail.</p>
<p>The results reported in the paper are striking. Across all PTM types tested, including the eight held-out modifications, GenPTM achieved F1-scores ranging from 92 percent to 96 percent across three different evaluation categories. The F1-score is a standard metric that balances precision, the fraction of predicted modifications that are correct, against recall, the fraction of true modifications that are found. Scores in the mid-nineties for modification types the model was never explicitly trained on indicate that the placeholder-based representation genuinely transfers knowledge about how modification events are expressed in text. In practical terms, a curator or database builder could point GenPTM at a PTM type with little or no dedicated training data and still expect reliable extraction of modified proteins and their sites from PubMed abstracts.</p>
<p>The implications for proteomics research extend beyond convenience. PTM-agnostic information extraction means that newly discovered or obscure modification types can be surveyed systematically as soon as reports appear, without waiting for a specialized tool or a funded curation project. It also means that existing PTM databases could be updated far more frequently, with automated systems flagging candidate modifications from the newest literature for expert review. Because the framework operates on abstracts, the most accessible layer of the literature, it can scale to the full breadth of PubMed rather than being restricted to a handful of model organisms or heavily studied protein families. The authors position the work as a viable solution for automated PTM knowledge discovery, and the architecture suggests a path toward text-mining systems that generalize across related biomedical relation-extraction problems as well.</p>
<p>The project reflects a substantial collaborative effort. The work was carried out by Shovan Bhowmik, Chuming Chen, Cathy Wu, and K. Vijay-Shanker at the University of Delaware, together with Karen Ross of Georgetown University Medical Center. The team acknowledges the valuable expertise in data curation contributed by Dr. Cecilia Arighi of the University of Delaware, as well as the dedicated curation efforts of Amos Nyabuti and Winnie Aketch, master&#8217;s students at Delaware&#8217;s Center for Bioinformatics and Computational Biology, who helped develop the training and testing corpora on which the framework depends. High-quality annotated corpora are the quiet foundation of every successful biomedical language model, and the paper is a reminder that advances in artificial intelligence for biology rest on meticulous human annotation work.</p>
<p>Financial support came from multiple National Institutes of Health grants, including R35GM141873, U54GM104941, and P20GM103446, along with 1R24GM146616-01 and, from the NIH Office of Strategic Coordination and the Common Fund, 1OT2OD032092. The article is published open access under a Creative Commons Attribution 4.0 license, meaning that any research group, database curator, or tool developer can read, reuse, and build upon the work without restriction. The paper was received on 30 April 2026, accepted on 14 September 2026, and published on 24 September 2026, carrying the DOI 10.1186/s12859-026-06663-1.</p>
<p>For a field drowning in its own literature, GenPTM offers something rarer than another incremental benchmark improvement: a demonstration that a single, thoughtfully designed system can serve the entire spectrum of post-translational modifications, from the phosphorylation events catalogued in thousands of papers to the citrullination and AMPylation events scattered across a comparative handful. If the framework&#8217;s generalization holds as it is applied more broadly, the tedious bottleneck between publication and curated knowledge could begin to close, letting the proteins&#8217; chemical vocabulary be read as fast as scientists can write it.</p>
<p><strong>Subject of Research:</strong> Automated information extraction of protein post-translational modifications from scientific literature using a generalizable natural language processing framework</p>
<p><strong>Article Title:</strong> GenPTM: a generalizable framework for protein post-translational modification information extraction from the scientific literature</p>
<p><strong>Article References:</strong> Bhowmik, S., Ross, K., Chen, C., Wu, C., &amp; Vijay-Shanker, K. (2026). GenPTM: a generalizable framework for protein post-translational modification information extraction from the scientific literature. <em>BMC Bioinformatics</em>. <a href="https://doi.org/10.1186/s12859-026-06663-1" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06663-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06663-1" rel="noopener noreferrer">10.1186/s12859-026-06663-1</a></p>
<p><strong>Keywords:</strong> post-translational modification, information extraction, text mining, BiomedBERT, phosphorylation, ubiquitination, proteomics, PubMed, natural language processing, protein sites, BMC Bioinformatics, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225386</post-id>	</item>
	</channel>
</rss>
