Friday, October 2, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On

October 2, 2026
in Biology
Drew Townsend
By Drew Townsend Scienmag Editorial Profile - Cell Biology
Reading Time: 5 mins read
0
AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On

AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On

AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Proteins are not static machines. After they are assembled inside a cell, most of them are chemically decorated with small molecular tags that switch them on, shut them down, send them to different compartments, or mark them for destruction. These decorations, known as post-translational modifications, or PTMs, are among the most important regulatory mechanisms in biology, and they are described constantly in the scientific literature. Yet capturing that knowledge in a usable, structured form has remained a stubborn bottleneck for the field of proteomics. A team of researchers at the University of Delaware and Georgetown University Medical Center now reports a solution that could change how PTM knowledge is harvested from millions of published papers. Their tool, called GenPTM, is described in an open-access paper published in BMC Bioinformatics on 24 September 2026.

The scale of the problem is easy to underestimate. Curated databases do exist that catalogue which proteins carry which modifications and at which amino acid positions, and these resources are indispensable for experimentalists planning experiments or interpreting mass-spectrometry data. But most of those databases cover only a limited number of PTM types, and they are not updated regularly. Meanwhile, the volume of PTM-related findings buried in PubMed abstracts keeps growing year after year. The result is a widening information gap: knowledge that has already been discovered and published sits locked inside unstructured text, inaccessible to computational analysis, while the manual curation effort needed to unlock it falls further behind. Building a separate, dedicated text-mining system for every one of the hundreds of known PTM types is simply not feasible.

GenPTM attacks this problem with a deliberately generalizable design. The central insight behind the framework is that, despite the bewildering chemical diversity of PTMs, the sentences scientists use to describe them are remarkably similar in structure. A paper might report that a protein is ubiquitinated at a lysine residue, phosphorylated on a serine, or methylated at an arginine, but the underlying linguistic pattern, in which a protein, a modification event, and a site are linked together, repeats across modification types. Rather than teaching a model the vocabulary of each modification separately, the Delaware team replaces PTM-specific modification names and chemical group mentions with generic placeholders. This unified text representation strategy forces the model to focus on the shared textual patterns that express modification events, regardless of which particular modification is being discussed.

On top of this representation, the researchers fine-tuned a classifier built on BiomedBERT, a language model pre-trained on biomedical text. The classifier’s job is deceptively simple: given a candidate protein or a candidate amino acid site in a sentence, decide whether that protein or site is genuinely modified in the context described. This is a harder judgment than it sounds, because scientific prose is full of negations, speculations, and references to other studies. A sentence saying that a site was not found to be phosphorylated, or that phosphorylation at a position is suspected but unconfirmed, must be distinguished from a clean report of an observed modification. The fine-tuned BiomedBERT classifier makes that determination for each candidate, and a post-processing module then assembles the individual decisions into final predictions: modified proteins, modified sites, or complete protein-site pairs.

The training and evaluation strategy is what makes the claim of generalizability convincing. The model was trained on data covering five major PTM types, including well-studied modifications such as ubiquitination and phosphorylation, which dominate the literature and the curated databases. The real test, however, came from evaluation on eight additional PTM types that the model had never seen during training. These included PTMs that are mentioned only rarely in scientific articles, such as citrullination and AMPylation, precisely the kinds of modifications for which dedicated curation resources are weakest and the need for automated extraction is greatest. A system that only works on abundant, well-documented modifications would offer little advantage over existing tools; GenPTM was designed to work on the long tail.

The results reported in the paper are striking. Across all PTM types tested, including the eight held-out modifications, GenPTM achieved F1-scores ranging from 92 percent to 96 percent across three different evaluation categories. The F1-score is a standard metric that balances precision, the fraction of predicted modifications that are correct, against recall, the fraction of true modifications that are found. Scores in the mid-nineties for modification types the model was never explicitly trained on indicate that the placeholder-based representation genuinely transfers knowledge about how modification events are expressed in text. In practical terms, a curator or database builder could point GenPTM at a PTM type with little or no dedicated training data and still expect reliable extraction of modified proteins and their sites from PubMed abstracts.

The implications for proteomics research extend beyond convenience. PTM-agnostic information extraction means that newly discovered or obscure modification types can be surveyed systematically as soon as reports appear, without waiting for a specialized tool or a funded curation project. It also means that existing PTM databases could be updated far more frequently, with automated systems flagging candidate modifications from the newest literature for expert review. Because the framework operates on abstracts, the most accessible layer of the literature, it can scale to the full breadth of PubMed rather than being restricted to a handful of model organisms or heavily studied protein families. The authors position the work as a viable solution for automated PTM knowledge discovery, and the architecture suggests a path toward text-mining systems that generalize across related biomedical relation-extraction problems as well.

The project reflects a substantial collaborative effort. The work was carried out by Shovan Bhowmik, Chuming Chen, Cathy Wu, and K. Vijay-Shanker at the University of Delaware, together with Karen Ross of Georgetown University Medical Center. The team acknowledges the valuable expertise in data curation contributed by Dr. Cecilia Arighi of the University of Delaware, as well as the dedicated curation efforts of Amos Nyabuti and Winnie Aketch, master’s students at Delaware’s Center for Bioinformatics and Computational Biology, who helped develop the training and testing corpora on which the framework depends. High-quality annotated corpora are the quiet foundation of every successful biomedical language model, and the paper is a reminder that advances in artificial intelligence for biology rest on meticulous human annotation work.

Financial support came from multiple National Institutes of Health grants, including R35GM141873, U54GM104941, and P20GM103446, along with 1R24GM146616-01 and, from the NIH Office of Strategic Coordination and the Common Fund, 1OT2OD032092. The article is published open access under a Creative Commons Attribution 4.0 license, meaning that any research group, database curator, or tool developer can read, reuse, and build upon the work without restriction. The paper was received on 30 April 2026, accepted on 14 September 2026, and published on 24 September 2026, carrying the DOI 10.1186/s12859-026-06663-1.

For a field drowning in its own literature, GenPTM offers something rarer than another incremental benchmark improvement: a demonstration that a single, thoughtfully designed system can serve the entire spectrum of post-translational modifications, from the phosphorylation events catalogued in thousands of papers to the citrullination and AMPylation events scattered across a comparative handful. If the framework’s generalization holds as it is applied more broadly, the tedious bottleneck between publication and curated knowledge could begin to close, letting the proteins’ chemical vocabulary be read as fast as scientists can write it.

Subject of Research: Automated information extraction of protein post-translational modifications from scientific literature using a generalizable natural language processing framework

Article Title: GenPTM: a generalizable framework for protein post-translational modification information extraction from the scientific literature

Article References: Bhowmik, S., Ross, K., Chen, C., Wu, C., & Vijay-Shanker, K. (2026). GenPTM: a generalizable framework for protein post-translational modification information extraction from the scientific literature. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06663-1

Image Credits: AI Generated

DOI: 10.1186/s12859-026-06663-1

Keywords: post-translational modification, information extraction, text mining, BiomedBERT, phosphorylation, ubiquitination, proteomics, PubMed, natural language processing, protein sites, BMC Bioinformatics, machine learning

Cite Scienmag News

Drew Townsend. (October 2, 2026). AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On. Scienmag. https://scienmag.com/ai-framework-reads-the-literature-to-map-protein-modifications-it-was-never-trained-on/

Drew Townsend. "AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On." Scienmag, 2 October 2026, https://scienmag.com/ai-framework-reads-the-literature-to-map-protein-modifications-it-was-never-trained-on/. Accessed 2 October 2026.

Drew Townsend. "AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On." Scienmag. October 2, 2026. https://scienmag.com/ai-framework-reads-the-literature-to-map-protein-modifications-it-was-never-trained-on/

Tags: advancements in proteomics researchAI-based literature miningautomated annotation of protein modificationsbioinformatics tools for PTMsBiomedBERTBMC Bioinformaticsdiscovery of novel PTMsGenPTM software for PTM mappinghandling large-scale scientific literatureinformation extractionMachine learningmachine learning in biological researchnatural language processingphosphorylationpost-translational modificationprotein post-translational modificationsprotein sitesProteomicsproteomics data curationPTM knowledge extractionPubMedstructured biological knowledge databasestext miningubiquitination
Share26Tweet16
Previous Post

Gut Microbes to the Rescue: Probiotics Shield Kidneys and Sperm From E. coli Damage in Rats

Next Post

Gut Hormone Ghrelin Rejuvenates Aging Brain Immune Cells to Restore Memory

Related Posts

Gut Hormone Ghrelin Rejuvenates Aging Brain Immune Cells to Restore Memory
Biology

Gut Hormone Ghrelin Rejuvenates Aging Brain Immune Cells to Restore Memory

October 2, 2026
Gut Microbes to the Rescue: Probiotics Shield Kidneys and Sperm From E. coli Damage in Rats
Biology

Gut Microbes to the Rescue: Probiotics Shield Kidneys and Sperm From E. coli Damage in Rats

October 2, 2026
Scientists Reveal Hidden Clustering Switch That Amplifies Lymphatic Vessel Growth Signals
Biology

Scientists Reveal Hidden Clustering Switch That Amplifies Lymphatic Vessel Growth Signals

October 2, 2026
Scientists Track Down the Genes That Let Maize Survive Drought in China’s Arid Northwest
Biology

Scientists Track Down the Genes That Let Maize Survive Drought in China’s Arid Northwest

October 2, 2026
Chromatin Remodeler INO80 Emerges as Master Gatekeeper of Lymphatic Vessel Development
Biology

Chromatin Remodeler INO80 Emerges as Master Gatekeeper of Lymphatic Vessel Development

October 2, 2026
Loss of Protective Bacteria Marks the Microbial Landscape of Cervical Cancer
Biology

Loss of Protective Bacteria Marks the Microbial Landscape of Cervical Cancer

October 2, 2026
Next Post
Gut Hormone Ghrelin Rejuvenates Aging Brain Immune Cells to Restore Memory

Gut Hormone Ghrelin Rejuvenates Aging Brain Immune Cells to Restore Memory

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Gut Hormone Ghrelin Rejuvenates Aging Brain Immune Cells to Restore Memory
  • AI Framework Reads the Literature to Map Protein Modifications It Was Never Trained On
  • Gut Microbes to the Rescue: Probiotics Shield Kidneys and Sperm From E. coli Damage in Rats
  • Teaching Patients to Manage Kidney Disease May Lift Quality of Life, Review Finds

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading