Friday, October 2, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

AI Model Fills the Gaps in Protein Mutation Maps

October 2, 2026
in Biology
Drew Townsend
By Drew Townsend Scienmag Editorial Profile - Cell Biology
Reading Time: 4 mins read
0
AI Model Fills the Gaps in Protein Mutation Maps

AI Model Fills the Gaps in Protein Mutation Maps

AI Model Fills the Gaps in Protein Mutation Maps

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Deep mutational scanning has transformed the way biologists interrogate proteins, allowing thousands of amino acid substitutions to be tested in parallel and assembled into detailed maps of variant effects. Yet for all its power, the technique rarely delivers a complete picture. Low sequencing depth, limited coverage, and assay dropout mean that many mutations in a typical experiment simply have no measured score. A new study published in Molecular Systems Biology tackles this persistent problem head-on, introducing a machine learning framework called VEFill that can accurately fill in the missing values of deep mutational scanning datasets and generalize to proteins it has never seen before.

Developed by Polina Polunina and Wolfgang Maier of the University of Freiburg together with Alan Rubin of the Walter and Eliza Hall Institute of Medical Research, VEFill is built on LightGBM, a gradient boosting framework well suited to structured, high-dimensional tabular data. Rather than relying on a single source of information, the model integrates a rich set of biologically informed features: evolutionary conservation scores from the EVE generative model, amino acid substitution matrices such as BLOSUM62 and PAM250, physicochemical descriptors of the wild-type and variant residues, and contextual sequence embeddings from ESM-1v, a transformer-based protein language model trained on millions of natural sequences. The authors trained and optimized the model using Bayesian hyperparameter search and stored all underlying data in a structured PostgreSQL database to ensure reproducibility.

The training ground for VEFill was the Human Domainome 1 dataset, a standardized collection of stability measurements generated with an abundance-based protein fragment complementation assay across 522 human protein domains, using site-saturation mutagenesis to introduce every possible amino acid substitution. From this resource, the team assembled a feature-complete subset of 140 domains comprising 136,854 mutations for the full model, and used the broader set of 521 domains, totaling 562,208 mutations, to train reduced-feature versions that depend only on ESM-1v embeddings and positional mean scores.

The performance figures are striking. In a leave-protein-out evaluation, where the model must predict variant effects for protein domains entirely absent from training, the best configuration achieved a coefficient of determination of 0.64 and a Pearson correlation of 0.80. Feature ablation experiments revealed a clear hierarchy of information: models relying solely on evolutionary scores, substitution matrices, or biochemical descriptors performed poorly, with R-squared values below 0.15. Adding ESM-1v embeddings lifted performance substantially, and combining the embeddings with the mean experimental score per position proved even more powerful, indicating that learned sequence representations and data-driven positional priors capture complementary aspects of mutational tolerance.

Perhaps the most practically important finding concerns data scarcity. In controlled subsampling experiments on 28 high-quality datasets, VEFill consistently outperformed a battery of existing imputation methods, including Envision, FactorizeDMS, AALasso, and nearest-neighbor approaches, once at least 20 percent of variants had been experimentally measured. Below that threshold, predictions became markedly harder, and the authors are candid that true zero-shot prediction without any positional context remains challenging, particularly for functionally complex proteins. This matters because a survey of 841 public score sets from the MaveDB repository showed that fully complete datasets are vanishingly rare, with only 44 out of 841 achieving complete coverage of all possible substitutions. The vast majority of real-world experiments therefore fall squarely within the regime where VEFill delivers its greatest benefit.

The team also probed how few measurements are actually needed per position. When per-protein models were trained with a restricted number of substitutions per site, accuracy rose rapidly and began to plateau at roughly four mutations per position. Intriguingly, models trained exclusively on five carefully chosen amino acids, histidine, glutamic acid, asparagine, isoleucine, and glycine, each representing a distinct physicochemical class, matched the performance of models trained on randomly selected substitutions at equivalent coverage. This suggests that sparse, information-efficient mutational libraries could be designed deliberately, prioritizing a small representative set of substitutions per site without sacrificing predictive power.

Noise ceiling analyses added an important dose of realism. Because experimental variability imposes a fundamental upper bound on achievable accuracy, the researchers simulated replicate experiments using reported measurement uncertainties and found that estimated ceilings ranged from 0.68 to 0.99 across domains. VEFill’s per-protein correlations, between 0.52 and 0.91 under an 80/20 split, sat consistently below but generally close to these limits, indicating that much of the remaining discrepancy between predicted and observed scores reflects measurement noise rather than model failure. Error profiles were lowest for near-neutral variants, where data are densest, and predictions at the extreme deleterious and high-activity tails should be read as conservative approximations rather than precise estimates.

Generalization beyond the training data showed both promise and limits. When the cross-protein model was applied to eight full-length proteins from independent MaveDB datasets, including Parkin, PTEN, aspartoacylase, calmodulin, TPK1, and TP53, it performed substantially better on stability-based assays than on activity-based ones, a mismatch the authors attribute to differences between training and test assay modalities and the complexity of cellular phenotypes. A lightweight two-feature version using only ESM-1v embeddings and positional mean scores performed comparably to the full model in most settings, offering a practical alternative when evolutionary scores or other annotations are unavailable. Analysis of protein family composition revealed that a few Pfam families dominated the training data, but retraining on a Pfam-unique subset increased error only modestly, suggesting the findings are robust rather than inflated by family-level redundancy.

Fine-grained evaluations exposed the model’s blind spots in biologically meaningful ways. Proline substitutions consistently produced elevated errors, likely because their unusual conformational constraints disrupt secondary structure in ways that sequence-based features do not fully capture. In the TRIM44 zinc finger domain, histidine-to-cysteine mutations were poorly predicted when entire positions were held out but accurately recovered when single variants were withheld, hinting at context-dependent chemistry, such as zinc coordination, that only fine-grained positional information can resolve. The authors suggest that future versions could incorporate explicit structural features, cross-species transfer learning, or systematic assay harmonization to close these gaps.

Overall, VEFill arrives at a moment when the variant effect field is scaling rapidly, with MaveDB now listing more than 2,600 public datasets and over 1,100 for human proteins. By providing an interpretable, scalable tool that turns partially complete mutational maps into denser, more usable resources, the study lowers the experimental barrier for variant prioritization, protein engineering, and the clinical interpretation of missense mutations. The code and trained models have been released openly, inviting the community to apply sparse-library design strategies and, ultimately, to build the comprehensive Atlas of Variant Effects that the field has long envisioned.

Subject of Research: Machine learning-based imputation of missing scores in deep mutational scanning datasets across human protein domains

Article Title: VEFill: accurate and generalizable deep mutational scanning score imputation across protein domains

Article References: Polunina, P. V., Maier, W., & Rubin, A. F. (2026). VEFill: accurate and generalizable deep mutational scanning score imputation across protein domains. Molecular Systems Biology, 22(6), 979-1002. https://doi.org/10.1038/s44320-026-00203-y

Image Credits: AI Generated

DOI: 10.1038/s44320-026-00203-y

Keywords: deep mutational scanning, variant effect prediction, machine learning, gradient boosting, protein stability, ESM-1v, protein language models, MaveDB, Human Domainome, imputation, missense variants, bioinformatics

Cite Scienmag News

Drew Townsend. (October 2, 2026). AI Model Fills the Gaps in Protein Mutation Maps. Scienmag. https://scienmag.com/ai-model-fills-the-gaps-in-protein-mutation-maps/

Drew Townsend. "AI Model Fills the Gaps in Protein Mutation Maps." Scienmag, 2 October 2026, https://scienmag.com/ai-model-fills-the-gaps-in-protein-mutation-maps/. Accessed 2 October 2026.

Drew Townsend. "AI Model Fills the Gaps in Protein Mutation Maps." Scienmag. October 2, 2026. https://scienmag.com/ai-model-fills-the-gaps-in-protein-mutation-maps/

Tags: amino acid substitution matricesbioinformaticsbioinformatics protein mutation toolsdeep mutational scanningESM-1vevolutionary conservation in proteinsgradient boostinghandling missing data in mutational datasetsHuman DomainomeimputationLightGBM protein analysisMachine learningmachine learning in protein analysisMaveDBmissense variantsprotein language modelsprotein mutation mappingprotein stabilityprotein variant effect mapsprotein variant effect predictiontransformer-based protein language modelsvariant effect predictionVEFill protein mutation prediction
Share26Tweet16
Previous Post

Qatar’s Radiation Landscape: Soils Stay Safe While Oil-Field Sludge Raises Red Flags

Next Post

Noise-Proof AI Translation Brings Low-Resource Languages Into the Digital Age

Related Posts

India’s Native Dogs Carry a Genetic Signature All Their Own, Landmark SNP Study Reveals
Biology

India’s Native Dogs Carry a Genetic Signature All Their Own, Landmark SNP Study Reveals

October 2, 2026
Infrared Light and Machine Learning Reveal Which Animals Mosquitoes Bite
Biology

Infrared Light and Machine Learning Reveal Which Animals Mosquitoes Bite

October 2, 2026
Wine Waste Gets a Second Life: Grape Pomace Kombucha Wins Over Tasters
Biology

Wine Waste Gets a Second Life: Grape Pomace Kombucha Wins Over Tasters

October 2, 2026
Microplastics and BPA Quietly Rewire Fish Immunity, Raising Viral Disease Risk in Aquaculture
Biology

Microplastics and BPA Quietly Rewire Fish Immunity, Raising Viral Disease Risk in Aquaculture

October 2, 2026
New Web Platform Brings Structural Variant Benchmarking to the Browser
Biology

New Web Platform Brings Structural Variant Benchmarking to the Browser

October 2, 2026
Zebu Cattle Milk Protein Reveals Stable Structure in Computational Model
Biology

Zebu Cattle Milk Protein Reveals Stable Structure in Computational Model

October 2, 2026
Next Post
Noise-Proof AI Translation Brings Low-Resource Languages Into the Digital Age

Noise-Proof AI Translation Brings Low-Resource Languages Into the Digital Age

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Microbial Inoculants Give Teak Cuttings a Head Start in Reforestation
  • Electronic Noses Learn to Sniff Out Food Fraud and Spoilage
  • Noise-Proof AI Translation Brings Low-Resource Languages Into the Digital Age
  • AI Model Fills the Gaps in Protein Mutation Maps

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading