<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>ESM-1v &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/esm-1v/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 17:08:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>ESM-1v &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Model Fills the Gaps in Protein Mutation Maps</title>
		<link>https://scienmag.com/ai-model-fills-the-gaps-in-protein-mutation-maps/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 17:08:11 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[amino acid substitution matrices]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[bioinformatics protein mutation tools]]></category>
		<category><![CDATA[deep mutational scanning]]></category>
		<category><![CDATA[ESM-1v]]></category>
		<category><![CDATA[evolutionary conservation in proteins]]></category>
		<category><![CDATA[gradient boosting]]></category>
		<category><![CDATA[handling missing data in mutational datasets]]></category>
		<category><![CDATA[Human Domainome]]></category>
		<category><![CDATA[imputation]]></category>
		<category><![CDATA[LightGBM protein analysis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in protein analysis]]></category>
		<category><![CDATA[MaveDB]]></category>
		<category><![CDATA[missense variants]]></category>
		<category><![CDATA[protein language models]]></category>
		<category><![CDATA[protein mutation mapping]]></category>
		<category><![CDATA[protein stability]]></category>
		<category><![CDATA[protein variant effect maps]]></category>
		<category><![CDATA[protein variant effect prediction]]></category>
		<category><![CDATA[transformer-based protein language models]]></category>
		<category><![CDATA[variant effect prediction]]></category>
		<category><![CDATA[VEFill protein mutation prediction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=228707</guid>

					<description><![CDATA[A new gradient boosting framework called VEFill accurately imputes missing deep mutational scanning scores across protein domains, outperforming existing methods whenever at least 20 percent of variants have been measured.]]></description>
										<content:encoded><![CDATA[<p>Deep mutational scanning has transformed the way biologists interrogate proteins, allowing thousands of amino acid substitutions to be tested in parallel and assembled into detailed maps of variant effects. Yet for all its power, the technique rarely delivers a complete picture. Low sequencing depth, limited coverage, and assay dropout mean that many mutations in a typical experiment simply have no measured score. A new study published in Molecular Systems Biology tackles this persistent problem head-on, introducing a machine learning framework called VEFill that can accurately fill in the missing values of deep mutational scanning datasets and generalize to proteins it has never seen before.</p>
<p>Developed by Polina Polunina and Wolfgang Maier of the University of Freiburg together with Alan Rubin of the Walter and Eliza Hall Institute of Medical Research, VEFill is built on LightGBM, a gradient boosting framework well suited to structured, high-dimensional tabular data. Rather than relying on a single source of information, the model integrates a rich set of biologically informed features: evolutionary conservation scores from the EVE generative model, amino acid substitution matrices such as BLOSUM62 and PAM250, physicochemical descriptors of the wild-type and variant residues, and contextual sequence embeddings from ESM-1v, a transformer-based protein language model trained on millions of natural sequences. The authors trained and optimized the model using Bayesian hyperparameter search and stored all underlying data in a structured PostgreSQL database to ensure reproducibility.</p>
<p>The training ground for VEFill was the Human Domainome 1 dataset, a standardized collection of stability measurements generated with an abundance-based protein fragment complementation assay across 522 human protein domains, using site-saturation mutagenesis to introduce every possible amino acid substitution. From this resource, the team assembled a feature-complete subset of 140 domains comprising 136,854 mutations for the full model, and used the broader set of 521 domains, totaling 562,208 mutations, to train reduced-feature versions that depend only on ESM-1v embeddings and positional mean scores.</p>
<p>The performance figures are striking. In a leave-protein-out evaluation, where the model must predict variant effects for protein domains entirely absent from training, the best configuration achieved a coefficient of determination of 0.64 and a Pearson correlation of 0.80. Feature ablation experiments revealed a clear hierarchy of information: models relying solely on evolutionary scores, substitution matrices, or biochemical descriptors performed poorly, with R-squared values below 0.15. Adding ESM-1v embeddings lifted performance substantially, and combining the embeddings with the mean experimental score per position proved even more powerful, indicating that learned sequence representations and data-driven positional priors capture complementary aspects of mutational tolerance.</p>
<p>Perhaps the most practically important finding concerns data scarcity. In controlled subsampling experiments on 28 high-quality datasets, VEFill consistently outperformed a battery of existing imputation methods, including Envision, FactorizeDMS, AALasso, and nearest-neighbor approaches, once at least 20 percent of variants had been experimentally measured. Below that threshold, predictions became markedly harder, and the authors are candid that true zero-shot prediction without any positional context remains challenging, particularly for functionally complex proteins. This matters because a survey of 841 public score sets from the MaveDB repository showed that fully complete datasets are vanishingly rare, with only 44 out of 841 achieving complete coverage of all possible substitutions. The vast majority of real-world experiments therefore fall squarely within the regime where VEFill delivers its greatest benefit.</p>
<p>The team also probed how few measurements are actually needed per position. When per-protein models were trained with a restricted number of substitutions per site, accuracy rose rapidly and began to plateau at roughly four mutations per position. Intriguingly, models trained exclusively on five carefully chosen amino acids, histidine, glutamic acid, asparagine, isoleucine, and glycine, each representing a distinct physicochemical class, matched the performance of models trained on randomly selected substitutions at equivalent coverage. This suggests that sparse, information-efficient mutational libraries could be designed deliberately, prioritizing a small representative set of substitutions per site without sacrificing predictive power.</p>
<p>Noise ceiling analyses added an important dose of realism. Because experimental variability imposes a fundamental upper bound on achievable accuracy, the researchers simulated replicate experiments using reported measurement uncertainties and found that estimated ceilings ranged from 0.68 to 0.99 across domains. VEFill&#8217;s per-protein correlations, between 0.52 and 0.91 under an 80/20 split, sat consistently below but generally close to these limits, indicating that much of the remaining discrepancy between predicted and observed scores reflects measurement noise rather than model failure. Error profiles were lowest for near-neutral variants, where data are densest, and predictions at the extreme deleterious and high-activity tails should be read as conservative approximations rather than precise estimates.</p>
<p>Generalization beyond the training data showed both promise and limits. When the cross-protein model was applied to eight full-length proteins from independent MaveDB datasets, including Parkin, PTEN, aspartoacylase, calmodulin, TPK1, and TP53, it performed substantially better on stability-based assays than on activity-based ones, a mismatch the authors attribute to differences between training and test assay modalities and the complexity of cellular phenotypes. A lightweight two-feature version using only ESM-1v embeddings and positional mean scores performed comparably to the full model in most settings, offering a practical alternative when evolutionary scores or other annotations are unavailable. Analysis of protein family composition revealed that a few Pfam families dominated the training data, but retraining on a Pfam-unique subset increased error only modestly, suggesting the findings are robust rather than inflated by family-level redundancy.</p>
<p>Fine-grained evaluations exposed the model&#8217;s blind spots in biologically meaningful ways. Proline substitutions consistently produced elevated errors, likely because their unusual conformational constraints disrupt secondary structure in ways that sequence-based features do not fully capture. In the TRIM44 zinc finger domain, histidine-to-cysteine mutations were poorly predicted when entire positions were held out but accurately recovered when single variants were withheld, hinting at context-dependent chemistry, such as zinc coordination, that only fine-grained positional information can resolve. The authors suggest that future versions could incorporate explicit structural features, cross-species transfer learning, or systematic assay harmonization to close these gaps.</p>
<p>Overall, VEFill arrives at a moment when the variant effect field is scaling rapidly, with MaveDB now listing more than 2,600 public datasets and over 1,100 for human proteins. By providing an interpretable, scalable tool that turns partially complete mutational maps into denser, more usable resources, the study lowers the experimental barrier for variant prioritization, protein engineering, and the clinical interpretation of missense mutations. The code and trained models have been released openly, inviting the community to apply sparse-library design strategies and, ultimately, to build the comprehensive Atlas of Variant Effects that the field has long envisioned.</p>
<p><strong>Subject of Research:</strong> Machine learning-based imputation of missing scores in deep mutational scanning datasets across human protein domains</p>
<p><strong>Article Title:</strong> VEFill: accurate and generalizable deep mutational scanning score imputation across protein domains</p>
<p><strong>Article References:</strong> Polunina, P. V., Maier, W., &amp; Rubin, A. F. (2026). VEFill: accurate and generalizable deep mutational scanning score imputation across protein domains. <em>Molecular Systems Biology, 22</em>(6), 979-1002. <a href="https://doi.org/10.1038/s44320-026-00203-y" rel="noopener noreferrer">https://doi.org/10.1038/s44320-026-00203-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44320-026-00203-y" rel="noopener noreferrer">10.1038/s44320-026-00203-y</a></p>
<p><strong>Keywords:</strong> deep mutational scanning, variant effect prediction, machine learning, gradient boosting, protein stability, ESM-1v, protein language models, MaveDB, Human Domainome, imputation, missense variants, bioinformatics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">228707</post-id>	</item>
	</channel>
</rss>
