<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>computational biology methods &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/computational-biology-methods/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 01 Apr 2026 18:40:25 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>computational biology methods &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Protein Language Model Accuracy Test Sheds Light on AI’s &#8216;Black Box&#8217;</title>
		<link>https://scienmag.com/protein-language-model-accuracy-test-sheds-light-on-ais-black-box/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 01 Apr 2026 18:40:25 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in biology]]></category>
		<category><![CDATA[AI interpretability in bioinformatics]]></category>
		<category><![CDATA[AI model confidence assessment]]></category>
		<category><![CDATA[biological sequence analysis]]></category>
		<category><![CDATA[computational biology methods]]></category>
		<category><![CDATA[molecular biology and machine learning]]></category>
		<category><![CDATA[Nature Methods protein research]]></category>
		<category><![CDATA[protein embedding evaluation]]></category>
		<category><![CDATA[protein language models]]></category>
		<category><![CDATA[protein structure prediction]]></category>
		<category><![CDATA[synthetic vs natural protein sequences]]></category>
		<category><![CDATA[trustworthiness of AI predictions]]></category>
		<guid isPermaLink="false">https://scienmag.com/protein-language-model-accuracy-test-sheds-light-on-ais-black-box/</guid>

					<description><![CDATA[In recent years, artificial intelligence (AI) language models have become ubiquitous tools in generating human-like text, powering chatbots, and automating content creation across various domains. Fascinatingly, these advances have spilled over into the realm of biology, where researchers are harnessing language models to decipher the complex information encoded in DNA and proteins. By conceptualizing biological [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In recent years, artificial intelligence (AI) language models have become ubiquitous tools in generating human-like text, powering chatbots, and automating content creation across various domains. Fascinatingly, these advances have spilled over into the realm of biology, where researchers are harnessing language models to decipher the complex information encoded in DNA and proteins. By conceptualizing biological sequences as a form of language, these models analyze patterns and relationships within the vast diversity of biomolecules, accelerating predictions and providing fresh insights into the intricacies of life’s molecular machinery. Despite their promise, a critical challenge has persisted—determining the reliability and confidence of the predictions generated by these models remains elusive.</p>
<p>Addressing this significant gap, computational biologists at Emory University have introduced an innovative approach that sets out to quantify the trustworthiness of protein language model embeddings. Published in <em>Nature Methods</em>, their novel framework evaluates the quality of the model’s internal representations by contrasting embeddings of natural proteins with those generated from synthetic, random sequences. This comparative strategy enables researchers to discern how confidently the model distinguishes biologically meaningful signals from noise, marking a transformative step in understanding and validating AI-driven biological inferences.</p>
<p>Embeddings refer to the numerical representation language models assign to data, condensing complex inputs into abstract vectors within a latent space where proximity implies similarity. In protein language models, this latent space metaphorically catalogues protein sequences based on structural and functional features discerned during training. Emory’s team visualized this latent space as a scatter plot, observing that natural proteins cluster according to their evolutionary and functional subtypes, while synthetic sequences, devoid of biological relevance, occupy distinctly separate regions. They coined this latter region the “junkyard,” positing it as a repository of low-quality embeddings that reflect the model’s unfamiliarity with non-biological sequences.</p>
<p>Central to their methodology is the concept of a “random neighbor score,” a metric that quantifies the proximity between a given protein’s embedding and those of synthetic, random sequences within the latent space. A low score indicates few or no synthetic neighbors nearby, suggesting the model’s high confidence in the biological validity of that protein’s embedding. Conversely, a high score implies the embedding is closer to the “junkyard,” signaling uncertainty or diminished reliability. This quantitative measure provides a computationally efficient and biologically grounded proxy for evaluating the model’s predictive certainty across diverse proteins and subsequences.</p>
<p>The implications of this development reach far beyond theoretical considerations. Protein sequences, encoded by DNA, fold into intricate three-dimensional structures that underpin nearly all cellular processes—from catalysis and signaling to defense mechanisms. With over 200 million protein sequences cataloged in databases such as UniProt, language models have a substantial foundation for training. Still, the true diversity of proteins extends into the trillions, much of it residing in the enigmatic microbial world that community metagenomes represent. Emory’s framework offers a vital means to assess whether inferences drawn from limited sample sets can generalize reliably to this vast, largely uncharted biosphere.</p>
<p>The endeavor to understand metagenomic complexity is vital because microorganisms do not exist in isolation; they form dynamic communities that profoundly impact host health and ecosystem functions. Characterizing the proteins encoded by these communities uncovers biochemical pathways and interactions that conventional experimental methods struggle to elucidate at scale. By sharpening the lens through which AI models interpret protein sequences, the newly introduced confidence metric enhances the promise of computational biology to unlock unprecedented biological insights.</p>
<p>This advancement also illuminates the intricate process of evolution etched into protein sequences. Evolution conserves amino acid residues essential for a protein’s function, imprinting a signature that language models learn to recognize during training. Natural proteins, therefore, exhibit coherent embedding patterns, reflective of their functional and structural constraints shaped over billions of years. Synthetic random sequences, bereft of adaptive significance, lack these signatures and cluster apart in embedding space. By leveraging this evolutionary contrast, the Emory team has devised a method to “peer inside” the black box of AI models, exposing the hallmark features that guide their predictions.</p>
<p>Further validation demonstrated that embeddings flagged as low-quality or uncertain by the random neighbor score tended to perform poorly in downstream biological tasks. These included function prediction, structural modeling, and interaction inference—core applications where misinterpretation can misguide research efforts. Therefore, employing this uncertainty measure not only improves model interpretability but also safeguards the fidelity of scientific conclusions derived from AI predictions, fostering greater trust in computational methodologies.</p>
<p>The simplicity and elegance of the approach belie its broad applicability across the burgeoning landscape of biological language models. As new architectures and training paradigms emerge, providing an intrinsic metric for embedding quality becomes crucial for optimizing model design and application. The concept of a biologically anchored uncertainty score represents a paradigm shift, moving away from generic metrics borrowed from computer science toward domain-specific criteria that align more closely with empirical biological evidence.</p>
<p>Just as a surgeon relies on the sharpest instruments to maximize precision and minimize risk, computational biologists can now choose and refine AI models with enhanced awareness of their limitations and strengths. This precision becomes exceptionally vital when extrapolating to the unknown proteomic “dark matter” found in environmental and clinical microbiomes, where experimental validation lags and computational predictions hold the key to discovery.</p>
<p>By fostering heightened quality control at every stage of protein data modeling—from sequence input through embedding to downstream prediction—this method mitigates the compounding of errors that can arise when working with noisy or unrepresentative datasets. Accurate uncertainty quantification is paramount in a field where even slight errors can propagate through complex biological networks, leading to misleading interpretations or missed opportunities.</p>
<p>This pioneering research was supported by the National Science Foundation and represents a milestone in integrating AI reliability with molecular biology. The work emboldens the interface between computational simulation and empirical biology, highlighting the immense potential of language models while tempering enthusiasm with rigorous validation. As AI-driven biology continues to evolve, transparency and confidence measures like the random neighbor score will shape the trajectory toward more robust and insightful discoveries.</p>
<hr />
<p><strong>Subject of Research</strong>: Not applicable</p>
<p><strong>Article Title</strong>: Quantifying uncertainty in protein representations across models and tasks</p>
<p><strong>News Publication Date</strong>: 1-Apr-2026</p>
<p><strong>Web References</strong>: <a href="http://dx.doi.org/10.1038/s41592-026-03028-7">DOI link</a></p>
<p><strong>Image Credits</strong>: Bromberg lab</p>
<h4><strong>Keywords</strong></h4>
<p>Bioinformatics, Sequence analysis, Research methods, Complex networks, Computers, Metagenomics, Protein functions</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">148277</post-id>	</item>
		<item>
		<title>New Computational Method Promises to Compress Decades of Disease Biology Research into Days</title>
		<link>https://scienmag.com/new-computational-method-promises-to-compress-decades-of-disease-biology-research-into-days/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Tue, 11 Nov 2025 22:07:44 +0000</pubDate>
				<category><![CDATA[Cancer]]></category>
		<category><![CDATA[accelerated disease research techniques]]></category>
		<category><![CDATA[advancements in disease biology]]></category>
		<category><![CDATA[biochemical processes in human cells]]></category>
		<category><![CDATA[cancer and Alzheimer’s disease studies]]></category>
		<category><![CDATA[cellular pH impact on health]]></category>
		<category><![CDATA[computational biology methods]]></category>
		<category><![CDATA[high-throughput protein analysis techniques]]></category>
		<category><![CDATA[Notre Dame research innovations]]></category>
		<category><![CDATA[pH-sensitive proteins research]]></category>
		<category><![CDATA[protein activity modulation]]></category>
		<category><![CDATA[protein structure dynamics]]></category>
		<category><![CDATA[therapeutic interventions in cell biology]]></category>
		<guid isPermaLink="false">https://scienmag.com/new-computational-method-promises-to-compress-decades-of-disease-biology-research-into-days/</guid>

					<description><![CDATA[At the microscopic scale of biology, the smallest components often exert the most profound influences. Human cells, measuring approximately ten micrometers across, host an intricate network of biochemical processes that dictate life at the cellular level. Among the most critical yet underappreciated factors shaping these processes is the concentration of protons, or pH, within cells. [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>At the microscopic scale of biology, the smallest components often exert the most profound influences. Human cells, measuring approximately ten micrometers across, host an intricate network of biochemical processes that dictate life at the cellular level. Among the most critical yet underappreciated factors shaping these processes is the concentration of protons, or pH, within cells. Slight fluctuations in pH can dramatically alter cellular functions such as movement, division, and signal transduction. These pH changes are not merely biochemical trivia; they have been implicated as accelerants in the progression of severe illnesses including cancer, Alzheimer’s disease, and Huntington’s disease.</p>
<p>Understanding how protein structures respond dynamically to pH changes has remained a challenging frontier in cell biology. Proteins, the molecular workhorses of cells, often undergo conformational shifts that modulate their activity in response to the acidic or basic environment. Discerning which proteins are sensitive to these pH variations is paramount, as it may unlock new pathways for therapeutic intervention. Currently, experimental approaches to identify pH-sensitive proteins are labor-intensive and time-consuming, often requiring painstaking analyses of individual proteins in isolation.</p>
<p>In a groundbreaking advancement, researchers at the University of Notre Dame have introduced a powerful computational pipeline capable of scanning hundreds of proteins within days, rather than years. This novel method accelerates the identification of pH-sensitive domains within proteins, revolutionizing the initial screening phase of biomolecular research. By leveraging existing structural data and experimental insights, the team created an algorithmic process that predicts specific residues within proteins that could mediate pH-dependent allosteric regulation.</p>
<p>Dr. Katharine White, Clare Boothe Luce Assistant Professor in the Department of Chemistry and Biochemistry at Notre Dame, emphasized the transformative nature of this technology. “Prior to this development, scientists were searching for a needle in a haystack when identifying pH-responsive proteins,” she remarked. The computational pipeline effectively refines that haystack into a manageable collection of candidate proteins, setting the stage for focused experimentation and drug design.</p>
<p>Historically, only a handful of cytoplasmic proteins — approximately seventy — have been validated as pH-sensitive via experimental studies despite the hypothesis that many more possess this characteristic. Moreover, detailed mechanistic insights exist for fewer than a third of these known proteins. The challenge stems from the complexity of measuring pH-dependent conformational changes, which often involve subtle shifts in ionizable amino acid networks that are difficult to capture through traditional experimental modalities.</p>
<p>The new study, recently published in the journal Science Signaling, represents an important leap forward. With funding support from the National Science Foundation and the National Institutes of Health, White and her team formed a modular pipeline adept at integrating conformational data from protein crystal structures, pKa predictions of ionizable groups, and bioinformatic annotations. The pipeline can systematically identify so-called “ionizable networks,” clusters of amino acids whose protonation states modulate protein structure and function in response to pH changes.</p>
<p>A particularly salient application of this method was the analysis of the Src homology 2 (SH2) domain, a conserved protein module central to signal transduction pathways regulating cell growth, differentiation, and immune responses. The SH2 domain is recurrently mutated in various cancers, making it a prime target for understanding pH-mediated regulatory mechanisms. White’s team experimentally validated the in silico prediction that the SH2 domain exhibits marked pH sensitivity, confirming both its biological relevance and the accuracy of the computational model.</p>
<p>Further insights emerged concerning c-Src, a non-receptor tyrosine kinase with pivotal roles in oncogenic signaling. The study elucidated the precise molecular locale where pH influences c-Src activity, underscoring how acid-base chemistry interfaces with protein allosteric regulation. Such mechanistic clarity holds promise for the development of precision therapeutics that exploit the protonation states of key residues to modulate enzyme function selectively.</p>
<p>Papa Kobina Van Dyck, lead author and recent doctoral graduate in biophysics at Notre Dame, reflected on the magnitude of the achievement: “We condensed what would have taken decades of biochemical experimentation into a matter of weeks using computational methods.” This acceleration dramatically enhances the pace at which research can move from hypothesis to experimental validation and, eventually, clinical application.</p>
<p>Beyond cancer and neurodegeneration, the implications of mapping pH-sensitive protein networks extend to a broad spectrum of medical conditions characterized by dysregulated pH dynamics, including diabetes, autoimmune diseases, and traumatic brain injury. The Notre Dame pipeline therefore represents a versatile tool not only for fundamental biological discovery but also for translational efforts aimed at drug discovery and personalized medicine.</p>
<p>In summary, this pioneering work exemplifies how integrative computational biology can circumvent traditional experimental bottlenecks, offering new vistas for exploring the complex molecular choreography dictated by pH fluctuations in cells. By illuminating the ionizable networks that govern protein allostery, the study provides a foundation for innovative therapies targeting diseases that span oncology, neurology, and beyond.</p>
<p>For readers seeking to delve deeper into this transformative research, the full article titled “Ionizable networks mediate pH-dependent allostery in the SH2 domain–containing signaling proteins SHP2 and SRC” is accessible through Science Signaling. The comprehensive study meticulously outlines the computational methodologies and experimental validations that underpin this advancement, heralding a new era in cellular physiology and disease biology.</p>
<hr />
<p><strong>Subject of Research</strong>: pH-dependent regulation of protein structure and function in cellular signaling pathways.</p>
<p><strong>Article Title</strong>: Ionizable networks mediate pH-dependent allostery in the SH2 domain–containing signaling proteins SHP2 and SRC</p>
<p><strong>News Publication Date</strong>: 11-Nov-2025</p>
<p><strong>Web References</strong>:</p>
<ul>
<li>Original article: <a href="https://www.science.org/doi/10.1126/scisignal.adt3018">https://www.science.org/doi/10.1126/scisignal.adt3018</a>  </li>
<li>University of Notre Dame overview: <a href="https://research.nd.edu/news-and-events/news/new-computational-process-could-help-condense-decades-of-disease-biology-research-into-days/">https://research.nd.edu/news-and-events/news/new-computational-process-could-help-condense-decades-of-disease-biology-research-into-days/</a></li>
</ul>
<p><strong>Image Credits</strong>: Photo by Peter Ringenberg/University of Notre Dame</p>
<p><strong>Keywords</strong>: Cellular processes, Life sciences, Diseases and disorders, Breast cancer, Signaling pathways</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">104256</post-id>	</item>
	</channel>
</rss>
