<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>sequence analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/sequence-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 06:10:01 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>sequence analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns the Language of Hormones to Predict Peptide Drugs</title>
		<link>https://scienmag.com/ai-learns-the-language-of-hormones-to-predict-peptide-drugs/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 06:10:01 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advances in peptide therapeutics]]></category>
		<category><![CDATA[AI for hormone sequence analysis]]></category>
		<category><![CDATA[AI-driven discovery of peptide hormones]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[BiLSTM neural networks for peptide function prediction]]></category>
		<category><![CDATA[bioinformatics]]></category>
		<category><![CDATA[class imbalance]]></category>
		<category><![CDATA[computational biology]]></category>
		<category><![CDATA[computational biology for hormone research]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning in drug discovery]]></category>
		<category><![CDATA[drug discovery]]></category>
		<category><![CDATA[ESM2]]></category>
		<category><![CDATA[hormone peptide identification techniques]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning models for biological sequence analysis]]></category>
		<category><![CDATA[peptide drugs]]></category>
		<category><![CDATA[peptide hormone prediction]]></category>
		<category><![CDATA[peptide hormone regulation and signaling]]></category>
		<category><![CDATA[peptide hormones]]></category>
		<category><![CDATA[peptide-based drug development]]></category>
		<category><![CDATA[protein language models]]></category>
		<category><![CDATA[protein language models for peptide classification]]></category>
		<category><![CDATA[sequence analysis]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233814</guid>

					<description><![CDATA[A new deep learning framework called pLM-HP combines the ESM2 protein language model with a BiLSTM network to predict peptide hormones with record accuracy, outperforming existing methods across every major metric.]]></description>
										<content:encoded><![CDATA[<p>Peptide hormones are among the body&#8217;s most powerful chemical messengers. These short chains of amino acids regulate metabolism, growth, development, and the delicate balance of homeostasis in organisms ranging from plants to humans. When their regulation goes awry, the consequences can be serious, which is precisely why they have become prized targets in drug development. Unlike small-molecule drugs, peptides carry more hydrogen bond donors and acceptors, allowing them to bind their targets with higher specificity and fewer off-target side effects. They are also readily degraded by enzymes in the body, show low metabolic toxicity, and exhibit good biocompatibility. Yet identifying which of the countless short peptide sequences in a genome actually function as hormones has long been a slow, expensive, and experimentally demanding task.</p>
<p>Now, a research team led by Chunyan Ao, Shihu Jiao, Xi Su, and Huan Yang has unveiled a deep learning framework called pLM-HP that promises to change that. Writing in the Journal of Advanced Research, the authors describe a system that combines a pre-trained protein language model, ESM2, with a bidirectional long short-term memory network, or BiLSTM, to predict whether a given peptide sequence is a hormone. On an independent test set, the model achieved a balanced accuracy of 95.64 percent, with sensitivity of 94.48 percent and specificity of 96.79 percent, a Matthews correlation coefficient of 0.824, and an area under the ROC curve of 0.991. Those numbers place it well ahead of existing tools, and they hint at a broader shift in how computational biology identifies functional molecules.</p>
<p>The challenge the researchers set out to solve is not new. Traditional identification of peptide hormones relies on experimental techniques such as mass spectrometry, affinity purification, and activity screening, all of which are complex and time-consuming. Computational methods emerged as a complementary approach, but early efforts leaned on sequence similarity searches or motif recognition algorithms, which falter when homologous sequences are scarce. Machine learning brought improvements: the HOPPred model, for example, integrates BLAST similarity, MERCI motifs, and logistic regression, achieving an AUROC of 0.96 and an MCC of 0.80 on an independent test set, while mHPpred adopted a multi-view feature representation and a meta-modeling strategy to build a more robust ensemble. Still, many methods remained tethered to hand-crafted features or predefined motifs, limiting their ability to represent sequences fully, and class imbalance in real data often biased models toward the majority class, dulling their sensitivity to genuine hormone peptides.</p>
<p>Protein language models have upended that paradigm. Models such as ProtBERT, ProtT5, and ESM2 are pre-trained on enormous collections of unlabeled protein sequences using masked language modeling, in which the model learns to predict hidden amino acids from their surrounding context. In doing so, they acquire high-dimensional embeddings that encode evolutionary, structural, and physicochemical properties of proteins. These embeddings have already powered advances in post-translational modification prediction, peptide structure modeling, and functional peptide classification. But until now, their use in peptide hormone prediction had been limited, and no systematic study had evaluated how well they perform on this specific task.</p>
<p>The pLM-HP framework treats a peptide sequence like a sentence and each amino acid like a word. ESM2, built on the Transformer encoder architecture and pre-trained on UniRef50 sequences, processes the sequence with a multi-head self-attention mechanism that captures long-range dependencies between residues, combined with feed-forward networks for nonlinear feature transformation. Each input is augmented with special beginning-of-sequence and end-of-sequence tokens, and token and positional embeddings are combined to capture residue order and context. Crucially, the researchers used ESM2 solely as a frozen feature extractor, without fine-tuning, and passed the residue-level embeddings directly to the downstream classifier.</p>
<p>One of the study&#8217;s most striking findings concerns model size. The team benchmarked four ESM2 variants, ranging from the 8-million-parameter esm2_t6_8M to the 650-million-parameter esm2_t33_650M, with output dimensions of 320, 480, 640, and 1280 respectively. All achieved balanced accuracy of roughly 95 percent in both five-fold cross-validation and independent testing, but the mid-sized esm2_t12_35M model, with 480-dimensional outputs, performed best, reaching a balanced accuracy of 95.61 percent and an MCC of 0.802 in cross-validation and 95.64 percent and 0.824 on the independent test set. The largest model, despite its far greater capacity, actually performed slightly worse, with a balanced accuracy of 95.15 percent and an MCC of 0.804, suggesting that for short peptide sequences of 11 to 41 residues, bigger is not necessarily better and may even introduce redundant features or mild overfitting.</p>
<p>The choice of classifier mattered just as much. When the researchers compared deep learning architectures and traditional machine learning methods on top of the ESM2 features, the BiLSTM emerged as the clear winner. By running two recurrent networks in forward and backward directions and concatenating their hidden states, the BiLSTM exploits both past and future context around each residue, capturing contextual dependencies that other architectures miss. It outperformed convolutional neural networks, multilayer perceptrons, support vector machines, logistic regression, random forests, XGBoost, and LightGBM across the board. Tree-based models fared worst: random forest achieved only about 73 percent balanced accuracy on the independent test set, with a strong bias toward predicting negative samples, while XGBoost and LightGBM reached high specificity near 98 percent but sensitivity below 81 percent, limiting their ability to spot true hormone peptides.</p>
<p>Class imbalance posed another formidable obstacle, since hormone peptides are vastly outnumbered by non-hormone sequences in real datasets. The team built their training data from 5,729 plant- and animal-derived peptide hormone sequences curated in the Hmrbase2 database, reduced to 1,174 non-redundant positives using CD-HIT at a 60 percent similarity threshold, paired with 11,740 non-hormone peptides drawn from PeptideAtlas. Rather than relying on resampling tricks, they incorporated class weights into a weighted binary cross-entropy loss, boosting the contribution of the rare hormone class. When they stress-tested the model on training sets with positive-to-negative ratios ranging from 1:1 to 1:10, performance remained remarkably stable, with balanced accuracy consistently above 94 percent, MCC values between 0.741 and 0.824, AUC values from 0.983 to 0.992, and AUPRC values from 0.841 to 0.915. Notably, the gap between sensitivity and specificity stayed small, at just 2.31 percentage points on the final test set, indicating a balanced classifier rather than one that simply favors the majority class.</p>
<p>The advantages over traditional approaches were dramatic. Handcrafted sequence descriptors such as pseudo amino acid composition, composition of k-spaced amino acid pairs, and quasi-sequence-order descriptors, when paired with classifiers like XGBoost, showed high specificity but weak sensitivity, and classical imbalance-handling strategies such as SMOTE and ADASYN often improved sensitivity at the cost of specificity, producing unstable results. In head-to-head comparisons on the same independent test set, pLM-HP outperformed HOPPred on every metric, improving balanced accuracy by 24.33 percentage points, sensitivity by 35.00 points, and specificity by 13.64 points, while raising the MCC by 0.522 and the AUC by 0.264. Against mHPpred, it improved balanced accuracy by 9.95 points, specificity by 18.51 points, MCC by 0.368, and AUC by 0.055. Visualization analyses using t-SNE and UMAP reinforced the story: while handcrafted features left positive and negative samples substantially overlapping, and raw ESM2 embeddings only partly separated them, the representations refined by the BiLSTM formed compact, well-separated clusters.</p>
<p>The implications reach well beyond one prediction task. As high-throughput omics and peptide-based drug development accelerate, efficiently and reliably identifying hormone-functional peptides from vast sequence spaces has become a key bridge between basic research and drug discovery. Peptide hormones, capable of probing and modulating protein-protein interactions, are considered ideal scaffolds for drug design, and a tool that can screen candidate sequences with near-99 percent AUC could dramatically narrow the experimental search space. The authors are candid about limitations: the current framework relies on linear sequence information and does not explicitly integrate structural context that shapes peptide conformation and signaling. Future work, they suggest, could incorporate structure predictions from AlphaFold2, apply geometric deep learning with graph neural networks, and explore focal loss, contrastive learning, and fine-tuning strategies to improve robustness under extreme imbalance and cross-dataset transfer. Even so, pLM-HP offers a compelling demonstration that when protein language models meet the right sequence-aware classifier, the hidden language of hormonal peptides becomes far easier to read.</p>
<p><strong>Subject of Research:</strong> Computational prediction of peptide hormones using pre-trained protein language model embeddings and deep learning</p>
<p><strong>Article Title:</strong> pLM-HP: Peptide hormone prediction using pre-trained protein language model representations</p>
<p><strong>Article References:</strong> Ao, C., Jiao, S., Su, X., &amp; Yang, H. (2026). pLM-HP: Peptide hormone prediction using pre-trained protein language model representations. <em>Journal of Advanced Research</em>. <a href="https://doi.org/10.1016/j.jare.2026.10.002" rel="noopener noreferrer">https://doi.org/10.1016/j.jare.2026.10.002</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.jare.2026.10.002" rel="noopener noreferrer">10.1016/j.jare.2026.10.002</a></p>
<p><strong>Keywords:</strong> peptide hormones, protein language models, ESM2, BiLSTM, deep learning, drug discovery, machine learning, bioinformatics, class imbalance, sequence analysis, peptide drugs, computational biology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233814</post-id>	</item>
	</channel>
</rss>
