<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>protein embeddings &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/protein-embeddings/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 19:15:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>protein embeddings &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Hybrid AI Model Blends Transformer and BiLSTM to Predict Cancer Drug Synergy</title>
		<link>https://scienmag.com/hybrid-ai-model-blends-transformer-and-bilstm-to-predict-cancer-drug-synergy/</link>
		
		<dc:creator><![CDATA[Nathaniel Bowman]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:15:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI in Oncology]]></category>
		<category><![CDATA[benchmark dataset classification accuracy]]></category>
		<category><![CDATA[BiLSTM]]></category>
		<category><![CDATA[BT-Synergy]]></category>
		<category><![CDATA[cancer cell lines]]></category>
		<category><![CDATA[Cancer drug synergy prediction]]></category>
		<category><![CDATA[combination cancer therapy]]></category>
		<category><![CDATA[combination therapy]]></category>
		<category><![CDATA[computational drug discovery]]></category>
		<category><![CDATA[computational pharmacology]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[drug interaction modeling]]></category>
		<category><![CDATA[drug pair screening automation]]></category>
		<category><![CDATA[drug synergy prediction]]></category>
		<category><![CDATA[DrugCombDB]]></category>
		<category><![CDATA[hybrid deep learning models]]></category>
		<category><![CDATA[in vitro synergy assay limitations]]></category>
		<category><![CDATA[molecular and cell-line data encoding]]></category>
		<category><![CDATA[protein embeddings]]></category>
		<category><![CDATA[ProteinBERT]]></category>
		<category><![CDATA[representational learning in pharmacology]]></category>
		<category><![CDATA[SELFIES]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[transformer and BiLSTM integration]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201632</guid>

					<description><![CDATA[Researchers at the University of Qom have developed BT-Synergy, a hybrid BiLSTM-Transformer deep learning model that predicts synergistic cancer drug combinations with 0.8458 accuracy by integrating SELFIES molecular encodings with protein language model cell-line representations.]]></description>
										<content:encoded><![CDATA[<p>One of the most stubborn bottlenecks in modern oncology is not finding new drugs, but finding the right pairs of existing drugs that work better together than either does alone. Combination therapy can amplify treatment efficacy, slow the emergence of resistance, and reduce systemic toxicity, yet the number of possible drug pairings across thousands of compounds and hundreds of cancer cell lines grows so quickly that laboratory screening cannot keep pace. In vitro synergy assays remain slow, expensive, and labor-intensive, leaving most of the chemical space of possible combinations unexplored. A new study published in Discover Artificial Intelligence by Sahar Abbasi Rostami and Amir Lakizadeh of the University of Qom in Iran addresses this gap with a hybrid deep learning architecture called BT-Synergy, which the researchers report achieved an accuracy of 0.8458 on benchmark datasets for classifying synergistic drug combinations.</p>
<p>The central problem that BT-Synergy tackles is representational. Earlier computational approaches to synergy prediction, including AuDNNsynergy, SynPathy, and the widely used DeepSynergy model, relied on engineered molecular descriptors or structured multi-omics inputs. More recent frameworks such as SynergyX, DFFNDDS, SYNPRED, PRODeepSyn, DeepTraSynergy, and CFSSynergy jointly encode drug structures and cell-line characteristics using attention mechanisms, feature fusion modules, or protein-protein interaction networks. Yet many of these methods depend on SMILES string encodings or protein similarity matrices, which can miss higher-order chemical and biological dependencies. The Qom team argues that what is needed is an encoder that simultaneously understands the sequential grammar of a molecule and the long-range contextual relationships distributed across it, joined with a biologically grounded picture of the cell in which the interaction takes place.</p>
<p>To build that encoder, the researchers turned to SELFIES, a self-referencing molecular string representation that guarantees chemically valid outputs, unlike SMILES, which can produce structurally impossible sequences that corrupt downstream learning. Each drug is tokenized into SELFIES subunits, truncated or padded to a fixed length, and then processed by a hybrid module in which a bidirectional long short-term memory network is integrated directly into a Transformer block. In the final configuration, the BiLSTM actually replaces the conventional feed-forward sublayer inside the Transformer encoder. This design choice is deliberate: the Transformer&#8217;s multi-head self-attention excels at capturing long-range, non-local dependencies across a molecular sequence, while the BiLSTM contributes sequential inductive biases, reading the token stream in both forward and backward directions to preserve local structural patterns that attention alone can dilute.</p>
<p>The architecture was not chosen blindly. The team systematically compared variants, including a GRU-Transformer hybrid, a parallel configuration in which BiLSTM and Transformer pathways process the sequence simultaneously before element-wise fusion and layer normalization, and pure BiLSTM or pure Transformer baselines. Model depth also mattered: reducing the Transformer to two layers slightly degraded accuracy, while four or five layers inflated computational cost without commensurate gains. Three layers emerged as the optimal balance and were adopted in the final model. Ablation experiments confirmed that the hybrid design outperformed each single-architecture variant under identical conditions, supporting the premise that global contextual modeling and sequential dependency learning are complementary rather than redundant.</p>
<p>Equally important is how BT-Synergy represents the cellular context. Rather than relying on manually curated similarity networks, the model constructs cell-line embeddings from pre-trained protein language models. For each drug-cell-line instance, the researchers compile the union of proteins that are either annotated drug targets or observed as expressed in the relevant cancer cell line, drawing on drug-protein interaction and cell-line protein expression matrices inherited from the DeepTraSynergy dataset. This union typically spans between 54 and 1,479 proteins per sample, averaging roughly 354. Each protein&#8217;s canonical amino acid sequence is retrieved from UniProt and encoded with ProteinBERT from the TAPE suite, which produces dense vectors capturing local residue motifs and longer-range sequence dependencies. A learnable attention-based pooling layer then weights each protein embedding by its relevance, aggregating them into a single fixed-size cell representation that can be trained end-to-end with the rest of the network.</p>
<p>Fusion of the chemical and biological streams happens through a dual-fusion module designed to capture higher-order cross-modal interactions. The two drug embeddings are concatenated, then combined with the cell-line embedding via element-wise multiplication, addition, and subtraction. Multiplication emphasizes synergistic effects, addition captures complementary relationships, and subtraction highlights contrastive signals between molecular and cellular modalities. This interaction-aware scheme replaces naive concatenation, which a baseline variant confirmed is less effective. Because drug combinations are biologically symmetric, the team also applied order-invariance augmentation, generating mirrored training samples in which the two drugs are swapped. The augmentation paid off: across five cross-validation folds, predictions for original and reversed pairs showed a correlation of 0.9721 with a mean absolute difference of just 0.0511, indicating the model treats drug order symmetrically as biology demands.</p>
<p>Training and evaluation relied on two heterogeneous benchmarks. DrugCombDB contributed 69,436 drug-pair-cell-line observations spanning 764 compounds and 76 cancer cell lines, scored with the Zero Interaction Potency metric, whose values cluster tightly around zero. OncologyScreen, by contrast, contains 4,176 observations across 29 compounds and 21 cell lines, scored with the Loewe additivity model, which spans a far wider numerical range. To harmonize these divergent scales and combat class imbalance, the researchers adopted a quantile-based discretization: pairs in the upper quartile of each dataset&#8217;s score distribution were labeled synergistic, those in the lower quartile non-synergistic, and the ambiguous middle half was excluded. A sensitivity analysis comparing 50/50, 33/67, and 25/75 thresholds showed that including low-confidence pairs introduces substantial label noise. The strictest 25/75 configuration delivered the best trade-off, with accuracy of 0.8458, AUC-ROC of 0.9229, and F1 of 0.8422 on DrugCombDB.</p>
<p>The model also held up under punishing robustness protocols. In leave-one-drug-out evaluation, where all combinations involving held-out drugs are removed from training, BT-Synergy achieved an AUC-ROC of 0.8349; in leave-one-cell-line-out testing it reached 0.8357, suggesting genuine resilience to unseen drugs and biological contexts. When trained exclusively on DrugCombDB and tested on the entirely non-overlapping OncologyScreen dataset, the model retained encouraging predictive performance, providing preliminary evidence of cross-dataset transfer, though the authors caution that differing synergy-scoring systems limit strong generalizability claims. Interpretability analyses reinforced the picture: attention heatmaps revealed both globally distributed attention, integrating distant structural components, and sharply localized focus on chemically salient SELFIES symbols such as branching indicators, double-bond notations, and heteroatom tokens. In a token ablation experiment, masking the highest-attention fragments dropped one predicted synergy probability from 0.476 to 0.175, a striking decrease that suggests the model&#8217;s decisions hinge on specific molecular motifs, although the researchers stress that attention weights are proxy indicators rather than proven mechanisms.</p>
<p>Per-drug subgroup analysis added a biologically coherent note. Among the 29 OncologyScreen compounds, the model performed best on drugs with well-characterized mechanisms of action: 5-fluorouracil, an antimetabolite targeting thymidylate synthase, achieved a per-drug AUC-ROC of 0.9354, methotrexate, which inhibits dihydrofolate reductase, scored 0.9055, and doxorubicin, a DNA-targeting agent, reached 0.8850. Compounds with broad, pleiotropic, or poorly defined pharmacology fared noticeably worse. Across all 21 cancer cell lines, performance remained stable, with AUC-ROC values generally between 0.72 and 0.85, indicating the protein-informed cell representations prevent over-specialization to particular cellular backgrounds. Compared against DeepSynergy, GraphSynergy, NEXGB, DeepTraSynergy, and CFSSynergy, BT-Synergy delivered competitive performance on both benchmarks, an outcome the authors attribute to the combination of chemically valid SELFIES encoding, the BiLSTM-Transformer hybrid, and biologically informed protein embeddings.</p>
<p>The limitations are candidly acknowledged. Quantile-based binarization excludes half of the experimental spectrum, so reported performance reflects clearly defined observations rather than the full continuous distribution of synergy. Differences between ZIP and Loewe scoring constrain interpretations of transfer learning, and data sparsity plus the multi-target nature of complex biology mean performance will vary across contexts. The authors call for future validation using harmonized synergy measurements, continuous-label prediction, and additional independent pharmacological benchmarks, alongside extensions to multi-drug combinations and richer omics modalities. Even with those caveats, BT-Synergy demonstrates that fusing sequence-aware molecular encoders with protein language model embeddings can push drug synergy prediction toward the accuracy and robustness that precision oncology demands, and with the source code released on GitHub and both datasets publicly available, the framework is positioned to be tested, extended, and potentially deployed in the search for the next life-extending drug combination.</p>
<p><strong>Subject of Research:</strong> A hybrid deep learning model combining BiLSTM and Transformer architectures with protein embeddings to predict synergistic cancer drug combinations</p>
<p><strong>Article Title:</strong> A hybrid BiLSTM transformer model for drug synergy prediction</p>
<p><strong>Article References:</strong> Rostami, S. A., &amp; Lakizadeh, A. (2026). A hybrid BiLSTM transformer model for drug synergy prediction. <em>Discover Artificial Intelligence, 6</em>(1), Article 1182. <a href="https://doi.org/10.1007/s44163-026-02262-4" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02262-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02262-4" rel="noopener noreferrer">10.1007/s44163-026-02262-4</a></p>
<p><strong>Keywords:</strong> drug synergy prediction, BT-Synergy, BiLSTM, Transformer, SELFIES, ProteinBERT, combination therapy, cancer cell lines, DrugCombDB, deep learning, computational pharmacology, protein embeddings</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201632</post-id>	</item>
		<item>
		<title>New AI model maps the entire protein universe in a single view</title>
		<link>https://scienmag.com/new-ai-model-maps-the-entire-protein-universe-in-a-single-view/</link>
		
		<dc:creator><![CDATA[Drew Townsend]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 01:50:55 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI-driven understanding of cellular functions]]></category>
		<category><![CDATA[amino acid sequence]]></category>
		<category><![CDATA[amino acid sequence and 3D structure integration]]></category>
		<category><![CDATA[artificial intelligence in biochemistry]]></category>
		<category><![CDATA[bioinformatics tools for protein research]]></category>
		<category><![CDATA[CATH]]></category>
		<category><![CDATA[CLSS]]></category>
		<category><![CDATA[CLSS model for protein analysis]]></category>
		<category><![CDATA[contrastive learning]]></category>
		<category><![CDATA[deep learning for protein analysis]]></category>
		<category><![CDATA[ECOD]]></category>
		<category><![CDATA[evolution of protein families]]></category>
		<category><![CDATA[evolutionary biochemistry]]></category>
		<category><![CDATA[Institute of Science Tokyo]]></category>
		<category><![CDATA[interdisciplinary approaches in molecular biology]]></category>
		<category><![CDATA[mapping biological diversity]]></category>
		<category><![CDATA[protein classification]]></category>
		<category><![CDATA[protein embeddings]]></category>
		<category><![CDATA[protein evolution]]></category>
		<category><![CDATA[protein folding and molecular tasks]]></category>
		<category><![CDATA[protein language model]]></category>
		<category><![CDATA[protein structure]]></category>
		<category><![CDATA[protein structure prediction]]></category>
		<category><![CDATA[protein universe mapping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=200584</guid>

					<description><![CDATA[An international research team has developed CLSS, a protein language model that unites amino acid sequence and structural information into a single map of protein space, revealing evolutionary relationships across billions of years.]]></description>
										<content:encoded><![CDATA[<p>Every living cell depends on thousands of distinct protein families, each folding into precise three-dimensional shapes to carry out the molecular tasks that sustain life. Where all of this diversity came from, and how the different families relate to one another across billions of years of evolution, remains one of the deepest open questions in biochemistry. An international team of researchers, including the Earth-Life Science Institute (ELSI) at Institute of Science Tokyo, has now unveiled a new artificial intelligence tool that brings scientists closer to an answer by fusing the two fundamental languages of proteins—amino acid sequence and three-dimensional structure—into a single, unified representation. The work, published in Proceedings of the National Academy of Sciences, promises to transform how researchers explore the vast and largely unmapped protein universe.</p>
<p>The study was led by Professor Rachel Kolodny and PhD candidate Guy Yanai of the University of Haifa, together with Professor Nir Ben-Tal and graduate student Gabriel Axel of Tel Aviv University, and Specially Appointed Associate Professor Liam M. Longo of ELSI. Kolodny also spent five months as a visiting researcher at ELSI, developing methods to analyze the new model. Their creation, dubbed CLSS for Contrastive Learning Sequence-Structure, is a protein language model designed to overcome a stubborn problem that has limited previous computational approaches: the awkward relationship between what a protein&#8217;s sequence says and what its structure actually does.</p>
<p>Scientists have long organized proteins into hierarchical groups based on relatedness, much like the genus and species categories biologists use to classify organisms. These curated systems, such as the widely used ECOD and CATH databases, distill decades of expert knowledge. But with artificial intelligence now capable of generating &#8217;embeddings&#8217;—numerical representations in which proteins with similar properties receive nearby coordinates, like postal codes on a map—researchers can visualize relationships across millions of proteins at once, producing what the team calls a protein world map. The catch is that sequence and structure do not map neatly onto each other. Unrelated sequences can fold into similar shapes, while even identical sequences can sometimes adopt wildly different structures.</p>
<p>Most existing protein language models treat sequence and structure as separate worlds, processing one or the other independently. Even hybrid models that incorporate both kinds of data rarely place the sequence and the structure of the same protein at the same location on a global map, leaving researchers with two conflicting atlases of protein space. CLSS was engineered specifically to resolve this discordance. Using a machine learning strategy known as contrastive learning, the model is trained on pairs of protein sequences and their corresponding structures, learning to pull matching sequence-structure pairs together in the embedding space while pushing unrelated pairs apart.</p>
<p>The result is a single shared map in which a protein occupies essentially the same location whether the model is given its sequence or its structure. When benchmarked against other state-of-the-art protein language models, CLSS succeeded in producing a cohesive unified representation, something its predecessors could not achieve. Remarkably, the model&#8217;s maps closely reproduced the relationships recorded in the expert-curated ECOD and CATH classification systems, even though those classifications were never shown to the model during training. In direct classification tests, CLSS also performed strongly, demonstrating that merging sequence and structure information yields genuinely more informative protein representations.</p>
<p>Perhaps the most exciting feature of CLSS is its ability to handle fragments. Most protein language models require a complete sequence or structure to generate a meaningful embedding, but CLSS showed that short sequence fragments can in many cases be positioned meaningfully alongside full-length proteins and structures. This capability matters enormously for evolutionary studies, because small pieces of proteins have been repeatedly reused and rearranged throughout the history of life. Some fragments may even have served as the primordial building blocks from which the earliest protein domains were assembled, meaning that similar fragments appearing in otherwise unrelated proteins can hint at ancient evolutionary connections.</p>
<p>The maps produced by CLSS also revealed sweeping patterns across protein space that were previously difficult to see. When the researchers overlaid biological properties onto the maps, proteins associated with organic cofactors turned out to cluster in particular regions, while metal-binding proteins were scattered more broadly. Such patterns illustrate how global protein maps can serve not only as classification tools but as instruments for exploring the interplay between sequence, structure, function, and deep evolutionary history, potentially exposing large-scale patterns invisible to conventional pairwise comparison methods.</p>
<p>&#8216;This gives us a way to look at the protein universe through sequence and structure at the same time, rather than treating them as separate worlds,&#8217; said Longo. &#8216;What is particularly exciting for us is the possibility of using these maps to uncover large-scale evolutionary patterns that are difficult to recognise using conventional approaches.&#8217; The team ultimately envisions unified sequence-structure representations opening new frontiers in database searches, protein engineering, and the reconstruction of evolutionary trajectories—offering a fresh window onto how the staggering diversity of proteins found in life today emerged over nearly four billion years of evolution.</p>
<p><strong>Subject of Research:</strong> A contrastive-learning protein language model that unifies protein sequence and structure representations to map the protein universe</p>
<p><strong>Article Title:</strong> Uniting sequence and structure to map the protein universe</p>
<p><strong>Article References:</strong> Uniting sequence and structure to map the protein universe. (n.d.). <a href="https://www.eurekalert.org/news-releases/1142950" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> protein language model, CLSS, protein evolution, contrastive learning, protein structure, amino acid sequence, ECOD, CATH, protein embeddings, evolutionary biochemistry, protein classification, Institute of Science Tokyo</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">200584</post-id>	</item>
	</channel>
</rss>
