<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>protein sequence analysis &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/protein-sequence-analysis/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 28 Aug 2026 02:53:31 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>protein sequence analysis &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>PhaseOM Unifies Phase-Separation Analysis and Key-Residue Detection in One Framework</title>
		<link>https://scienmag.com/phaseom-unifies-phase-separation-analysis-and-key-residue-detection-in-one-framework/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Fri, 28 Aug 2026 02:53:27 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[biomolecular condensate formation]]></category>
		<category><![CDATA[biomolecular condensate formation mechanisms]]></category>
		<category><![CDATA[biomolecular condensates]]></category>
		<category><![CDATA[computational analysis of phase separation]]></category>
		<category><![CDATA[computational framework for phase separation analysis]]></category>
		<category><![CDATA[disordered protein regions]]></category>
		<category><![CDATA[droplet dynamics in cells]]></category>
		<category><![CDATA[faster detection of phase separation propensity]]></category>
		<category><![CDATA[identification of phase separation switches]]></category>
		<category><![CDATA[identifying phase separation switches]]></category>
		<category><![CDATA[key-residue detection in phase separation]]></category>
		<category><![CDATA[key-residue detection in proteins]]></category>
		<category><![CDATA[liquid-liquid phase separation]]></category>
		<category><![CDATA[membrane-free cellular compartments]]></category>
		<category><![CDATA[phase separation in cell biology]]></category>
		<category><![CDATA[protein disorder regions]]></category>
		<category><![CDATA[protein disordered regions]]></category>
		<category><![CDATA[protein role in phase separation]]></category>
		<category><![CDATA[protein sequence analysis]]></category>
		<category><![CDATA[residue-driven phase separation]]></category>
		<category><![CDATA[RNA involvement in phase separation]]></category>
		<guid isPermaLink="false">https://scienmag.com/phaseom-unifies-phase-separation-analysis-and-key-residue-detection-in-one-framework/</guid>

					<description><![CDATA[Liquid–liquid phase separation, the process by which molecules spontaneously condense into concentrated droplets without a surrounding membrane, has become one of biology’s most closely watched phenomena. A new computational framework called PhaseOM promises to make the process easier to dissect by determining not only whether a protein is likely to participate in phase separation, but [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Liquid–liquid phase separation, the process by which molecules spontaneously condense into concentrated droplets without a surrounding membrane, has become one of biology’s most closely watched phenomena. A new computational framework called PhaseOM promises to make the process easier to dissect by determining not only whether a protein is likely to participate in phase separation, but also what role it plays, which regions of its sequence are disordered, and which individual residues may help drive the formation of biomolecular condensates. The framework, described by researchers from multiple institutions in a study published online in the Journal of Advanced Research, could give scientists a faster way to identify the molecular “switches” controlling membrane-free compartments inside cells. Such compartments include nucleoli, P bodies and other condensates that organize biochemical reactions without the walls of a traditional organelle.</p>
<p>The biological principle behind the work is deceptively simple. In a liquid–liquid phase separation event, proteins, RNA and other molecules separate from the surrounding cellular fluid much as oil separates from water, creating a dense phase enriched in selected components and a dilute phase containing the remainder. Unlike a crystal or a permanently aggregated protein, a liquid condensate can remain dynamic: molecules enter and leave, droplets fuse, and interactions can be rapidly remodeled. These properties allow cells to concentrate enzymes, RNA-processing factors and DNA-repair proteins precisely where they are needed. The molecular forces involved are generally weak when considered individually, but they become powerful when repeated across many interaction sites. Electrostatic attraction, cation–π and π–π interactions, transient binding motifs and flexible protein segments can collectively create the multivalent network needed for condensation.</p>
<p>A central feature of many phase-separating proteins is the intrinsically disordered region, or IDR. Unlike a folded domain, which adopts a relatively stable three-dimensional structure, an IDR samples a constantly shifting ensemble of conformations. Its flexibility and composition can make it especially effective at forming numerous temporary contacts with other proteins or nucleic acids. IDRs are often enriched in charged, polar or interaction-prone residues, although no single sequence pattern explains every condensate. This diversity has made prediction difficult. Earlier computational tools typically searched for narrow signatures, such as prion-like amino-acid composition, aromatic interaction potential or the presence of arginine and tyrosine. More recent systems have incorporated evolutionary conservation, structural predictions, protein–protein interactions, imaging data and machine-learning embeddings, but many still treat all phase-separating proteins as if they perform the same job.</p>
<p>PhaseOM was designed to address that missing distinction. Its first task is to classify a candidate protein as either a scaffold or a client. Scaffolds are the core components that initiate, organize or maintain a condensate. They provide much of the interaction network that gives the assembly its physical integrity. Clients, by contrast, are recruited into an existing condensate and become selectively concentrated there, but generally do not serve as the principal drivers of its formation. The difference matters because two proteins can both be found inside the same droplet while contributing to it in fundamentally different ways. Misclassifying a client as a scaffold could lead researchers toward the wrong experiments, obscure the mechanism of condensation and complicate efforts to target disease-associated condensates.</p>
<p>The framework combines several machine-learning strategies in a sequential workflow. For scaffold detection, PhaseOM uses embeddings generated by ProtT5-XL-U50, a protein-language model that represents sequence information numerically, together with a graph-attention network. The graph incorporates predicted spatial relationships between amino acids, connecting residues whose alpha-carbon atoms lie within 10 angstroms of one another. This allows the model to combine sequence-derived information with a representation of local three-dimensional proximity. The scaffold classifier achieved an area under the receiver operating characteristic curve, or AUC, of 0.9954. An AUC of 1.0 indicates perfect separation between classes, although performance measured on curated data does not guarantee identical accuracy on proteins outside the training distribution.</p>
<p>Client proteins are evaluated with a separate ensemble classifier built from 22 optimized physicochemical and structural features. These features are intended to capture properties such as composition, charge, flexibility and predicted structural behavior that may distinguish recruited proteins from condensate-driving scaffolds. The client model reached an AUC of 0.9474. The system applies a probability threshold of 0.5 to both categories; if a protein exceeds that threshold for both scaffold and client, the class with the higher probability is selected. This arrangement reflects the biological reality that classification is not always clean. Some proteins may participate in more than one condensate, switch roles depending on cellular context or act as a scaffold in one environment and a client in another.</p>
<p>After assigning a functional category, PhaseOM searches for IDR-associated phase-separation regions. Its disorder model uses 17 features and achieved an AUC of 0.9835. The researchers then apply a fourth model to locate key residues within those regions. This residue-level predictor is a multilayer perceptron equipped with multi-head self-attention and processes 1,058-dimensional input data. Attention mechanisms allow a model to weigh relationships among different positions in a sequence rather than treating every residue as an isolated feature. The key-residue model produced an AUC of 0.8092, lower than the framework’s protein- and region-level tasks but potentially useful for prioritizing residues for laboratory testing. In practice, the output could direct mutagenesis experiments toward a small number of candidate positions rather than requiring researchers to alter an entire disordered region.</p>
<p>The predictions also exposed biochemical differences between the two functional classes. Client proteins tended to adopt more expanded conformations and contained higher frequencies of cysteine and histidine, residues that may contribute to π-related or electrostatic interactions in particular molecular settings. Their IDRs were associated with an excess of negative charge and increased flexibility. Scaffold proteins, in contrast, showed greater representation of tyrosine and arginine and more complex interaction patterns. These observations do not imply that a single amino acid determines a protein’s role. Phase separation depends on the combined behavior of many residues, the presence of RNA or other binding partners, post-translational modifications, concentration, temperature, salt conditions and the cellular environment. Instead, the patterns provide statistical clues that can be integrated into a broader mechanistic model.</p>
<p>The study reports that PhaseOM outperformed Seq2Phase on independent tests, improving client-protein AUC by 0.21 and scaffold-protein AUC by more than 0.07. The authors attribute much of the gain to the framework’s integrated structural and ensemble-based architecture, which connects functional classification to disorder mapping and residue identification. They also tested the system on alpha-synuclein isoforms, proteins of particular interest because abnormal assemblies of alpha-synuclein are linked to neurodegenerative disease. Residue-level validation showed 76 to 91 percent accuracy for IDR assignments, with stronger agreement in longer disordered regions. Predicted scaffold probabilities exceeded 0.91 across the isoforms, consistent with the biological consensus used in the analysis. These results suggest that the system can reproduce known features, although experimental validation across a wider range of proteins will be essential.</p>
<p>The potential applications extend beyond cataloguing proteins. Aberrant condensates have been associated with cancer, neurodegeneration and failures in RNA processing, and their formation may sometimes precede irreversible aggregation. A tool that identifies the residues and regions most responsible for condensation could help researchers test whether a disease-linked mutation changes a protein’s scaffold activity, client recruitment or interaction landscape. It might also support the design of molecules that selectively disrupt pathological condensates while preserving normal ones. PhaseOM is not itself a treatment and cannot establish causation from sequence alone. Its predictions require biochemical and cellular experiments, particularly because phase behavior is strongly context-dependent. Even so, by organizing the analysis into a hierarchy—from scaffold or client, to disordered region, to critical residue—the framework offers a practical route from a raw protein sequence to testable molecular hypotheses. The researchers have made the model architecture and feature configurations available through a public GitHub repository, potentially allowing other groups to evaluate and extend the approach as the rapidly expanding field of biomolecular condensates moves toward more precise, residue-level biology.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Phase separation analysis and key residue detection in proteins</p>
<p><strong>Article Title:</strong> PhaseOM: an integrated multi-task framework for phase separation analysis and key residue detection</p>
<p><strong>Article References:</strong> Xu, L., Zhou, S., Ran, Z., Qin, X., Liu, T., Zou, Q., Li, F., &amp; Jia, C. (2026). PhaseOM: an integrated multi-task framework for phase separation analysis and key residue detection. <em>Journal of Advanced Research</em>. <a href="https://doi.org/10.1016/j.jare.2026.08.052" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/j.jare.2026.08.052</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.jare.2026.08.052" target="_blank" rel="noopener noreferrer">10.1016/j.jare.2026.08.052</a></p>
<p><strong>Keywords:</strong> liquid–liquid phase separation, biomolecular condensates, intrinsically disordered regions, scaffold proteins, client proteins, machine learning, key residues, protein prediction, PhaseOM, alpha-synuclein</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">183280</post-id>	</item>
		<item>
		<title>DC-BiGAN-IR Predicts Insulin Receptors Using Protein Language Models and Wavelet-Enhanced PSSM</title>
		<link>https://scienmag.com/dc-bigan-ir-predicts-insulin-receptors-using-protein-language-models-and-wavelet-enhanced-pssm/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 02:32:29 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[amino acid sequence representation]]></category>
		<category><![CDATA[biomedical data integration]]></category>
		<category><![CDATA[computational protein annotation]]></category>
		<category><![CDATA[deep learning in bioinformatics]]></category>
		<category><![CDATA[generative adversarial networks in biology]]></category>
		<category><![CDATA[Insulin receptor prediction]]></category>
		<category><![CDATA[metabolic disease mechanisms]]></category>
		<category><![CDATA[molecular target identification]]></category>
		<category><![CDATA[protein language models]]></category>
		<category><![CDATA[protein sequence analysis]]></category>
		<category><![CDATA[transmembrane receptor modeling]]></category>
		<category><![CDATA[wavelet-enhanced PSSM]]></category>
		<guid isPermaLink="false">https://scienmag.com/dc-bigan-ir-predicts-insulin-receptors-using-protein-language-models-and-wavelet-enhanced-pssm/</guid>

					<description><![CDATA[Insulin receptor prediction has entered a new computational phase with the proposed DC-BiGAN-IR framework, a deep-learning system designed to identify and characterize insulin receptors from protein sequences. The method combines several advanced technologies that are rarely integrated in a single prediction pipeline: an ensemble of pre-trained protein language models, an integrated discrete wavelet transformation, a [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Insulin receptor prediction has entered a new computational phase with the proposed DC-BiGAN-IR framework, a deep-learning system designed to identify and characterize insulin receptors from protein sequences. The method combines several advanced technologies that are rarely integrated in a single prediction pipeline: an ensemble of pre-trained protein language models, an integrated discrete wavelet transformation, a tri-blocked position-specific scoring matrix, and a dual-channel bidirectional generative adversarial network. Together, these components aim to capture the chemical, evolutionary, and structural signals hidden inside amino-acid sequences. The approach arrives at a moment when researchers are seeking faster ways to annotate proteins, understand metabolic disease mechanisms, and identify molecular targets without relying entirely on slow and expensive laboratory experiments. By translating biological sequences into multiple complementary digital representations, DC-BiGAN-IR offers a strategy for turning the vast and still largely unexplored protein universe into actionable biomedical information.</p>
<p>The insulin receptor is a particularly important target because it sits at the center of glucose regulation. It is a transmembrane receptor tyrosine kinase that responds to insulin and triggers a cascade of intracellular events controlling glucose uptake, lipid metabolism, protein synthesis, and cell growth. When insulin-receptor signaling is weakened or disrupted, the consequences can include insulin resistance, type 2 diabetes, metabolic syndrome, and other chronic disorders. Although the receptor is well studied, distinguishing insulin receptors and related proteins from sequence data remains a demanding computational problem. Protein sequences may share partial similarities while performing very different biological functions, and evolutionary changes can obscure the motifs that define receptor identity. A reliable predictor must therefore recognize more than short sequence patterns. It must understand broader relationships involving residue composition, evolutionary conservation, local sequence order, and long-range dependencies.</p>
<p>DC-BiGAN-IR addresses this challenge through an ensemble of pre-trained protein language models. These models are trained on enormous collections of protein sequences and learn statistical representations that reflect how amino acids interact across biological evolution. In much the same way that language models learn relationships between words, protein language models learn relationships between residues and sequence regions. Their internal representations can capture information associated with secondary structure, domain organization, functional motifs, and evolutionary constraints, even when explicit structural data are unavailable. Using an ensemble rather than a single model allows the system to combine different learned perspectives. Each model may emphasize distinct patterns, and their fused outputs can provide a richer feature space for downstream classification. This is especially valuable for membrane proteins, whose functional signatures may be distributed across several regions rather than concentrated in one easily recognizable motif.</p>
<p>The second major component is integrated discrete wavelet transformation, a signal-processing technique adapted here for biological sequence analysis. A protein sequence can be converted into numerical signals using features such as amino-acid physicochemical properties, residue frequencies, or model-derived embeddings. The wavelet transform then decomposes these signals into components operating at different scales. Broad, low-frequency components can represent gradual trends across a sequence, while high-frequency components may reveal abrupt changes, local motifs, or boundaries between functional regions. Unlike conventional methods that examine sequence data only in the original representation, wavelet analysis can expose patterns that are difficult to see through direct inspection. Integrating these multiscale features with language-model embeddings may help the predictor distinguish global architecture from local biochemical signals, improving its ability to identify proteins that belong to the insulin-receptor family.</p>
<p>Evolutionary information enters the framework through a tri-blocked position-specific scoring matrix, commonly known as PSSM. A PSSM is generated by comparing a query sequence with related proteins and estimating how frequently particular amino acids appear at each position. Conserved positions receive strong statistical signatures, while variable positions provide information about regions that tolerate evolutionary change. In DC-BiGAN-IR, the PSSM information is divided into three blocks, creating separate feature groups that can preserve different aspects of evolutionary preference and sequence context. This tri-blocked design is intended to prevent the rich but high-dimensional PSSM signal from being compressed into a single undifferentiated representation. Instead, the model can process multiple evolutionary views and compare them with features obtained from language models and wavelet decomposition. The result is a multimodal description of each protein, combining what the sequence looks like, how it varies across evolution, and how its patterns unfold at different scales.</p>
<p>At the heart of the architecture is a dual-channel bidirectional generative adversarial network. Generative adversarial networks traditionally consist of a generator and a discriminator engaged in a competitive learning process. The generator attempts to produce realistic synthetic feature representations, while the discriminator tries to distinguish artificial features from genuine examples. Through this contest, the system can learn a more informative decision boundary, particularly when training data are limited or unevenly distributed. The bidirectional design extends the concept by allowing information to move in both forward and reverse directions through the sequence representation. This can help capture dependencies that begin near the amino-terminal region but influence residues much farther toward the carboxyl terminus, as well as the reverse relationship. The dual-channel structure separates or complements distinct feature streams, allowing sequence-derived and evolutionary or transformed signals to be processed before they are jointly interpreted.</p>
<p>This architecture could be especially useful because protein datasets often contain a serious imbalance between positive and negative examples. Confirmed insulin receptors may be relatively scarce compared with unrelated proteins, and the available sequences may not represent the full diversity found across species. A model trained on imbalanced data can become biased toward the majority class, producing apparently strong accuracy while missing biologically important receptors. Adversarial learning may help enrich the minority-class representation by generating plausible feature patterns, while the combined channels can preserve independent evidence from different sources. However, synthetic data do not automatically equal biological truth. Any generated representation must be evaluated against experimentally verified sequences, independent test sets, and external databases. Performance should also be measured using sensitivity, specificity, precision, recall, Matthews correlation coefficient, and area under the precision-recall curve, rather than accuracy alone.</p>
<p>The potential impact extends beyond annotation. A faster and more accurate insulin-receptor predictor could assist researchers in screening newly sequenced organisms, prioritizing candidate proteins for laboratory testing, and studying how receptor families evolved. It could also support investigations into mutations that alter receptor activity, contribute to drug resistance, or affect the molecular pathways associated with diabetes. In pharmaceutical research, computational filtering can reduce the number of sequences requiring experimental characterization and help identify related receptors for comparative analysis. The same design principles may be transferable to other protein families, including transporters, enzymes, immune receptors, and viral proteins. That broader adaptability is one reason hybrid architectures are attracting attention: biological function is encoded at multiple levels, and a single representation may overlook critical evidence.</p>
<p>Yet DC-BiGAN-IR should be understood as a predictive tool, not a replacement for experiments. Computational models can be influenced by the quality of their training data, the choice of negative examples, the evolutionary databases used to create PSSMs, and the possibility that benchmark sequences are too closely related. Data leakage, in which similar sequences appear in both training and testing collections, can make a model seem more capable than it is in real-world use. Independent validation on geographically, taxonomically, and experimentally diverse datasets will be essential. Researchers will also need to determine whether the system can explain its predictions by identifying influential residues, conserved regions, or sequence segments associated with receptor classification. Interpretability matters because a prediction that cannot be biologically examined is difficult to translate into a laboratory hypothesis.</p>
<p>The arrival of DC-BiGAN-IR reflects a larger transformation in molecular biology, where artificial intelligence is moving from simple pattern recognition toward integrated biological reasoning. By combining learned protein representations with wavelet-based multiscale analysis, evolutionary scoring, and adversarial feature generation, the framework attempts to read protein sequences as layered biological messages rather than strings of isolated characters. Its promise lies in this convergence: language models provide contextual knowledge, PSSMs contribute evolutionary memory, wavelets reveal hidden structure across scales, and the dual-channel bidirectional network unites these signals into a single prediction system. If rigorous external testing confirms its effectiveness, the approach could become a valuable component of computational protein annotation and metabolic-disease research. For now, its most important message is clear: the next breakthroughs in insulin biology may emerge not only from the laboratory bench, but also from algorithms capable of decoding the complex language of proteins.</p>
<p><strong>Subject of Research</strong>: Computational prediction and identification of insulin receptor proteins using deep learning and protein-sequence analysis.</p>
<p><strong>Article Title</strong>: DC-BiGAN-IR: Prediction of Insulin Receptor Using an Ensemble of Pre-Trained Protein Language Models and Integrated Discrete Wavelet Transformation with Tri-Blocked PSSM in a Dual-Channel Bidirectional Generative Adversarial Network</p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>Keywords</strong>: insulin receptor, protein language models, deep learning, generative adversarial network, BiGAN, discrete wavelet transformation, PSSM, protein sequence analysis, bioinformatics, computational biology, diabetes research</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">181468</post-id>	</item>
	</channel>
</rss>
