Tuesday, October 6, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Biology

AI Reads Protein Sequences to Find Domain Boundaries Without Any Structural Data

October 6, 2026
in Biology
Jason Bradley
By Jason Bradley Scienmag Editorial Profile - Structural Biology
Reading Time: 6 mins read
0
AI Reads Protein Sequences to Find Domain Boundaries Without Any Structural Data

AI Reads Protein Sequences to Find Domain Boundaries Without Any Structural Data

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every protein is a story told in chapters. Biologists call those chapters domains: compact, semi-independent units of three-dimensional structure that fold, function, and evolve largely on their own terms. Knowing where one domain ends and the next begins is not a trivial detail. It determines how researchers chop a protein into stable fragments for crystallization, how they design constructs for cryo-electron microscopy, and how they engineer new enzymes and therapeutics. Yet for the vast majority of proteins, especially those from bacteria that have never been crystallized or imaged, domain boundaries remain an educated guess. A new machine learning framework called Seq2DomML, developed by Alana Monks, Azam Asilian Bidgoli, and Michael D. L. Suits at Wilfrid Laurier University and published in BMC Bioinformatics, now offers a way to read those boundaries directly from the amino acid sequence, with no structural coordinates and no homologous domain assignments required as input.

The problem the Canadian team set out to solve is deceptively simple to state and notoriously hard to crack. When a protein’s structure has been solved experimentally, domain boundaries can be traced by eye or by algorithm through the folded chain. But structural determination remains slow, expensive, and unevenly distributed across the tree of life. For everything else, scientists lean on conserved domain annotations from resources such as Pfam, InterPro, and the NCBI Conserved Domain Database, or they resort to iterative truncation screening, building a panel of constructs and testing which ones express, fold, and behave well. Each of these strategies has blind spots. Annotation databases capture only domains that have already been characterized in other proteins, leaving novel or fast-evolving domains invisible. Truncation screening is empirical, laborious, and often burns months of laboratory time on constructs that fail. Seq2DomML was designed to complement these approaches by learning, from data, what a domain boundary looks like in the raw language of sequence.

The foundation of the framework is the Evolutionary Classification of Domains, or ECOD, a curated database that organizes protein domains by their evolutionary and structural relationships. The researchers converted ECOD’s structural domain annotations into binary residue-level labels: each amino acid in a bacterial protein sequence was marked as belonging to an annotated domain, or alternatively as part of a domain boundary, an inter-domain linker, or a terminal flanking region at the ends of the chain. This transformation turns a structural biology question into a supervised sequence labeling problem, the same class of task that powers speech recognition and named-entity extraction in language technology. Every residue becomes a token to be classified, and the model must learn the contextual cues that signal whether a given position sits inside a folded unit or in the flexible connective tissue between them.

Feeding that problem required an unusually rich encoding of each residue. The team initially described every amino acid with 2,389 features spanning amino acid identity, positional information along the chain, physicochemical properties, and interaction-related descriptors. That breadth is a deliberate hedge: no single property cleanly distinguishes a boundary residue from an interior one. Hydrophobicity, conservation patterns, local sequence composition, and position relative to the termini all carry partial signals, and a boundary might be legible only in their combination. But a feature space of nearly 2,400 dimensions also brings the classic curse of dimensionality, inviting overfitting and burying informative signals in noise. The researchers therefore invested heavily in feature selection, and this is where Seq2DomML acquires its most distinctive technical character.

Feature reduction proceeded in two stages. First, correlation-based pruning removed redundant descriptors that carried overlapping information. Then came the evolutionary step: a multi-objective genetic algorithm that treated the feature set itself as a genome to be optimized. In the initialization stage, the algorithm randomly assigned numbers of active features across a population and generated 100 unique binary feature vectors, each one a candidate recipe of which descriptors the model should see. Pairs of parents were then matched at random and combined through two-point crossover to produce two offspring per couple, with bit-flip mutation applied to the offspring to preserve diversity across the population. Each generation, every parent and offspring was trained and evaluated on three objectives simultaneously: the macro F1-score, the prediction rate for the minority boundary class, and the recall for that class. A non-dominated sorting algorithm, borrowed from multi-objective optimization theory, ranked individuals by their best achievable trade-offs across the three objectives, and the top individuals, equal in number to the original population, survived into the next generation. Crossover, mutation, and selection repeated for 50 generations, gradually sculpting a compact feature set that balanced overall accuracy against sensitivity to the rare and difficult boundary residues.

The classifier at the heart of the framework is a bidirectional long short-term memory network, a recurrent architecture well suited to sequential data. Bidirectionality matters here for a reason rooted in protein biology: the signal that a residue sits at a domain boundary depends on context from both sides of the chain, the tail end of one folded unit and the beginning of the next. An LSTM reading the sequence only left to right would encounter boundary evidence after the fact; a bidirectional model integrates information flowing in both directions, allowing the hidden state at each position to reflect the entire local sequence environment. The LSTM’s gating mechanisms, which learn to retain or forget information over long stretches of input, help the network bridge the dozens or hundreds of residues that may separate a boundary from the nearest obvious sequence landmark.

Evaluation was designed to guard against the most common pitfall in protein machine learning: leakage of evolutionary memory. The researchers removed from the held-out test set any sequences sharing at least 35 percent global sequence identity with the training set, ensuring the model could not simply recognize close relatives of proteins it had already seen. Against this stringent benchmark, Seq2DomML was compared with established annotation tools including InterPro, Pfam, and NCBI CDD, and it outperformed them on the classification metrics. It achieved a macro F1-score of 0.734, a figure that weights performance across all residue classes equally and therefore reflects genuine skill on the underrepresented boundary and linker categories, alongside a within-domain F1-score of 0.924, showing near-expert reliability at recognizing residues inside annotated domains. Its domain count mean absolute error of 0.952 was the second-lowest among the evaluated methods, trailing only Pfam.

Perhaps the most telling statistic is the mean domain Intersection over Union, a metric borrowed from computer vision that measures the spatial overlap between predicted and reference regions. Seq2DomML posted the highest mean IoU of the field, 0.846, indicating the strongest spatial agreement with the ECOD-derived domain regions among all methods tested. In practical terms, this means the model does not merely tally the right number of domains; it draws their edges in roughly the right places. For an experimentalist deciding where to cut a protein to isolate a single domain for structural studies, edge placement is everything. A prediction that captures the correct domain count but misplaces a boundary by thirty residues can yield a construct that aggregates, misfolds, or refuses to crystallize. The IoU result suggests Seq2DomML’s boundaries are tight enough to be actionable.

The authors are careful to position the tool as a complement rather than a replacement for the annotation infrastructure that has served biology for decades. Conserved domain databases rest on curated evolutionary evidence and remain authoritative where homologs exist. Seq2DomML’s value emerges precisely where that evidence runs out: in bacterial proteins with no experimentally determined structure, no recognizable domain homologs, and no reliable annotation to guide construct design. In those cases, the model’s residue-level partitions can help prioritize candidate splice regions, flagging the most promising places to truncate a protein before a single clone is made. Because it requires only sequence as input, it scales to the enormous and rapidly growing output of bacterial genome sequencing projects, where the majority of encoded proteins still await structural or functional characterization.

The broader significance of the work lies in its demonstration that structurally defined partitions are recoverable from sequence-derived features alone. The multi-objective evolutionary feature selection adds a methodological lesson for the field: when the class of interest is rare and the feature space is vast, optimizing a single accuracy metric can quietly sacrifice exactly the predictions that matter most. By forcing the feature set to balance macro F1, boundary prediction rate, and boundary recall simultaneously, the Wilfrid Laurier team built a model that stays honest about its hardest task. As structural biology pipelines increasingly depend on computational triage to decide which proteins and which fragments are worth experimental effort, frameworks like Seq2DomML point toward a future in which the chapter breaks of every protein, however novel, can be read straight from its sequence.

Subject of Research: Machine learning prediction of protein domain boundaries from bacterial amino acid sequences

Article Title: Seq2DomML: Protein domain boundary prediction from bacterial protein sequences using a machine learning framework

Article References: Monks, A., Asilian Bidgoli, A., & Suits, M. D. L. (2026). Seq2DomML: Protein domain boundary prediction from bacterial protein sequences using a machine learning framework. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06666-y

Image Credits: AI Generated

DOI: 10.1186/s12859-026-06666-y

Keywords: protein domains, domain boundary prediction, machine learning, LSTM, feature selection, ECOD, bioinformatics, bacterial proteins, construct design, sequence annotation, Pfam, InterPro

Cite Scienmag News

Jason Bradley. (October 6, 2026). AI Reads Protein Sequences to Find Domain Boundaries Without Any Structural Data. Scienmag. https://scienmag.com/ai-reads-protein-sequences-to-find-domain-boundaries-without-any-structural-data/

Jason Bradley. "AI Reads Protein Sequences to Find Domain Boundaries Without Any Structural Data." Scienmag, 6 October 2026, https://scienmag.com/ai-reads-protein-sequences-to-find-domain-boundaries-without-any-structural-data/. Accessed 6 October 2026.

Jason Bradley. "AI Reads Protein Sequences to Find Domain Boundaries Without Any Structural Data." Scienmag. October 6, 2026. https://scienmag.com/ai-reads-protein-sequences-to-find-domain-boundaries-without-any-structural-data/

Tags: bacterial proteinsbioinformaticsbioinformatics without structural dataComputational protein engineeringconstruct designcryo-electron microscopy construct designdeep learning for protein structuredomain boundary detection algorithmsdomain boundary predictionECODenzyme and therapeutic engineeringfeature selectionInterProLSTMMachine learningmachine learning in structural biologyPfamprotein domain boundary predictionprotein domainsprotein folding and evolutionprotein sequence analysis toolssequence annotationsequence-based protein analysisstructural genomics and protein annotation
Share26Tweet16
Previous Post

Cloud Computing Maps Soil Erosion Crisis in Ethiopia’s Cereal Heartland

Next Post

Ghana’s Traditional Rice Landraces Outshine Improved Varieties in Key Minerals, Study Finds

Related Posts

Gut Bacterium Found Migrating to the Liver May Drive Autoimmune Bile Duct Disease
Biology

Gut Bacterium Found Migrating to the Liver May Drive Autoimmune Bile Duct Disease

October 6, 2026
Ghana’s Traditional Rice Landraces Outshine Improved Varieties in Key Minerals, Study Finds
Biology

Ghana’s Traditional Rice Landraces Outshine Improved Varieties in Key Minerals, Study Finds

October 6, 2026
Why Some IVF Calves Grow Too Big: New Multi-Omics Study Offers Clues
Biology

Why Some IVF Calves Grow Too Big: New Multi-Omics Study Offers Clues

October 6, 2026
Emmy Noether Award Brings Insect Brain Evolution Researcher to Mainz
Biology

Emmy Noether Award Brings Insect Brain Evolution Researcher to Mainz

October 6, 2026
Blood Cells Rewound: Direct Neural Conversion Slowly Erases Epigenetic Age
Biology

Blood Cells Rewound: Direct Neural Conversion Slowly Erases Epigenetic Age

October 6, 2026
CT Scans Reveal Hidden Bone Secrets in the Long-Legged Buzzard
Biology

CT Scans Reveal Hidden Bone Secrets in the Long-Legged Buzzard

October 6, 2026
Next Post
Ghana’s Traditional Rice Landraces Outshine Improved Varieties in Key Minerals, Study Finds

Ghana's Traditional Rice Landraces Outshine Improved Varieties in Key Minerals, Study Finds

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Nanovaccine Built From Metal–Organic Frameworks Drives Tumours Into Full Retreat
  • Gut Bacterium Found Migrating to the Liver May Drive Autoimmune Bile Duct Disease
  • Ghana’s Traditional Rice Landraces Outshine Improved Varieties in Key Minerals, Study Finds
  • AI Reads Protein Sequences to Find Domain Boundaries Without Any Structural Data

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading