For decades, the molecular blueprints of life have been drawn mostly for a handful of model organisms: humans, mice, fruit flies, and yeast. Meanwhile, the animals that actually feed the world—chickens, ducks, pigs, cattle, and sheep—have remained largely in the dark at the level of their protein interaction networks. That gap has now been substantially closed. A team of Chinese researchers has unveiled AASIA, the Agricultural Animals Structural Interactome Atlas, the first comprehensive database of structure-informed protein–protein interactions covering five of the world’s most economically important livestock species. The resource, described in the Journal of Advanced Research, contains 410,421 high-confidence predicted interactions, each accompanied by a computationally generated three-dimensional model of the protein complex involved.
Why does this matter? Proteins almost never work alone. They form intricate interaction networks that govern everything from immune responses to reproduction, growth, and product quality. Understanding these networks at a systems level is essential for improving traits such as disease resistance, fertility, and meat or egg quality—the cornerstones of food security and sustainable farming. Yet existing interaction databases like STRING, BioGRID, IntAct, and MINT offer only sparse coverage of agricultural species, and they typically lack the structural annotations needed to identify binding interfaces, functional hotspots, or the molecular consequences of mutations. Without such detail, linking a genetic variant to a measurable trait—an essential step in molecular breeding—remains a matter of guesswork.
The new database was built through an ambitious integrative pipeline that combines three complementary prediction strategies. The first, interolog mapping, exploits evolutionary conservation: if two proteins in a model organism are known to interact, their homologs in a farm animal are likely to interact as well. The team assembled a rigorously filtered library of experimentally supported interactions from seven model organisms, including human, mouse, fly, worm, yeast, E. coli, and Arabidopsis, retaining only high-confidence interactions scored by the HIPPIE scheme. The second strategy, domain–domain interaction inference, predicts interactions based on conserved protein domain pairings drawn from the 3did catalog and expanded with an expectation–maximization algorithm applied to curated interaction data.
The third and most modern component is a deep learning model. The researchers retrained PPITrans, a transformer-based architecture built on the ProtT5 protein language model, using a curated set of 44,035 high-confidence human interactions as positive samples and 430,370 negative pairs. Training on an NVIDIA A30 GPU with the AdamW optimizer, the model learned sequence-derived interaction patterns that generalize across species. The scores from all three methods were then fused through a logistic regression model trained on a mouse benchmark dataset, producing a single interaction probability for every protein pair. On independent testing, the integrated model achieved an area under the ROC curve of 0.914 and an area under the precision–recall curve of 0.679, outperforming each individual component. Interactions were classified as high-confidence only above a stringent threshold calibrated at a 0.1 percent false positive rate.
The resulting species-specific networks are striking in scale. Cattle lead the pack with 173,331 predicted interactions covering 34.0 percent of the proteome, followed by pig with 90,987 interactions, chicken with 61,593, duck with 43,886, and sheep with 40,624. In total, these interactomes span proteomes at coverage rates ranging from 19.5 to 34.0 percent—a dramatic expansion of molecular data for organisms that have historically been interaction-data deserts. Cross-validation against independent resources lent credibility to the predictions: the cattle network shared 8,127 interactions with a D-SCRIPT-based dataset and captured 207 of 632 experimentally solved cattle complex structures in Interactome3D, both overlaps highly significant by Fisher’s exact test.
What truly distinguishes AASIA, however, is its structural dimension. The team used ESMFold, a cutting-edge structure prediction system powered by the ESM-2 protein language model, to generate three-dimensional models for every protein monomer across the five proteomes and, in multimer mode, for all 410,000-plus interacting pairs. Unlike earlier approaches requiring multiple sequence alignments or structural templates, ESMFold predicts structures directly from amino acid sequences, enabling proteome-scale modeling at unprecedented speed. Most monomer models showed high confidence, with pLDDT scores predominantly above 70. Each complex was then annotated with interface residues defined by geometric contact and hotspot residues identified through computational alanine scanning with the FoldX force field, which flags positions whose mutation would destabilize the interface by at least 0.5 kilocalories per mole.
To demonstrate the database’s power, the researchers dove deep into the chicken interactome. Its network topology follows classic small-world, scale-free architecture, with an average of 24 interactions per protein, a mean shortest path length of 3.3, and a clustering coefficient of 0.63—far higher than randomized networks of the same size. Fifty predicted chicken interactions overlapped significantly with experimentally supported ones, and predicted pairs showed Gene Ontology semantic similarity comparable to known interactions and far above random pairs. A flagship case study examined the actin-capping protein assembly: the predicted CAPZA1–CAPZB heterodimer aligned with the experimentally solved crystal structure with a root-mean-square deviation of just 1.6 angstroms, and the predicted binding interfaces for both subunits with ACTIN matched mutagenesis-verified actin-binding motifs at the C-termini of the CapZ subunits.
Perhaps the most compelling illustration of AASIA’s practical value involves a predicted interaction between XRCC5, a DNA repair factor, and H2AC39, a histone variant, in chicken. The interaction was independently supported by all three prediction methods and scored 0.94 overall. Remarkably, the predicted interface overlaps a known quantitative trait locus associated with eggshell strength. A residue at that locus, arginine 10 of XRCC5, sits directly at the interaction interface and was flagged as a putative hotspot: FoldX predicted that mutating it to alanine would destabilize the complex by 1.09 kilocalories per mole, and the database’s variant effect module predicted a corresponding drop in the interaction score. This offers a testable, structure-based hypothesis that genetic variation at this position could alter chromatin dynamics during eggshell mineralization—precisely the kind of genotype-to-phenotype bridge molecular breeders have long sought.
The platform itself is designed for broad usability, built on a Vue.js frontend, Express.js backend, and MySQL metadata management, with Cytoscape.js for network visualization and Mol* and NGL Viewer for interactive 3D structure exploration. Users can search by UniProt, Ensembl, or Entrez identifiers, gene names, or keywords; browse protein and interaction pages with functional annotations from Gene Ontology and KEGG; and, crucially, submit their own protein sequences to the Predict module for de novo interaction inference or quantitative assessment of mutation effects on interaction scores. This last capability turns AASIA from a static repository into an active hypothesis-generating engine.
The authors acknowledge limitations: public databases mix direct binary interactions with association-level evidence, and ESMFold does not explicitly model intrinsically disordered regions that mediate many interactions. Still, performance on a stricter binary-only test subset was even stronger, with an AUROC of 0.928. Looking forward, the team plans to expand AASIA with host–pathogen interaction data—potentially illuminating conserved interfaces exploited by avian influenza and African swine fever viruses—and to integrate transcriptomic, epigenomic, and genome-wide variant data. By transforming raw omics information into spatially resolved, residue-level insight, AASIA positions itself as a foundational tool for the next generation of livestock genomics and biotech breeding.
Subject of Research: A structure-informed protein–protein interactome database for five agricultural animal species
Article Title: AASIA: A comprehensive protein structural interactome database for agricultural animals
Article References: Li, J., Yuan, M., Jiang, L., Li, D., Yin, Z., Zhu, F., Shi, W., Hou, Z., & Zhang, Z. (2026). AASIA: A comprehensive protein structural interactome database for agricultural animals. Journal of Advanced Research. https://doi.org/10.1016/j.jare.2026.09.015
Image Credits: AI Generated
DOI: 10.1016/j.jare.2026.09.015
Keywords: AASIA, protein-protein interactions, agricultural animals, structural biology, ESMFold, deep learning, interolog mapping, molecular breeding, livestock genomics, variant effect prediction, disease resistance, food security
Cite Scienmag News
Jason Bradley. (September 22, 2026). New Protein Structure Database Maps 410,000 Interactions Across Farm Animals. Scienmag. https://scienmag.com/new-protein-structure-database-maps-410000-interactions-across-farm-animals/
Jason Bradley. "New Protein Structure Database Maps 410,000 Interactions Across Farm Animals." Scienmag, 22 September 2026, https://scienmag.com/new-protein-structure-database-maps-410000-interactions-across-farm-animals/. Accessed 22 September 2026.
Jason Bradley. "New Protein Structure Database Maps 410,000 Interactions Across Farm Animals." Scienmag. September 22, 2026. https://scienmag.com/new-protein-structure-database-maps-410000-interactions-across-farm-animals/

