Deep in the genomes of fungi lies a vast, unexplored library of molecular switches that control which genes are turned on and when. These switches, known as transcriptional activators, carry short protein segments called activation domains that recruit the cellular machinery needed to read a gene. Now, a team of researchers led by scientists at the University of California, Berkeley, has combined machine learning with high-throughput experiments to map these activation domains across the entire fungal kingdom, publishing their findings in Genome Biology. The work delivers the first functional annotation for thousands of proteins from non-model fungi and demonstrates a powerful new strategy for building predictive models that do not collapse when confronted with evolutionary diversity.
The challenge the team set out to solve is a familiar one in computational biology. Predictive models trained on data from model organisms such as baker’s yeast often perform poorly when applied to species that sit far away on the tree of life. Existing datasets are heavily biased toward a handful of well-studied organisms, and this overrepresentation causes algorithms to fail in evolutionary studies of non-model species. The problem is compounded for transcriptional activators specifically, because their activation domains are intrinsically disordered, meaning they lack a fixed three-dimensional structure, and they are poorly conserved across evolution. Comparative genomics, which relies on sequence similarity to transfer functional annotations between species, simply cannot see these domains clearly.
To break through this barrier, the researchers developed a hybrid framework that pairs high-capacity molecular assays with active learning, an iterative machine learning strategy in which the model itself helps decide which experiments to run next. The centerpiece of their approach is ADhunter, a regression model designed to identify transcriptional activators and quantify how strong they are. In head-to-head benchmarking, ADhunter outperformed state-of-the-art algorithms at both tasks, distinguishing genuine activators from inactive sequences and ranking them by the intensity of the gene expression they drive.
What makes the study remarkable is the sheer scale of the evolutionary space it explores. The team deployed ADhunter’s uncertainty estimates to guide sampling across 7,842,516 proteins drawn from 2,400 fungal genomes. Rather than measuring every candidate, an impossible task even with modern robotics, the model flagged the sequences about which it was least confident, ensuring that each round of experiments delivered maximum informational value. This active learning loop, in which predictions generate experiments and experimental results refine predictions, allowed the researchers to efficiently probe the darkest corners of fungal protein space.
The experimental campaign paid off on an extraordinary scale. The team functionally characterized 9,836 activation domains from 1,071 fungal genomes, representing a 15.5-fold expansion in genome representation compared with existing datasets. For 3,416 proteins in non-model fungi, these measurements provide the first functional annotation ever recorded. Each of these domains was tested in a high-throughput molecular assay that quantifies its ability to promote gene expression, generating precisely calibrated training data rather than the noisy, indirect labels that have limited previous efforts.
Technically, the workflow begins with a training set of 17,609 protein tiles from fungal and plant proteins, each tile carrying a measured activity value. ADhunter learns the sequence-to-activity relationship from this initial dataset and then applies uncertainty sampling to select new tiles from the fungal proteomes for experimental testing. The measurements from the active learning round are harmonized onto a common activity distribution with the initial data, allowing the model to be retrained on a dataset that spans a far wider swath of evolutionary history. This closed loop of prediction, measurement, and retraining is what gives the framework its power to generalize beyond the organisms that dominated earlier datasets.
Beyond simply cataloging activators, the study probes the biophysical logic that underlies activation domain function. Interpretability analysis of ADhunter, which examines what sequence features the model relies on when making predictions, aligned with established biophysical models of how activation domains work. More intriguingly, the analysis revealed novel protein codes that had been underrepresented in previous studies, suggesting that the expanded evolutionary sampling uncovered sequence patterns that model-organism datasets had never captured. Because activation domains are disordered and rapidly evolving, these newly detected codes may represent alternative biochemical solutions that fungi have repeatedly invented to accomplish the same regulatory task.
The implications extend well beyond fungal biology. Activation domains are essential tools in synthetic biology and metabolic engineering, where researchers stitch together genetic circuits to coax microbes into producing fuels, medicines, and materials. A model that works across the fungal kingdom gives engineers a vastly larger palette of regulatory parts to choose from, including domains optimized by evolution for organisms and conditions far removed from the laboratory. The work was conducted in part through the Department of Energy’s Joint BioEnergy Institute, reflecting its relevance to bioenergy applications, and the authors include prominent synthetic biologists such as Jay Keasling alongside computational biologist Max Staller and plant and microbial biologist Patrick Shih.
Perhaps the most important message of the paper is methodological. The results highlight the importance of sampling from non-model organisms to build evolutionarily robust functional genomics models, and the framework offers a general strategy that can be adapted to other classes of disordered proteins and other branches of life. As high-throughput assays become cheaper and machine learning models grow more capable, the bottleneck in biology is increasingly the design of informative experiments. Active learning addresses exactly this bottleneck by letting the model point to where ignorance is greatest, converting limited experimental budgets into maximal biological insight.
The study also serves as a corrective to a long-standing habit in genomics. For decades, functional annotation has flowed outward from a few model organisms, leaving the overwhelming majority of life’s diversity annotated by guesswork or silence. By demonstrating that a model trained with evolutionarily guided sampling can discover and characterize functional elements that comparative methods miss entirely, the Berkeley team has made a strong case that the future of functional genomics lies in deliberately embracing diversity rather than avoiding it. The 9,836 newly characterized activation domains, and the model that found them, are now available as an open resource, inviting researchers across microbiology, evolution, and bioengineering to put the fungal kingdom’s hidden switches to work.
Subject of Research: Active learning-guided discovery and functional characterization of fungal transcriptional activation domains
Article Title: Active learning enables evolutionary discovery and characterization of fungal transcriptional activators
Article References: Waldburger, L., Nisonoff, H., Zintel, M. A., Kirkpatrick, L. D., Lam, A. W. Y., Lanclos, N., Keasling, J. D., Staller, M. V., & Shih, P. M. (2026). Active learning enables evolutionary discovery and characterization of fungal transcriptional activators. Genome Biology. https://doi.org/10.1186/s13059-026-04269-7
Image Credits: AI Generated
DOI: 10.1186/s13059-026-04269-7
Keywords: active learning, machine learning, transcriptional activators, activation domains, fungal genomics, intrinsically disordered proteins, functional genomics, ADhunter, non-model organisms, evolutionary biology, synthetic biology, uncertainty sampling
Cite Scienmag News
Juliet Wilcox. (September 26, 2026). Machine Learning Hunts Down Hidden Gene Switches Across the Fungal Tree of Life. Scienmag. https://scienmag.com/machine-learning-hunts-down-hidden-gene-switches-across-the-fungal-tree-of-life/
Juliet Wilcox. "Machine Learning Hunts Down Hidden Gene Switches Across the Fungal Tree of Life." Scienmag, 26 September 2026, https://scienmag.com/machine-learning-hunts-down-hidden-gene-switches-across-the-fungal-tree-of-life/. Accessed 26 September 2026.
Juliet Wilcox. "Machine Learning Hunts Down Hidden Gene Switches Across the Fungal Tree of Life." Scienmag. September 26, 2026. https://scienmag.com/machine-learning-hunts-down-hidden-gene-switches-across-the-fungal-tree-of-life/








