Every gene in the human genome carries a hidden layer of instructions that determines which parts of its RNA message are kept and which are discarded. This process, known as alternative splicing, allows a single gene to produce many different proteins and is one of the main reasons an organism as complex as a human can function with a relatively modest number of genes. The decision to include or skip a particular segment of RNA is made by RNA-binding proteins, or RBPs, which recognize short sequence patterns, called motifs, embedded in the RNA molecule itself. A long-standing challenge in molecular biology has been to work backwards: given a set of RNA sequences and a change in splicing behavior, which protein is responsible? A team of Italian researchers now reports a substantially upgraded computational framework, RNAmotifs 2.0, that tackles precisely this problem, and their results suggest the method can reliably identify the correct regulator from sequence and splicing data alone.
The work, published in BMC Biology by Mariachiara Grieco, Tommaso Becchi, Gabriele Boscagli and colleagues under the supervision of Matteo Cereda, spans the University of Milan, the Italian Institute for Genomic Medicine, IFOM ETS, the University of Turin and the IRCCS E. Medea institute. The central difficulty the team set out to address is that RNA-binding proteins rarely act through a single, cleanly defined binding site. Instead, they often recognize multivalent RNA motifs, combinations of short sequence elements arranged in specific positions relative to the splice sites that mark exon boundaries. Connecting such a motif to its candidate trans-acting protein from sequence information alone has remained stubbornly difficult, because many proteins recognize similar sequence patterns and because the functional consequence of binding depends heavily on where in the transcript the motif sits.
RNAmotifs 2.0 approaches the problem by coupling motif discovery with two independent streams of experimental evidence: in vivo binding data and measured splicing responses. The heart of the new framework is a module called MaRs, short for Motifs and Regulators of splicing, which integrates these data types into a single association score linking each multivalent RNA motif to a candidate RBP. The method does not treat all binding evidence equally. It learns position-specific binding and regulatory expectations, meaning it develops a statistical model of where a given protein tends to bind relative to exons and what effect that binding typically has on splicing outcomes. It also weights binding quality, so that stronger, more confident cross-linking signals contribute more to the association than weak or ambiguous ones.
Technically, the framework divides the region around each cassette exon, the type of alternative exon that is either included or skipped in the mature message, into three splice-proximal regulatory regions: the upstream intron, the exon body itself, and the downstream intron. Enrichment of a motif in these regions is captured through a combined regional enrichment measure, and the positional information is critical because a motif that enhances exon inclusion when placed downstream may repress inclusion when placed upstream. The association score that MaRs computes blends several components, including an enrichment score reflecting how strongly a motif is overrepresented near regulated exons, a binding score derived from enhanced cross-linking and immunoprecipitation, or eCLIP, experiments that map where proteins physically touch RNA inside living cells, and a measure of similarity between the predicted splicing response and the observed changes in percent spliced-in values following protein depletion.
A major methodological advance in version 2.0 is the optimization of the framework’s parameters on a per-protein basis. Rather than applying a single fixed set of thresholds and weights to every RBP, the authors use a guided hyperparameter search to tune the analysis for each protein individually. They compared a traditional grid search with Bayesian optimization, a strategy that intelligently samples the parameter space instead of exhaustively testing every combination, and found the Bayesian approach efficient for navigating the large parameter grid. To ensure that these optimized parameters generalize rather than simply overfit the training data, the team employed leave-one-RBP-out cross-validation, in which the framework is tuned without access to the protein being tested and then asked to identify that held-out protein from scratch.
The benchmark for this validation came from the ENCODE project, which has systematically depleted hundreds of RNA-binding proteins in cultured human cells using siRNA and measured the resulting changes in gene expression and splicing through RNA sequencing, alongside eCLIP maps of where each protein binds. Across this large collection of knockdown experiments, RNAmotifs 2.0 consistently prioritized the perturbed regulator as the top candidate for the motifs it discovered, with performance particularly strong when the splicing effects were large, as quantified by the absolute change in mean percent spliced-in upon protein depletion. The framework’s ability to recover the true regulator was evaluated using the area under the receiver operating characteristic curve, a standard measure of ranking quality, and the results were summarized across dozens of proteins in both HepG2 and K562 cell lines.
One practical concern with any method that integrates RNA-seq and eCLIP data is that these datasets are often generated in different cell types, and the authors explicitly tested the robustness of their approach to such mismatches. In a cross-cell-line transfer analysis, they paired knockdown exon sets from one cell line with binding maps from another, for example HepG2 knockdown data analyzed against K562 eCLIP profiles and vice versa. The framework retained substantial predictive power under these mismatched conditions, indicating that the learned positional and regulatory expectations capture biology that is at least partly shared across cellular contexts rather than being an artifact of perfectly matched datasets. An ablation analysis of the association-score components further revealed which ingredients of the score contribute most to performance, providing a transparent account of why the method works rather than presenting it as an opaque black box.
To demonstrate that the framework generalizes beyond the ENCODE benchmark, the team performed their own experimental validation in PC3 prostate cancer cells depleted of HNRNPK, a member of the heterogeneous nuclear ribonucleoprotein family with well-documented roles in RNA processing. After knocking down HNRNPK and profiling splicing changes with rMATS, a standard tool for detecting differential splicing from replicate RNA-seq data, they ran RNAmotifs 2.0 on the resulting set of regulated cassette exons. The framework identified HNRNPK as the splicing regulator, confirming that the method can recover the correct answer in a completely new biological setting with data generated independently of the training resources. This out-of-sample success is arguably the most convincing evidence that the learned motif-to-regulator associations reflect genuine molecular recognition rules.
Interpretability is a defining feature of the new framework and one that distinguishes it from purely predictive machine-learning approaches. Because the association score is built from named, quantifiable components, enrichment, binding quality, positional specificity and response similarity, researchers can inspect exactly why a particular protein was nominated as a candidate regulator of a particular set of exons. The authors also condensed the output into a compact visualization that summarizes the results for each RBP, making it feasible to survey many proteins and cell lines at a glance. Supplementary analyses addressed additional design choices, including the effect of motif width on RBP identification and the robustness of eCLIP-derived binding profiles, with the motif-width benchmark showing how this parameter influences both the identification of the correct protein and the number of enriched multivalent motifs detected.
The significance of this work extends well beyond method development. Alternative splicing is frequently misregulated in cancer and in a wide range of genetic diseases, and mutations in splice-regulatory motifs are increasingly recognized as drivers of pathology. A tool that can connect altered splicing patterns to the proteins responsible, using only sequence data, binding maps and expression measurements, gives researchers a systematic way to dissect these regulatory networks and to generate hypotheses about which factors might be therapeutically targeted. The framework is openly available, and the study’s authors, whose work was supported by AIRC and the Compagnia di San Paolo Foundation, have released extensive supplementary documentation covering every analytical decision. As eCLIP datasets continue to accumulate across tissues and disease states, frameworks like RNAmotifs 2.0 are poised to become standard instruments for translating the growing mountain of RNA data into concrete statements about which molecular players control which splicing decisions.
Subject of Research: A computational framework for linking multivalent RNA motifs to their candidate RNA-binding protein regulators of alternative splicing
Article Title: RNAmotifs 2.0: a framework for inferring multivalent RNA motifs and their candidate regulators of splicing
Article References: Grieco, M., Becchi, T., Boscagli, G., Priante, F., Caizzi, L., Pozzoli, U., & Cereda, M. (2026). RNAmotifs 2.0: a framework for inferring multivalent RNA motifs and their candidate regulators of splicing. BMC Biology. https://doi.org/10.1186/s12915-026-02745-x
Image Credits: AI Generated
DOI: 10.1186/s12915-026-02745-x
Keywords: alternative splicing, RNA-binding proteins, multivalent RNA motifs, RNAmotifs 2.0, eCLIP, RNA sequencing, splicing regulation, ENCODE, Bayesian optimization, HNRNPK, motif discovery, computational biology
Cite Scienmag News
Juliet Wilcox. (October 1, 2026). New Computational Tool Links RNA Motifs to the Proteins That Control Splicing. Scienmag. https://scienmag.com/new-computational-tool-links-rna-motifs-to-the-proteins-that-control-splicing/
Juliet Wilcox. "New Computational Tool Links RNA Motifs to the Proteins That Control Splicing." Scienmag, 1 October 2026, https://scienmag.com/new-computational-tool-links-rna-motifs-to-the-proteins-that-control-splicing/. Accessed 1 October 2026.
Juliet Wilcox. "New Computational Tool Links RNA Motifs to the Proteins That Control Splicing." Scienmag. October 1, 2026. https://scienmag.com/new-computational-tool-links-rna-motifs-to-the-proteins-that-control-splicing/








