Metabolomics, the large-scale study of small molecules in biological systems, has long faced a stubborn interpretability problem. Researchers can measure thousands of metabolites at once, but deciding which of those compounds actually matter for distinguishing a diseased sample from a healthy one is far from straightforward. A new study published in BMC Bioinformatics by Francisco Traquete, Carlos Cordeiro, Marta Sousa Silva and António E. N. Ferreira of the Universidade de Lisboa proposes a fresh answer: instead of treating every metabolite as an isolated variable, the method lets a graph neural network see the biochemical web that connects them, and then ranks each compound by how much it moves the network’s final prediction.
Traditional approaches to metabolite importance ranking rely heavily on statistical significance testing or on measures of how much a compound contributes to a classification model’s decisions. The catch, as the Portuguese team points out, is that these approaches treat metabolites as independent variables. In reality, metabolism is anything but independent. Metabolites are linked through enzymatic reactions, and compounds that are close together in a biochemical network often change together when biology shifts. By ignoring this structure, conventional ranking methods can miss groups of interconnected metabolites that jointly drive the biological differences between sample classes, such as healthy versus diseased or treated versus untreated.
The new method begins with a data representation known as a Formula Difference Network, abbreviated FDiN. In this kind of network, each node represents a metabolite, and edges connect compounds whose chemical formulas differ in a meaningful and systematic way. This construction is closely related to mass-difference networks, which have become popular in the analysis of high-resolution mass spectrometry data, particularly the ultra-high-resolution measurements produced by Fourier-transform ion-cyclotron-resonance mass spectrometers. The idea is elegant in its simplicity: if two metabolites differ by a known chemical transformation, such as the addition of a specific functional group, that relationship is encoded as an edge, giving the machine learning model a scaffold of biochemical plausibility to work from.
Once the Formula Difference Network is built, it serves as the input to a predictive graph neural network, which the authors call a Formula Difference Graph Neural Network, or FDiGNN. Graph neural networks are a class of deep learning models designed to operate on graph-structured data. Instead of processing each data point in isolation, they pass information between neighbouring nodes, layer by layer, so that each metabolite’s internal representation comes to reflect not only its own measured abundance but also the abundances and properties of its biochemical neighbours. The architecture developed by the team includes a Global Attention Pooling layer, a component that weighs the contributions of different nodes when the network aggregates graph-level information to make its final class prediction, and the authors acknowledge Mário Nascimento for his help with that layer.
The genuinely novel step comes after the model has been fitted to the data. The researchers rank metabolites by their impact on the sample class prediction probabilities, in effect asking how much each node’s presence and behaviour influences the probability that the network assigns to a given class. This ranking strategy, which the authors refer to as a Probability-Impact based Network Node Importance method, abbreviated PINNI, turns the trained model itself into an interpretability tool. Rather than asking a separate statistical test which metabolites differ between groups, the approach interrogates the prediction machinery of the graph neural network and traces importance back to individual compounds within the network context.
To find out whether this strategy actually works, the team applied it to three benchmark datasets. The results showed a clear pattern: the ranking favoured subnetworks of metabolites, with connectivity acting as a driving factor in how importance was assigned. When the top-ranked compounds were examined, they turned out to be enriched in edges, meaning that the metabolites the method flagged as most important were more densely connected to one another than would be expected by chance. In practical terms, the method tends to highlight coherent biochemical neighbourhoods rather than scattered individual compounds, which is precisely the kind of output that biologists looking for altered pathways or reaction modules want to see.
The team also carried out a careful stress test of the method’s behaviour. Using datasets in which some metabolites had simulated significance artificially built into them, the researchers examined how the ranking distributed importance across the network. They found a clear bias toward assigning higher importance to nodes located in connected subgraphs. This is a double-edged finding. On one hand, it confirms that the method genuinely exploits network structure, which is its intended purpose. On the other hand, it means that isolated metabolites, however biologically relevant they might be, may be systematically under-ranked simply because they lack connections within the Formula Difference Network. Understanding this bias is essential for anyone hoping to apply the method to real experimental data, because it defines the kinds of biological signals the approach is best positioned to detect.
The authors position their strategy as an alternative to the common workflow in which important metabolites are mapped onto curated biological pathways and assessed through enrichment analysis. Pathway enrichment has served metabolomics well for years, but it depends on the completeness and quality of pathway databases, and it can struggle with compounds that are poorly annotated or absent from known pathways. A network-based ranking that emerges directly from a trained predictive model offers a complementary route: it does not require pre-existing pathway annotations, and it can, in principle, surface coordinated changes in metabolite connectivity that enrichment analysis might overlook. For researchers working with complex, high-resolution metabolomic datasets, particularly those generated in environmental, clinical or microbiome studies where many compounds remain unidentified, this independence from curated databases could prove valuable.
The work also speaks to a broader conversation in machine learning about the interpretability of graph neural networks. These models have become powerful tools across the life sciences, from drug discovery to protein structure prediction, but their decisions are notoriously difficult to explain. By building interpretability into the ranking of nodes according to their impact on prediction probabilities, the Lisbon team offers a template that could inspire similar approaches in other network-based biological applications, including gene regulatory networks, protein interaction networks and reaction networks in systems biology. The method demonstrates that a model’s internal prediction dynamics can be mined for biological insight rather than treated as an opaque black box.
The research was funded by the Fundação para a Ciência e a Tecnologia in Portugal through projects 2023.14744.PEX and 2024.06915.CPCA.A1, along with a PhD grant to Francisco Traquete, and was conducted within the European Network of Fourier-Transform Ion-Cyclotron-Resonance Mass Spectrometry Centers funded by the European Union’s Horizon 2020 programme, as well as the Portuguese Mass Spectrometry Network. Published open access under a Creative Commons Attribution licence, the study arrives at a moment when metabolomics is generating ever larger and more complex datasets, and when the demand for methods that can translate raw molecular measurements into biological understanding has never been higher. By teaching a neural network to respect the wiring diagram of metabolism and then asking it which wires matter most, the Portuguese team has added a genuinely new instrument to the metabolomics toolkit, one that promises rankings of metabolites that reflect not just statistical noise thresholds but the connected architecture of biochemistry itself.
Subject of Research: Graph neural network interpretability for ranking metabolites in metabolomics networks
Article Title: Enhancing interpretability in metabolomics: ranking metabolites by their impact on graph neural network predictions
Article References: Traquete, F., Cordeiro, C., Sousa Silva, M., & Ferreira, A. E. N. (2026). Enhancing interpretability in metabolomics: ranking metabolites by their impact on graph neural network predictions. BMC Bioinformatics. https://doi.org/10.1186/s12859-026-06633-7
Image Credits: AI Generated
DOI: 10.1186/s12859-026-06633-7
Keywords: metabolomics, graph neural networks, interpretability, Formula Difference Networks, biochemical networks, mass spectrometry, machine learning, metabolite ranking, BMC Bioinformatics, node importance, pathway analysis, computational biology
Cite Scienmag News
Cassandra Pierce. (October 3, 2026). New AI Method Ranks Metabolites by Their Impact on Graph Neural Network Predictions. Scienmag. https://scienmag.com/new-ai-method-ranks-metabolites-by-their-impact-on-graph-neural-network-predictions/
Cassandra Pierce. "New AI Method Ranks Metabolites by Their Impact on Graph Neural Network Predictions." Scienmag, 3 October 2026, https://scienmag.com/new-ai-method-ranks-metabolites-by-their-impact-on-graph-neural-network-predictions/. Accessed 3 October 2026.
Cassandra Pierce. "New AI Method Ranks Metabolites by Their Impact on Graph Neural Network Predictions." Scienmag. October 3, 2026. https://scienmag.com/new-ai-method-ranks-metabolites-by-their-impact-on-graph-neural-network-predictions/

