The human body is, in a very real sense, a chemical factory that never stops working. Alongside the molecules encoded by our own metabolism, the trillions of microbes that inhabit the gut churn out thousands of small organic compounds that influence immunity, metabolism, and a host of other physiological processes. Yet for all the sophistication of modern biomedical science, identifying what those molecules actually are remains one of the field’s most stubborn bottlenecks. By most estimates, more than 80 percent of the compounds detected in a typical biological sample cannot be matched to any known structure using current methods. They appear in datasets as anonymous peaks, chemically real but biologically silent, their identities locked away behind the limits of available reference data.
Researchers at the Boyce Thompson Institute and Cornell University have now developed a tool that begins to dismantle that bottleneck. The tool, called AIMe, short for AI Molecule Explorer, uses a form of artificial intelligence known as neuro-symbolic AI to predict, organize, and search the mass spectra of the entirety of known small organic molecules, a catalog of more than 100 million compounds. In doing so, it effectively builds a vast, searchable map of chemical space designed to accelerate both hypothesis generation and compound identification. The work is a collaboration between Frank Schroeder, Professor at the Boyce Thompson Institute and in Cornell’s Department of Chemistry and Chemical Biology, and Carla Gomes, Professor of Computing and Information Science and director of Cornell’s AI for Science Institute.
To understand why AIMe matters, it helps to understand how small molecule identification has worked for decades. Mass spectrometry is the workhorse technique, deployed everywhere from toxicology labs to food analysis facilities. When a compound is analyzed, the instrument fragments it and records the masses of the resulting pieces. That pattern of fragments, known as the tandem mass spectrum or MS2 spectrum, functions as a molecular fingerprint. To identify an unknown compound, researchers traditionally compare its spectrum against a reference library of spectra from known compounds, or they attempt to interpret the fragmentation pattern manually, developing hypotheses about the structure one painstaking step at a time.
The trouble is that experimental reference libraries remain sparse. Collectively, the available libraries cover fewer than 1 percent of known compounds, and resolving the structure of a single unknown can take days to months of iterative analysis and experimental validation. As a result, spectra without close library matches usually remain unannotated, no matter how abundant or biologically interesting they might be. Expert interpretation, meanwhile, does not scale. There are simply not enough specialists to work through the mountains of unannotated data accumulating in repositories around the world.
AIMe takes a fundamentally different approach. Rather than waiting for experimental spectra to accumulate in libraries, it predicts spectra computationally, then organizes those predictions into a searchable resource called MS2KOSMOS. Using this pipeline, the team generated more than 800 million predicted spectra covering essentially all known small organic molecules in PubChem, the largest publicly available chemical database. That represents roughly a thousandfold expansion of searchable chemical space relative to existing experimental libraries, a leap that transforms what is even possible to attempt when confronting an unknown spectrum.
At the heart of the system sits DeepMS2Reasoner, a model designed to simulate how molecules fragment inside a mass spectrometer. According to Gomes, the model builds fragmentation pathways step by step, using symbolic chemical rules to enumerate physically plausible fragmentation steps and a neural network to assign likelihoods to each step. The result is not merely a predicted spectrum but an annotated map of how a molecule came apart, a feature that makes AIMe’s outputs interpretable in chemical terms rather than just computationally useful. That interpretability matters enormously for working chemists, who need to understand why a prediction looks the way it does before they commit laboratory resources to testing it.
To demonstrate what the tool can accomplish in practice, the researchers applied it to a comparative metabolomics dataset from mice. The experiment compared germ-free mice, animals raised without any gut microbiota, against mice carrying a normal complement of gut bacteria. Several thousand chemical features differed between the two groups, and most of the abundant ones could not be identified using standard methods. The team used AIMe to query MS2KOSMOS with spectra from the 111 most abundant unidentified microbiota-dependent compounds. For roughly a third of them, AIMe retrieved close predicted spectral matches and related structural candidates, providing chemically interpretable leads for follow-up work. For the rest, the tool mapped the unknown spectra to molecular neighborhoods, sets of structurally related compounds whose shared fragmentation patterns could inform hypotheses about what the unknowns might be.
Two compounds in particular became a case study in what AI-guided structure elucidation can accomplish. Both produced spectra suggesting they were polyamine derivatives, a well-studied class of molecules, but their fragment patterns did not match anything previously described. Using AIMe’s output as a guide, Schroeder’s team assembled candidate structures combinatorially, predicted spectra for each candidate, and used the comparisons to narrow the field. For the first compound, the best candidate was a linear putrescine derivative, which was then easily verified by synthesizing an authentic standard. The second compound proved more elusive. The predicted spectra of potential candidates consistently failed to explain two prominent peaks, until the team expanded the search to include cyclized variants.
Synthesis ultimately confirmed what the predictions suggested. The second compound turned out to be a structurally unusual macrocyclic polyamine, a ring-shaped variant unlike any previously reported from mouse or human biology. Searches of a large public mass spectrometry database subsequently found the same compound in samples of human origin, detected in 57 of 99 human fecal samples examined. The discovery carries real biological weight. Polyamines occupy an important position in biology, sitting at the intersection of diet, the gut microbiome, and immune function, and these findings suggest that the catalog of microbiota-dependent polyamines is far less complete than scientists had assumed. A ring-shaped molecule that had gone unnoticed in human samples until now is a vivid reminder of how much chemistry remains hidden in plain sight.
The researchers also tested AIMe at repository scale, applying it to more than 7 million spectral clusters from the Global Natural Products Social Molecular Networking database, one of the largest publicly available repositories of mass spectrometry data. Prior annotation efforts had managed to annotate roughly 416,000 of those clusters. AIMe returned putative annotations for approximately 2.69 million, using the same similarity threshold applied in that earlier effort, a more than sixfold increase in coverage achieved computationally. AIMe is accessible through Cornell’s website, and the source code will be made publicly available upon publication of the study, with a preprint already posted on bioRxiv. If the tool performs as advertised across the wider community, the anonymous peaks that have long cluttered metabolomics datasets may finally begin to yield their secrets, and the hidden universe of small molecules that shapes human health may come into focus at a scale never before possible.
Subject of Research: Neuro-symbolic AI for predicting mass spectra and identifying unknown small molecules in metabolomics
Article Title: New AI tool maps the hidden universe of small molecules
Article References: New AI tool maps the hidden universe of small molecules. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: AIMe, mass spectrometry, metabolomics, neuro-symbolic AI, small molecules, gut microbiome, polyamines, MS2KOSMOS, PubChem, compound identification, GNPS, chemical space
Cite Scienmag News
Bethany Barker. (October 10, 2026). AI Molecule Explorer maps 100 million compounds to reveal hidden chemistry of life. Scienmag. https://scienmag.com/ai-molecule-explorer-maps-100-million-compounds-to-reveal-hidden-chemistry-of-life/
Bethany Barker. "AI Molecule Explorer maps 100 million compounds to reveal hidden chemistry of life." Scienmag, 10 October 2026, https://scienmag.com/ai-molecule-explorer-maps-100-million-compounds-to-reveal-hidden-chemistry-of-life/. Accessed 10 October 2026.
Bethany Barker. "AI Molecule Explorer maps 100 million compounds to reveal hidden chemistry of life." Scienmag. October 10, 2026. https://scienmag.com/ai-molecule-explorer-maps-100-million-compounds-to-reveal-hidden-chemistry-of-life/

