The biomedical literature is, in principle, one of the greatest data resources ever assembled. Decades of genomics, transcriptomics, and proteomics studies have deposited raw sequencing reads, mass spectrometry files, and abundance matrices in public repositories, and the sheer volume of this material grows every year. In practice, however, most of those published datasets are never reused. The information needed to reproduce a reported result—the exact parameters of a quantification pipeline, the location of supplementary files, the version of a software tool, the mapping between sample identifiers and experimental conditions—is scattered across main texts, supplements, and code repositories in ways that make systematic computational reuse extraordinarily laborious. A new study published in PLOS Computational Biology by Alexandre Hutton and Jesse G. Meyer argues that this bottleneck is now tractable, and demonstrates it with a framework in which large language model agents do the tedious work of finding, downloading, rerunning, and synthesizing published omics data.
The core idea is deceptively simple: rather than building a single monolithic program to mine the literature, the researchers assembled a set of LLM-based agents, each equipped with tools for specific tasks. One set of tools fetches omics studies and extracts article metadata. Another identifies and downloads published data files, whether raw instrument output or processed abundance tables. A third executes containerized quantification pipelines, so that the agents can re-derive abundance estimates from raw data rather than trusting deposited summaries. Finally, synthesis tools allow the agents to compare results across studies, including formal statistical combination of findings. The agents coordinate these capabilities through the Model Context Protocol, or MCP, a standard way of exposing containerized software as callable services so that a language model can invoke complex bioinformatics workflows without the workflow logic being hard-coded into the model itself.
Containerization is a crucial design choice. Reproducibility crises in computational biology often trace back to environment drift: a pipeline that ran on one machine with one version of a proteomics search engine may produce subtly different results elsewhere. By wrapping each analysis tool in a container with pinned dependencies, the framework ensures that whatever the agent decides to run, it runs the same way every time. The agent’s contribution is then decision-making—choosing which study to examine, which files to retrieve, which pipeline is appropriate for a given data type, and how to interpret the output—while the numerical heavy lifting remains the province of established, auditable software.
To test the system at scale, the researchers first ran the pipeline across thousands of articles in PubMed Central, cataloguing references to datasets and recording where published data could actually be located. The authors are careful about what this corpus-scale exercise demonstrates. They present the catalogue as a descriptive system output rather than a validated measurement of extraction accuracy, an honest framing that distinguishes feasibility from benchmarked performance. Even so, the exercise illustrates the kind of literature-wide survey that would be impractical to perform manually: a systematic map of which papers point to which repositories, and how often the essential files are actually findable.
The more demanding test came in the form of five end-to-end reanalyses, spanning data-dependent acquisition proteomics, data-independent acquisition proteomics, and bulk RNA sequencing. In each case, the agents had to locate the relevant publication, identify the deposited raw data, download it, execute an appropriate quantification pipeline, and compare their re-derived abundances against the values the original authors had deposited. All five reanalyses completed successfully, though the authors document that each required some human guidance and workflow accommodations along the way. This is an important caveat: the system is agent-supported rather than fully autonomous, and the paper does not claim otherwise. The human role was documented rather than hidden, which makes the results interpretable as a realistic picture of what current agentic systems can and cannot do unaided.
The quantitative agreement between agent-run pipelines and the original analyses is striking. Per-sample correlations between re-quantified and deposited abundances ranged from 0.85 to 0.997, indicating that the agents’ pipelines recovered essentially the same biological measurements. For differential expression—the comparison of features between experimental conditions that is usually the point of such studies—the fold-change concordance between reanalysis and original was measured by Spearman correlation at 0.88 to 0.91. Among features called differentially expressed in both the original and the reanalysis, there were no direction reversals: nothing that the original authors reported as going up appeared in the reanalysis as going down. Residual differences in the lists of significant features turned out to be attributable to threshold placement, tool-version differences, and preprocessing choices rather than to any discrepancy in the underlying quantities, which is precisely the kind of benign variation that experienced reanalysts expect.
Beyond reproducing single studies, the framework demonstrates capabilities that move toward genuine scientific synthesis. The agents were able to identify semantically similar studies—papers addressing comparable biological questions even when they use different vocabulary—and to judge whether datasets from different sources are compatible enough to be analyzed together. The capstone demonstration is a random-effects meta-analysis across multiple proteomics studies of liver fibrosis, in which the agents combined reanalyzed data from independent publications and recovered a consistent pattern of protein regulation. Meta-analysis is among the most valuable and most labor-intensive activities in biomedicine, and the prospect of automating the data-gathering and harmonization steps, while keeping the statistical model transparent and auditable, is among the most consequential implications of this work.
The authors are explicit that the study is a feasibility demonstration, not a validated benchmark of literature-wide performance. Five successful reanalyses, however carefully chosen, do not establish that the system would succeed on an unbiased sample of the literature, and the corpus-scale catalogue was not independently validated for accuracy. What the work does establish is an auditable, reusable toolset: the containers, the MCP-exposed tools, and the agent workflows are all designed to be inspected, rerun, and extended by others. That design philosophy matters because agentic systems introduce their own reproducibility risks. A language model’s behavior can vary across runs and model versions, and the authors’ decision to document every human intervention and workflow accommodation provides a template for how such systems should be reported if their results are to be trusted.
The broader significance lies in what it suggests about the future of data reuse. If the friction that keeps most published omics data unused can be reduced to a prompt, the effective size of the public data commons multiplies enormously. Individual laboratories could re-examine published datasets in light of new hypotheses without negotiating the plumbing of repositories and pipelines; meta-analyses could draw on far more studies; and negative results buried in supplementary files could be surfaced systematically. At the same time, the study’s honest accounting of human guidance serves as a reminder that the last mile of computational biology—judging whether a dataset truly fits a question, and whether a pipeline’s quirks matter for a given conclusion—still benefits from expert oversight. What Hutton and Meyer have shown is that the plumbing no longer has to be the hard part, and that the agents needed to do it are no longer hypothetical.
Subject of Research: Agent-supported retrieval, reanalysis, and synthesis of published omics data using large language model agents
Article Title: Omics data discovery agents: Agent-supported retrieval, reanalysis, and synthesis of published omics data
Article References: Hutton, A., & Meyer, J. G. (2026). Omics data discovery agents: Agent-supported retrieval, reanalysis, and synthesis of published omics data. PLOS Computational Biology, 22(10), e1014822. https://doi.org/10.1371/journal.pcbi.1014822
Image Credits: AI Generated
DOI: 10.1371/journal.pcbi.1014822
Keywords: omics, large language models, AI agents, reproducibility, proteomics, RNA-seq, meta-analysis, Model Context Protocol, data reuse, bioinformatics, PubMed Central, liver fibrosis
Cite Scienmag News
Drew Townsend. (October 9, 2026). AI Agents Learn to Find, Rerun, and Reanalyze Published Omics Data. Scienmag. https://scienmag.com/ai-agents-learn-to-find-rerun-and-reanalyze-published-omics-data/
Drew Townsend. "AI Agents Learn to Find, Rerun, and Reanalyze Published Omics Data." Scienmag, 9 October 2026, https://scienmag.com/ai-agents-learn-to-find-rerun-and-reanalyze-published-omics-data/. Accessed 9 October 2026.
Drew Townsend. "AI Agents Learn to Find, Rerun, and Reanalyze Published Omics Data." Scienmag. October 9, 2026. https://scienmag.com/ai-agents-learn-to-find-rerun-and-reanalyze-published-omics-data/

