<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>literate programming &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/literate-programming/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 02:23:22 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>literate programming &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Metabonaut: Free Workflows Teach Scientists Reproducible Metabolomics From Raw Data to Results</title>
		<link>https://scienmag.com/metabonaut-free-workflows-teach-scientists-reproducible-metabolomics-from-raw-data-to-results/</link>
		
		<dc:creator><![CDATA[Alexandra Wallace]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 02:23:22 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[Bioconductor]]></category>
		<category><![CDATA[data sharing]]></category>
		<category><![CDATA[FAIR principles]]></category>
		<category><![CDATA[LC-MS/MS]]></category>
		<category><![CDATA[literate programming]]></category>
		<category><![CDATA[mass spectrometry]]></category>
		<category><![CDATA[MassSpectrometry datasets sharing]]></category>
		<category><![CDATA[Metabolomics]]></category>
		<category><![CDATA[metabolomics data analysis training]]></category>
		<category><![CDATA[metabolomics data reanalysis]]></category>
		<category><![CDATA[metabolomics data repositories]]></category>
		<category><![CDATA[metabolomics data sharing]]></category>
		<category><![CDATA[metabolomics research transparency]]></category>
		<category><![CDATA[metabolomics standards and reporting]]></category>
		<category><![CDATA[Metabonaut]]></category>
		<category><![CDATA[mzTab-M]]></category>
		<category><![CDATA[open educational resource]]></category>
		<category><![CDATA[open-access metabolomics education]]></category>
		<category><![CDATA[open-source metabolomics tools]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[reproducibility challenges in metabolomics]]></category>
		<category><![CDATA[reproducible metabolomics workflows]]></category>
		<category><![CDATA[untargeted liquid chromatography mass spectrometry]]></category>
		<category><![CDATA[xcms]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233102</guid>

					<description><![CDATA[An open-access review introduces Metabonaut, a set of executable, containerised teaching workflows that show researchers how to make untargeted LC-MS/MS metabolomics analyses as findable, accessible, interoperable and reusable as the raw data they consume.]]></description>
										<content:encoded><![CDATA[<p>Two decades after the Metabolomics Standards Initiative laid down its minimum-reporting recommendations, the field of untargeted metabolomics has quietly built one of the most impressive data-sharing infrastructures in the life sciences. MetaboLights now hosts 3,080 public studies, Metabolomics Workbench holds 4,379, and GNPS/MassIVE spans 19,734 mass spectrometry datasets, all as of July 2026. Yet a striking paradox sits at the heart of this open-data revolution: even when every raw file is freely downloadable, almost nobody can recompute the results published in the accompanying paper. A new open-access review in the journal Metabolomics confronts that paradox head-on and introduces Metabonaut, an open educational resource designed to teach researchers how to make their analyses as shareable, auditable and reusable as their data.</p>
<p>The problem is not a shortage of tools. Open-source software now covers essentially every step of an untargeted liquid chromatography tandem mass spectrometry workflow, from raw vendor files to annotated feature tables. The problem, the authors argue, is training and practice. A reader who downloads a deposited set of mzML files can rarely reproduce the feature table reported in the corresponding study, because parameter values are only partially documented, manual exclusion and filtering steps are described in prose if at all, and the metadata needed for the downstream statistics is often missing. An evaluation of 399 public studies across four metabolomics repositories found that none of the Metabolomics Standards Initiative biological-context reporting standards was met by every study, with compliance for individual standards ranging from zero to 97 percent. Poor metadata makes a study hard to reuse even when its raw files are open.</p>
<p>The consequences of unreproducible analysis are not hypothetical. In one of the most notorious cases in computational biology, undocumented, spreadsheet-based errors in a set of genomic predictors went undetected until a forensic re-analysis, by which point the flawed models had already been used to assign patients in clinical trials. Scripted, transparent analyses make such errors visible and correctable before they propagate. The review&#8217;s authors, led by Philippine Louail and Johannes Rainer of Eurac Research in Bolzano, Italy, adopt the FAIR principles, originally formulated for data and later extended to research software and workflows, and apply them to the analysis itself. In their framing, an analysis is FAIR when its code, computational environment, narrative description and intermediate outputs are findable under a persistent identifier, accessible under clearly defined conditions, interoperable through open standards such as mzML, mzTab-M and HDF5, and reusable enough for others to adapt.</p>
<p>Metabonaut is built on the R for Mass Spectrometry initiative, part of the Bioconductor project, which enforces strict package review, shared base classes and regular release cycles. At the lowest level, the Spectra package abstracts mass spectrometry data independently of where they are stored, exposing the same R interface whether the data live in memory, on disk, in a database or in a public repository. Chromatograms plays the analogous role for chromatographic data, while MsExperiment ties spectra, chromatograms and sample metadata into a single experiment-level container. On this foundation sits xcms 4, one of the field&#8217;s main preprocessing tools, refactored to integrate natively with the rest of the ecosystem, alongside MetaboAnnotation and CompoundDb, which expose a uniform annotation interface across in-house and public reference data from MassBank, HMDB or GNPS. The pedagogical payoff of this shared architecture is substantial: a learner who masters one Spectra object and one MsExperiment object carries that mastery across preprocessing, annotation and statistics, and the same code runs unchanged on three files or on several thousand.</p>
<p>The medium that binds everything together is literate programming, an idea Donald Knuth introduced in the early 1990s and scientific computing has since adapted through tools such as R Markdown and, now, Quarto. In a Quarto document, explanatory prose, executable code and rendered output, including tables, figures and console results, are interleaved in a single source file that compiles to one document. The authors call this practice, applied to a complete data analysis, literate analysis, and they argue it changes the unit of scientific communication. In the conventional model, a metabolomics study is reported across separate artefacts, a methods section, figures and tables, supplementary parameter lists and, if the authors are diligent, a code repository, and these can drift apart with no reliable way to tell which is authoritative. A Quarto vignette collapses all four into one: because the document executes, the description of the analysis and the analysis itself are the same file. Writing an analysis as plain text also makes it amenable to version control, so any two versions can be compared line by line, and a tagged release deposited in Zenodo receives a DOI and can be retrieved unchanged.</p>
<p>Durability is handled by containerisation. A container image is a self-contained snapshot of the whole computational stack, the operating system libraries, the R version, the pinned Bioconductor release and every package the analysis depends on, frozen at publication time. For a user this replaces dependency management with two commands, installing a runtime such as Docker once and then pulling the archived image, and it is what allows a vignette to be re-executed years later in the same environment rather than merely read. There is even a contemporary twist: because a script is plain text, a large language model can read it, explain it and suggest fixes, so a learner who does not understand a peak-detection call can paste that code into an assistant and get an explanation in whichever language they prefer. A sequence of menu clicks in a graphical tool leaves no text to share and offers no equivalent.</p>
<p>Metabonaut version 1.6.2 comprises nine executable vignettes, all released under a CC BY-SA 4.0 licence and archived on Zenodo. The keystone is a complete end-to-end analysis that carries a small blood plasma dataset from cardiovascular disease patients and healthy controls, deposited in MetaboLights as MTBLS8735, from preprocessing through statistical analysis to annotation using xcms, MetaboAnnotation, SummarizedExperiment and limma. Three further vignettes cover preprocessing, including dataset investigation before peak detection and quality-control-driven normalisation and feature selection with the notame package, which handles drift correction, batch effects, QC-based filtering and random-forest imputation of missing values. Three cover annotation, from building reference libraries out of GNPS and MassBank to cross-language annotation with Python and advanced feature annotation through the SIRIUS framework. A further vignette merges new data with an already-processed reference dataset, and a ninth demonstrates import and export of the standardised mzTab-M reporting format. Most vignettes execute in under ten minutes; two require special conditions, including one that processes roughly 600 gigabytes of data from a 4,063-sample population study and takes more than seventeen hours.</p>
<p>One of the resource&#8217;s most persuasive demonstrations is that scripted analysis does not sacrifice the visual intuition graphical tools provide; it amplifies it. Because plots are generated from code, they can be reproduced, parameterised and applied uniformly across an entire dataset. A single command produces a grid of chromatograms across hundreds of samples or the extracted-ion chromatogram of one feature across an entire cohort, faceted by injection order or study group. Manually opening each sample in a graphical tool does not scale, whereas a scripted plot is written once and reused unchanged whatever the number of samples. The same logic extends to public archives: the MsBackendMetaboLights package turns an archive identifier directly into a Spectra object without manual download, so a deposited dataset becomes an input to a new analysis rather than merely the endpoint of the study that produced it. The Seamless Alignment vignette aligns features in mass-to-charge and retention-time space between a new cohort and a previously preprocessed public reference, enabling cross-study comparison without joint reprocessing. The Large Scale Processing vignette shows that the xcms idioms used on a dozen files are identical on thousands; only the storage backend changes.</p>
<p>Interoperability extends beyond R. The SpectriPy package bridges Bioconductor&#8217;s Spectra to Python&#8217;s matchms, letting an R workflow call the ModifiedCosine algorithm to filter and score MS2 spectra and bring the results back for interpretation, all within the same Quarto document. RuSirius gives programmatic access to SIRIUS, driving formula assignment, CSI:FingerID and COSMIC structure ranking and MSNovelist de novo structure generation, with a clear annotation logic taught in the vignette: library search first, COSMIC for orthogonal confirmation, and MSNovelist for genuinely novel structures. Results themselves are made FAIR through mzTab-M, the HUPO-PSI standard reporting format for metabolomics, which ensures that a quantitative result and the evidence behind it travel together and remain readable by other software without bespoke parsing. The authors are candid about limits: the resource covers LC-MS and LC-MS/MS only, leaving NMR and ion-mobility mass spectrometry out of scope, and R remains a genuine learning investment that supervisors should budget for, though they report that a few focused weeks are enough to make a newcomer productive. Their larger ambition is cultural. Much of the progress on the data side of FAIR followed from obligation rather than goodwill, as journals and funders began requiring raw data deposition. The authors expect the analysis side to follow the same path, and they consider it a reasonable goal that depositing a re-executable analysis becomes an expectation of publishing rather than an act of individual diligence. Metabonaut, freely available on GitHub, Docker Hub and the Bioconductor Workshop Platform, is their blueprint for what that expectation should look like.</p>
<p><strong>Subject of Research:</strong> Reproducible, script-based untargeted LC-MS/MS metabolomics data analysis and FAIR-principle education through the open Metabonaut resource</p>
<p><strong>Article Title:</strong> Metabonaut: an open educational resource for learning reproducible, script-based untargeted LC-MS/MS metabolomics</p>
<p><strong>Article References:</strong> Louail, P., De Graeve, M., Mahmoud, A., Marques de Sá e Silva, D., Nishida, K., Suksi, V., Tagliaferri, A., Tomè, G., Verri Hernandes, V., &amp; Rainer, J. (2026). Metabonaut: an open educational resource for learning reproducible, script-based untargeted LC-MS/MS metabolomics. <em>Metabolomics, 22</em>(5), Article 165. <a href="https://doi.org/10.1007/s11306-026-02538-x" rel="noopener noreferrer">https://doi.org/10.1007/s11306-026-02538-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11306-026-02538-x" rel="noopener noreferrer">10.1007/s11306-026-02538-x</a></p>
<p><strong>Keywords:</strong> metabolomics, Metabonaut, FAIR principles, reproducibility, LC-MS/MS, mass spectrometry, Bioconductor, literate programming, open educational resource, xcms, mzTab-M, data sharing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233102</post-id>	</item>
	</channel>
</rss>
