<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>reproducible research &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/reproducible-research/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 00:13:31 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>reproducible research &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Rainfall Before Flowering Predicts Honey Yields, Machine Learning Study Finds</title>
		<link>https://scienmag.com/rainfall-before-flowering-predicts-honey-yields-machine-learning-study-finds/</link>
		
		<dc:creator><![CDATA[Teresa Odom]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 00:13:31 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[beekeeping]]></category>
		<category><![CDATA[climate factors affecting honey harvests]]></category>
		<category><![CDATA[climate variables]]></category>
		<category><![CDATA[cross-validation]]></category>
		<category><![CDATA[feature importance]]></category>
		<category><![CDATA[global research on honey yields]]></category>
		<category><![CDATA[honey production forecasting methods]]></category>
		<category><![CDATA[honey yield prediction]]></category>
		<category><![CDATA[innovative beekeeping technology]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in agriculture]]></category>
		<category><![CDATA[machine learning workflows for beekeeping]]></category>
		<category><![CDATA[open-source agricultural data]]></category>
		<category><![CDATA[precision apiculture]]></category>
		<category><![CDATA[predictive modeling in apiculture]]></category>
		<category><![CDATA[rainfall]]></category>
		<category><![CDATA[rainfall impact on honey production]]></category>
		<category><![CDATA[Random Forest]]></category>
		<category><![CDATA[reproducible research]]></category>
		<category><![CDATA[seasonal honey yield classification]]></category>
		<category><![CDATA[SHAP explainability]]></category>
		<category><![CDATA[SMOTE]]></category>
		<category><![CDATA[weather data for beekeeping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224490</guid>

					<description><![CDATA[An international research team has built an open, reproducible machine learning workflow that classifies honey seasons as poor, moderate, or good using only freely available temperature and rainfall data, with rainfall before flowering emerging as the strongest predictor.]]></description>
										<content:encoded><![CDATA[<p>Honey yields are notoriously difficult to forecast. Beekeepers must decide months in advance whether to invest in supplementary feeding, extra hive boxes, labor, and the costly transport of colonies, all without knowing whether the coming season will reward them with a bumper harvest or a disappointing one. Now, a team of researchers from Chile, Peru, Brazil, and Australia has built a machine learning workflow that classifies each beekeeping season as poor, moderate, or good using nothing more than freely available weather data, and they have released everything, including the raw data and code, so that other scientists can reproduce and extend the work.</p>
<p>The study, published in the journal Smart Agricultural Technology, addresses a persistent gap in the literature on honey yield prediction. Previous efforts have ranged from simple regression equations linking climate to harvests in the United Kingdom two decades ago, to fuzzy inference systems, radial basis function interpolation, and modern algorithms such as random forests and gradient boosting applied in Australia, Spain, Italy, and Turkey. Yet the authors found that relatively few such studies are indexed in major databases like Scopus and Web of Science, and many of them keep their datasets closed, making it impossible for other researchers to verify results, compare models fairly, or build on the findings. The new work is explicitly designed to break that pattern.</p>
<p>At the heart of the study is a dataset of 49 records of average honey yield per hive, drawn from 15 apiaries in southwest Australia between 2011 and 2018. The main nectar source in that region is the marri tree, Corymbia calophylla, which flowers around February. Because reliable yield records are scarce and often confidential, the researchers deliberately kept the predictor set simple: minimum and maximum temperatures and rainfall, downloaded from weather stations of the Australian Government&#8217;s Bureau of Meteorology located near each apiary. From these they engineered 45 climatic features, including monthly mean temperatures, monthly rainfall, counts of days above 40 degrees Celsius and below 25 degrees Celsius, each indexed by how many months before flowering the measurement was taken.</p>
<p>The team formulated the problem as a supervised multiclass classification task, mirroring the methodology of the seminal 2020 study on marri honey harvests so that results could be compared directly. Yields below 20 kilograms per hive were labeled a poor year, between 20 and 40 kilograms a moderate year, and above 40 kilograms a good year. Exploratory analysis showed that the climatic variables overlap heavily across the three classes, which rules out simple threshold rules and statistical models, and instead points toward machine learning methods capable of capturing complex, nonlinear relationships.</p>
<p>A central innovation of the paper is its layered approach to interpretability. The researchers applied three complementary techniques to understand which variables drive the predictions. First, they computed feature importance in a random forest model based on the mean decrease in impurity, which revealed that rainfall variables dominate: rainfall five, one, and eight months before flowering emerged as the three most influential features, followed by minimum temperatures three and eleven months before flowering. Second, they quantified how frequently each variable appears at different depths of the decision trees inside the random forest, on the premise that the most relevant features sit closer to the root. This analysis largely corroborated the importance rankings, with rainfall eight and five months before flowering again standing out, while the count of extremely hot days above 40 degrees Celsius had almost no influence.</p>
<p>Third, the team used SHAP, or Shapley Additive Explanations, a method rooted in cooperative game theory that distributes the prediction among the input features according to their average marginal contribution. The SHAP analysis added class-specific nuance. For poor harvests, the minimum temperature eleven months before flowering was among the strongest explanatory variables, with low values pushing predictions away from that class. For moderate harvests, rainfall seven months before flowering carried the greatest weight, with high values contributing positively to that prediction. For good harvests, high values of rainfall eight and five months before flowering were the clearest signals. Taken together, the interpretability analyses consistently point to rainfall in the months leading up to flowering as the most reliable indicator of a strong honey season.</p>
<p>On the modeling side, the researchers benchmarked nine widely used algorithms: logistic regression, k-nearest neighbors, support vector machines, decision trees, multilayer perceptrons, random forests, linear discriminant analysis, gradient boosting, and Naive Bayes. They evaluated each under a battery of experimental conditions, including with and without feature selection, with four normalization schemes (no normalization, division by the maximum, min-max scaling, and z-score standardization), and under two cross-validation strategies, leave-one-out and stratified 10-fold. The effects were algorithm-dependent. Distance-based and margin-based methods such as k-nearest neighbors, support vector machines, logistic regression, and multilayer perceptrons benefited from min-max and z-score scaling, while tree-based models like random forests and decision trees were indifferent to normalization, and Naive Bayes sometimes performed worse when the data were rescaled.</p>
<p>Feature selection based on the interpretability results, using the nine most important variables, improved several models, most dramatically linear discriminant analysis, which jumped from an accuracy of 0.65 to 0.80 under both validation schemes. A k-nearest neighbors model with k set to three also reached roughly 0.80 accuracy. To squeeze out further gains on the small dataset, the team applied SMOTE, a synthetic minority over-sampling technique that balances the classes during training, along with bagging and a Voting Classifier that combines the probability outputs of multiple base models. The best configuration paired k-nearest neighbors with a support vector machine under soft voting, achieving an accuracy of 0.82 and a Matthews correlation coefficient above 0.71. That figure represents a substantial improvement over the 0.67 accuracy reported in the foundational Australian study, achieved here without any satellite or remote sensing data at all.</p>
<p>The authors are careful about what these numbers mean. With only 49 samples and no independent external test set, the reported performance reflects cross-validation estimates rather than proven generalization to new regions or production systems. Hyperparameter searches proved unstable on such a small dataset, so the team deliberately used default settings to keep the comparison controlled and reproducible. They frame the contribution not as a single performance record but as a replicable workflow: the raw climatic data, the yield database, the preprocessing scripts, the feature selection procedure, and the model evaluations are all openly available on GitHub, so that any research group can rerun the pipeline, adapt it to local flowering calendars, and retrain the models on regional data.</p>
<p>The practical implications reach beyond the laboratory. Because the model classifies seasons as low, moderate, or high yield potential using weather records that are essentially free, it could power a simple web or mobile tool in which a beekeeper selects an apiary location, the system links it to the nearest weather station, and the season&#8217;s outlook appears as an early alert. Such a signal could guide decisions on feeding schedules, which follow Farrar&#8217;s rule that a hive&#8217;s honey production capacity grows exponentially with bee population, as well as the preparation of honey supers, transhumance planning, labor allocation, and harvest logistics. Cooperatives, technical advisors, and public institutions could use territory-scale forecasts to plan assistance programs and climate adaptation strategies. The authors stress that the tool is meant to support, not replace, the beekeeper&#8217;s expert judgment, and that validation on new seasons and regions, along with the eventual addition of satellite vegetation data and hive sensor information, remains the necessary next step before the approach can be deployed at scale.</p>
<p><strong>Subject of Research:</strong> Machine learning classification of honey yield per hive using climatic variables</p>
<p><strong>Article Title:</strong> Machine learning models combined with feature importance methods for honey yield classification: A replicable approach</p>
<p><strong>Article References:</strong> Ahumada-García, R., Zabala-Blanco, D., Monzón, V. H., Sánchez, I., da Silva, N. F. F., Rosa, T. C., Ferreira, A. I. S., López-Cortés, X., Flores-Calero, M., &amp; Vasquez-Iglesias, P. (2026). Machine learning models combined with feature importance methods for honey yield classification: A replicable approach. <em>Smart Agricultural Technology, 15</em>, Article 102575. <a href="https://doi.org/10.1016/j.atech.2026.102575" rel="noopener noreferrer">https://doi.org/10.1016/j.atech.2026.102575</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.atech.2026.102575" rel="noopener noreferrer">10.1016/j.atech.2026.102575</a></p>
<p><strong>Keywords:</strong> machine learning, honey yield prediction, beekeeping, random forest, SHAP explainability, feature importance, climate variables, rainfall, SMOTE, cross-validation, precision apiculture, reproducible research</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224490</post-id>	</item>
		<item>
		<title>New Open Pipeline Speeds the Hunt for Unknown Viruses in Sequencing Data</title>
		<link>https://scienmag.com/new-open-pipeline-speeds-the-hunt-for-unknown-viruses-in-sequencing-data/</link>
		
		<dc:creator><![CDATA[Kristina Jarvis]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:02:00 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[assembly and taxonomic classification of viral sequences]]></category>
		<category><![CDATA[bioinformatics pipeline]]></category>
		<category><![CDATA[customizable bioinformatics pipelines]]></category>
		<category><![CDATA[detection of novel viruses]]></category>
		<category><![CDATA[emerging viruses]]></category>
		<category><![CDATA[environmental and clinical viral surveillance]]></category>
		<category><![CDATA[high-throughput viral discovery tools]]></category>
		<category><![CDATA[LazypipeX]]></category>
		<category><![CDATA[metagenomic data processing for virus discovery]]></category>
		<category><![CDATA[metagenomics]]></category>
		<category><![CDATA[next-generation sequencing]]></category>
		<category><![CDATA[next-generation sequencing data analysis]]></category>
		<category><![CDATA[NGS data analysis]]></category>
		<category><![CDATA[open-source viral detection pipelines]]></category>
		<category><![CDATA[pathogen detection]]></category>
		<category><![CDATA[reduced computational barriers in virology]]></category>
		<category><![CDATA[reproducible research]]></category>
		<category><![CDATA[speed and sensitivity in viral metagenomics]]></category>
		<category><![CDATA[viral metagenomics]]></category>
		<category><![CDATA[viral signal extraction from complex samples]]></category>
		<category><![CDATA[viral surveillance]]></category>
		<category><![CDATA[virome analysis]]></category>
		<category><![CDATA[virus discovery]]></category>
		<category><![CDATA[zoonotic spillover]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202368</guid>

					<description><![CDATA[Scientists have developed LazypipeX, a customizable and sensitive bioinformatics pipeline that accelerates the discovery of known and novel viruses from next-generation sequencing data.]]></description>
										<content:encoded><![CDATA[<p>Researchers have introduced LazypipeX, a customizable bioinformatics pipeline designed to make the discovery of novel viruses from next-generation sequencing (NGS) data faster, more sensitive, and more accessible to laboratories that lack large dedicated computational teams. Reported in npj Viruses, the work addresses one of the persistent bottlenecks in modern virology: the sheer difficulty of extracting meaningful viral signals from the enormous volumes of genetic sequence data that modern sequencing platforms generate. As sequencing costs continue to fall and metagenomic studies multiply, the ability to sift rapidly through millions of reads for traces of known and unknown viruses has become a defining capability for surveillance, diagnostics, and basic research alike.</p>
<p>The core problem LazypipeX tackles is well known to anyone who has worked in viral metagenomics. Sequencing a clinical sample, an environmental swab, or a pooled insect collection produces a mixture of host genetic material, bacterial genomes, and—usually in small proportions—viral sequences. Identifying those viral fragments requires a chain of computational steps: quality control of raw reads, removal of host and bacterial contamination, assembly of short reads into longer contiguous sequences, taxonomic classification, and comparison against reference databases to flag sequences that might represent novel agents. Each step has traditionally demanded separate tools, manual file handling, and considerable expertise in command-line computing, which has slowed analysis and introduced opportunities for error.</p>
<p>LazypipeX builds on the design philosophy of its predecessor, Lazypipe, which was developed to automate virome analysis in a single streamlined workflow. The new version extends that concept with a modular, customizable architecture intended to serve a much wider range of use cases. Users can tailor the pipeline to their specific data types, computational resources, and research questions, swapping components in and out without breaking the overall workflow. This flexibility matters because virome studies vary enormously: a hospital laboratory screening patient samples for known respiratory viruses has different needs from an ecology group cataloguing the viromes of wild rodents or an agricultural institute monitoring crops for emerging plant pathogens.</p>
<p>A central emphasis of the new pipeline is speed. The authors describe optimizations that allow rapid processing of large sequencing datasets, enabling iterative analysis in which researchers can screen samples, refine parameters, and re-analyze within a working session rather than waiting days for batch jobs to complete. In outbreak situations, where public health decisions depend on quickly knowing whether an unusual pathogen is present, that turnaround time can be decisive. Speed also changes the texture of exploratory research: when analysis cycles take hours rather than days, scientists can afford to ask more questions of their data, testing alternative assembly strategies or database configurations that a slower workflow would make impractical.</p>
<p>Sensitivity is the pipeline&#8217;s second headline virtue. Virus discovery often hinges on detecting sequences present at very low abundance in a background of overwhelming host DNA or RNA. Missing those faint signals can mean missing an emerging pathogen entirely. LazypipeX incorporates multiple complementary detection strategies, combining alignment-based approaches that find sequences resembling known viruses with assembly-based and similarity-based methods that can reveal more distant relatives or entirely novel agents. By running several strategies in parallel and consolidating their outputs, the pipeline increases the chance that something genuinely interesting will surface rather than be discarded as noise.</p>
<p>The pipeline&#8217;s classification stage draws on comprehensive protein and nucleotide sequence databases to assign likely identities to detected viral sequences, while explicitly flagging candidates that lack close matches—precisely the sequences most likely to represent new species or genera. This tiered reporting is a deliberate design choice. Rather than presenting a single flattened list of detections, LazypipeX helps researchers distinguish between routine findings, such as abundant bacteriophages or common plant viruses, and rare, divergent sequences that merit deeper investigation, such as de novo assembly, targeted PCR confirmation, or additional sampling.</p>
<p>Customizability extends beyond the choice of individual tools. The pipeline is structured so that laboratories can integrate their own reference databases, which is particularly valuable in regions or fields where locally relevant pathogens are underrepresented in public repositories. A laboratory in a dengue-endemic country, for example, can weight its analyses toward flavivirus references and local strain data, improving both sensitivity and interpretation. Similarly, groups studying wildlife viromes can add their own curated sets of viral genomes to reduce misclassification. This openness contrasts with rigid black-box solutions and reflects a broader movement in bioinformatics toward transparent, reproducible, and adaptable analytical frameworks.</p>
<p>Reproducibility receives careful attention as well. The workflow is implemented with containerization and dependency management practices that allow the exact computational environment to be shared alongside results, so that a colleague rerunning the analysis obtains the same outputs. In a field where publication reviews increasingly demand evidence that findings are not artifacts of particular software versions or parameter settings, this is more than a convenience. It also lowers the barrier for smaller institutions and research groups in resource-limited settings, since the pipeline is designed to run on modest hardware as well as on high-performance computing clusters, scaling with the data at hand.</p>
<p>The practical implications reach across several domains of viral science. In public health, faster and more sensitive virome screening strengthens surveillance for zoonotic spillover—the event in which a virus jumps from an animal reservoir into humans—a process that has driven pandemics from HIV to influenza to SARS-related coronaviruses. In clinical settings, unbiased metagenomic sequencing supported by pipelines like LazypipeX can identify unexpected pathogens in severely ill patients, guiding treatment when conventional tests fail. In ecology and evolution, comprehensive virome catalogs illuminate how viruses diversify, move between host species, and respond to environmental change. Agriculture and food security benefit too, since early detection of plant and livestock viruses can prevent costly outbreaks.</p>
<p>The release of LazypipeX arrives amid a striking expansion of virus discovery as a discipline. Large-scale projects sampling wildlife, livestock, and human populations have revealed that the virosphere is vastly richer than previously imagined, with potentially hundreds of thousands of vertebrate-infecting viruses awaiting description. Making sense of that torrent of data is fundamentally a computational challenge, and tools that lower the expertise threshold while maintaining scientific rigor will shape how quickly and how reliably the field progresses. By combining speed, sensitivity, and adaptability in a single open framework, LazypipeX positions itself as a practical workhorse for that effort—a pipeline intended not for a narrow niche but for the everyday work of turning raw sequencing reads into biological insight about the viral world.</p>
<p><strong>Subject of Research:</strong> A customizable bioinformatics pipeline for sensitive and rapid virus discovery from NGS data</p>
<p><strong>Article Title:</strong> LazypipeX: customizable virome analysis pipeline enabling fast and sensitive virus discovery from NGS data</p>
<p><strong>Article References:</strong> Weinstein, I., Vapalahti, O., Kant, R., &amp; Smura, T. (2026). LazypipeX: customizable virome analysis pipeline enabling fast and sensitive virus discovery from NGS data. <em>npj Viruses</em>. <a href="https://doi.org/10.1038/s44298-026-00237-x" rel="noopener noreferrer">https://doi.org/10.1038/s44298-026-00237-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44298-026-00237-x" rel="noopener noreferrer">10.1038/s44298-026-00237-x</a></p>
<p><strong>Keywords:</strong> LazypipeX, virus discovery, virome analysis, next-generation sequencing, metagenomics, bioinformatics pipeline, viral surveillance, pathogen detection, zoonotic spillover, NGS data analysis, emerging viruses, reproducible research</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202368</post-id>	</item>
	</channel>
</rss>
