<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>genomic prediction in crop breeding &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/genomic-prediction-in-crop-breeding/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 30 Aug 2026 08:21:48 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>genomic prediction in crop breeding &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>GWAS-informed ensembles boost genomic prediction of vitamin A carotenoids in cassava</title>
		<link>https://scienmag.com/gwas-informed-ensembles-boost-genomic-prediction-of-vitamin-a-carotenoids-in-cassava/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sun, 30 Aug 2026 08:21:45 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[accelerated breeding cycles]]></category>
		<category><![CDATA[advanced computational tools in agriculture]]></category>
		<category><![CDATA[biofortification for public health]]></category>
		<category><![CDATA[biofortified cassava varieties]]></category>
		<category><![CDATA[Cassava biofortification]]></category>
		<category><![CDATA[computational approaches to plant biofortification]]></category>
		<category><![CDATA[crop genetic diversity]]></category>
		<category><![CDATA[enhancing crop nutritional traits through genomics]]></category>
		<category><![CDATA[genetic markers for carotenoid content]]></category>
		<category><![CDATA[genome-wide association studies in crop breeding]]></category>
		<category><![CDATA[genomic estimated breeding values]]></category>
		<category><![CDATA[genomic prediction in crop breeding]]></category>
		<category><![CDATA[genomic prediction of vitamin A content]]></category>
		<category><![CDATA[GWAS and machine learning in plant breeding]]></category>
		<category><![CDATA[GWAS-informed machine learning models]]></category>
		<category><![CDATA[improving cassava nutritional quality]]></category>
		<category><![CDATA[improving nutrient content in staple crops]]></category>
		<category><![CDATA[plant genetics and genomics]]></category>
		<category><![CDATA[public health impact of vitamin A-rich crops]]></category>
		<category><![CDATA[rapid cassava breeding techniques]]></category>
		<category><![CDATA[sub-Saharan Africa staple crop improvement]]></category>
		<category><![CDATA[sub-Saharan Africa staple crops]]></category>
		<category><![CDATA[vitamin A carotenoid content in cassava]]></category>
		<guid isPermaLink="false">https://scienmag.com/gwas-informed-ensembles-boost-genomic-prediction-of-vitamin-a-carotenoids-in-cassava/</guid>

					<description><![CDATA[Cassava is a lifeline for millions of people across sub-Saharan Africa, but it is an imperfect lifeline. The starchy roots that anchor diets in Kenya, Uganda, Nigeria and beyond are notoriously poor sources of micronutrients, and vitamin A deficiency remains rampant in communities that depend on the crop as a staple. Biofortification—breeding cassava varieties whose [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Cassava is a lifeline for millions of people across sub-Saharan Africa, but it is an imperfect lifeline. The starchy roots that anchor diets in Kenya, Uganda, Nigeria and beyond are notoriously poor sources of micronutrients, and vitamin A deficiency remains rampant in communities that depend on the crop as a staple. Biofortification—breeding cassava varieties whose roots are rich in pro-vitamin A carotenoids—has long been championed as a public health solution. Now, a team of researchers working with Kenyan and Ugandan germplasm reports a substantial leap forward in the computational machinery of that effort, showing that a smarter combination of genome-wide association studies and machine learning can predict the nutritional quality of cassava roots with unprecedented accuracy.</p>
<p>The core obstacle is time. Cassava takes roughly twelve months to mature, which means breeders who wait to measure carotenoid content in harvested roots must commit years to each selection cycle. Genomic prediction offers a way around this bottleneck. The idea is to build statistical models that link tens of thousands of DNA markers to measured traits, then use those models to calculate genomic estimated breeding values—predictions of an individual&#8217;s additive genetic worth as a parent. Breeders can then select the most promising progenitors at the seedling stage, months before a single root is dug from the ground, compressing the breeding cycle and accelerating genetic gain.</p>
<p>In a study conducted at the Kenya Agricultural and Livestock Research Organization (KALRO) in Kakamega, researchers assembled a panel of 93 pro-vitamin A cassava genotypes derived from accessions held by Uganda&#8217;s National Crop Resource Research Institute. The plants were established as seedlings in January 2023, and in January 2024 two representative roots from each genotype were harvested under low-light evening conditions—a necessary precaution, because carotenoids degrade rapidly when exposed to light. Samples were rushed to the University of Nairobi within seven hours, where high-performance liquid chromatography quantified beta-carotene content in freshly harvested tissue. The measured values ranged from 0.04 to 1.03 mg per 100 g of root, with a mean of 0.52 mg/100 g. Notably, 68 percent of the population fell below the Institute of Medicine&#8217;s recommended benchmark of 6 micrograms of beta-carotene per gram, underscoring how much work remains for breeders.</p>
<p>The genotyping side of the study relied on DArTseq genotyping-by-sequencing against the Manihot esculenta 671_v8.0 reference genome, yielding 54,877 single nucleotide polymorphisms. After quality control filtering for call rate, minor allele frequency and heterozygosity using the snpReady package in R, 37,419 high-quality SNPs remained for downstream analysis. Five-fold cross-validation, in which 80 percent of the data trained each model and the remaining 20 percent served as an independent test, was used to evaluate prediction performance across all models.</p>
<p>The researchers&#8217; central innovation lay in how they identified trait-associated markers to feed into their prediction models. Conventional single-locus GWAS approaches, while widely used, suffer from limited statistical power in small populations—a persistent constraint in crop breeding programs. The team instead turned to multi-locus random marker effect GWAS (mr-GWAS), implemented through four algorithms in the mrMLM platform: mrMLM, FASTmrMLM, pLARmEB and ISIS EM-BLASSO. These methods treat SNP effects as random, shrinking estimates toward zero and applying a modified Bonferroni significance threshold of 0.05 divided by the effective number of markers, which corrects for multiple testing without the punishing conservatism of standard corrections. The payoff was clear: while the BLINK fixed-effect GWAS model detected only two significant SNPs for beta-carotene (on chromosomes 09 and 14) and one for flesh color on chromosome 01, the mr-GWAS pipeline uncovered five significant SNPs distributed across chromosomes 01, 03, 04, 14 and 18.</p>
<p>When these markers were incorporated as fixed-effect covariates into classical Bayesian genomic prediction models—Bayesian Ridge Regression, Bayesian Lasso, BayesA, BayesB and BayesC—the improvement was dramatic. Naive models struggled badly: prediction correlations on the validation set ranged from just 0.14 to 0.16 for beta-carotene. Adding the two BLINK SNPs lifted accuracy to between 0.41 and 0.43. But when all five mr-GWAS-derived SNPs were included, Bayesian Lasso reached a prediction ability of r = 0.81, with the remaining models clustered between 0.75 and 0.78. The pattern held for root flesh color as well, where a single chromosome 01 SNP boosted validation correlations from roughly 0.39–0.48 to 0.61–0.63. The lesson is that markers spanning more of the genome are more likely to be linked to the causal genes governing carotenoid accumulation, capturing a fuller share of the trait&#8217;s genetic architecture.</p>
<p>The study did not stop with parametric models, however. Traditional genomic prediction assumes purely additive gene action, yet real traits are shaped by dominance and epistasis—non-additive interactions that linear models systematically miss. To capture these effects, the researchers built non-parametric models using five machine learning algorithms: Random Forest, Support Vector Machine with radial kernel, Neural Networks, Extreme Gradient Boosting (XGBoost) and K-Nearest Neighbors, along with a stacked ensemble that used Random Forest as the meta-learner. Initially, these models performed modestly, with the ensemble topping out at r = 0.44 and the support vector machine limping along at r = 0.04.</p>
<p>The transformation came through feature selection. Using the Boruta algorithm, the team distilled 37,419 SNPs down to just 13 of the most informative predictors—a 99.97 percent reduction in dimensionality. With this streamlined marker set, performance soared: Random Forest and the ensemble model both achieved r = 0.79, while XGBoost and K-Nearest Neighbors reached r = 0.73. This result carries a practical message for computational breeders everywhere—curating the predictor set matters as much as choosing the algorithm, since removing noise variables sharpens model focus, reduces computational burden and prevents overfitting.</p>
<p>The distinction between the two modeling families maps directly onto breeding strategy. Genomic estimated breeding values from the parametric models reflect additive effects, the only genetic component parents reliably transmit to offspring, making them the right tool for early progenitor selection. Non-parametric machine learning models, by contrast, estimate total genomic value—additive and non-additive combined—which better predicts how a clone will actually perform as a finished variety. Together, the two approaches let breeders pick parents early and pick varieties accurately, shortening every stage of the pipeline from crossing to release.</p>
<p>The authors acknowledge limitations, including the modest population size of 93 genotypes and the fact that GWAS was performed on the full dataset before cross-validation, which may have introduced some optimistic bias into reported accuracies. They recommend larger datasets and nested cross-validation designs in future work. Even so, the results represent a meaningful advance for a crop that has long lagged behind maize and wheat in genomics-assisted breeding. For the millions of families whose primary food source falls short on vitamin A, faster breeding is not an abstraction—it is the difference between deficiency and health arriving sooner rather than later. With genomic prediction accuracies approaching r = 0.8 now achievable even in small breeding populations, cassava biofortification may finally be able to move at the speed the problem demands.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Genomic prediction of pro-vitamin A carotenoid (beta-carotene) content in cassava roots using GWAS-informed parametric Bayesian models and machine learning ensemble approaches</p>
<p><strong>Article Title:</strong> Multi-locus random marker effects GWAS, ensemble models and variable selection improve genomic prediction for pro-vitamin A carotenoids in cassava</p>
<p><strong>Article References:</strong> Abincha, W., Dzidzienyo, D. K., Tongoona, P., Ofori, K., Owor, B.-E., Kayondo, I. S., Ozimati, A., Mwale, S. E., &amp; Kivuva, B. M. (2026). Multi-locus random marker effects GWAS, ensemble models and variable selection improve genomic prediction for pro-vitamin A carotenoids in cassava. <em>Heliyon, 12</em>(14), Article e45360. <a href="https://doi.org/10.1016/j.heliyon.2026.e45360" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/j.heliyon.2026.e45360</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.heliyon.2026.e45360" target="_blank" rel="noopener noreferrer">10.1016/j.heliyon.2026.e45360</a></p>
<p><strong>Keywords:</strong> cassava biofortification, genomic prediction, pro-vitamin A carotenoids, beta-carotene, multi-locus GWAS, machine learning, Random Forest, ensemble models, variable selection, vitamin A deficiency, SNP markers, plant breeding</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">185362</post-id>	</item>
		<item>
		<title>United We Grow: Innovative Data Method Boosts Accuracy of Plant Predictions</title>
		<link>https://scienmag.com/united-we-grow-innovative-data-method-boosts-accuracy-of-plant-predictions/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Tue, 13 May 2025 16:32:23 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[academic data trusteeship in agriculture]]></category>
		<category><![CDATA[deep learning in agriculture]]></category>
		<category><![CDATA[enhancing crop yield predictions]]></category>
		<category><![CDATA[gene-by-environment interactions]]></category>
		<category><![CDATA[genomic prediction in crop breeding]]></category>
		<category><![CDATA[high-dimensional genetic data analysis]]></category>
		<category><![CDATA[innovative agricultural research methodologies]]></category>
		<category><![CDATA[integration of diverse agricultural datasets]]></category>
		<category><![CDATA[non-linear transformations in genetics]]></category>
		<category><![CDATA[overcoming data silos in research]]></category>
		<category><![CDATA[phenotypic traits prediction]]></category>
		<category><![CDATA[wheat breeding programs collaboration]]></category>
		<guid isPermaLink="false">https://scienmag.com/united-we-grow-innovative-data-method-boosts-accuracy-of-plant-predictions/</guid>

					<description><![CDATA[In recent years, the field of genomic prediction has undergone a transformative evolution, largely driven by advances in deep learning methodologies. Unlike traditional statistical models, which often rely on linear assumptions and pre-defined relationships, deep learning harnesses the power of flexible, non-linear transformations to capture complex patterns embedded within high-dimensional genetic data. This paradigm shift [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In recent years, the field of genomic prediction has undergone a transformative evolution, largely driven by advances in deep learning methodologies. Unlike traditional statistical models, which often rely on linear assumptions and pre-defined relationships, deep learning harnesses the power of flexible, non-linear transformations to capture complex patterns embedded within high-dimensional genetic data. This paradigm shift is particularly pertinent in crop breeding, where phenotypic traits such as yield, plant height, and heading date are influenced by intricate gene-by-environment interactions that conventional models struggle to accommodate effectively.</p>
<p>At the forefront of this cutting-edge research is a team from the Leibniz Institute of Plant Genetics and Crop Plant Research (IPK), who have undertaken a pioneering effort to integrate vast datasets spanning multiple wheat breeding programs. Acting as an academic data trustee, the IPK team successfully amalgamated information from four distinct breeding companies alongside extensive trial data accrued over a twelve-year period from various public-private partnerships. This unprecedented dataset collectively encompasses genotypic and phenotypic records from nearly 9,500 wheat genotypes evaluated across 168 diverse environmental conditions, providing a comprehensive foundation for genomic prediction endeavors.</p>
<p>One of the most formidable challenges the researchers faced was overcoming the notorious problem of data silos—isolated pools of proprietary company data that hamper large-scale analyses. By meticulously harmonizing heterogeneous phenotypic measurements and genotype-by-sequencing information, the team employed sophisticated data cleaning, standardization protocols, and imputation techniques to address missing single nucleotide polymorphism (SNP) markers. This careful curatorial process enabled the creation of a unified dataset amenable to advanced computational modelling, thereby facilitating cross-company collaboration without compromising data integrity or confidentiality.</p>
<p>Leveraging this rich resource, the researchers conducted an extensive comparative study between classical genomic prediction algorithms and modern deep learning frameworks based on artificial neural networks. Neural networks excel at discerning intricate, hierarchical patterns in structured datasets by iteratively adjusting internal parameters through backpropagation during model training. Crucially, the analyses demonstrated that by combining diverse test series flexibly, predictions could be substantially improved, reflecting a higher resolution understanding of genotype-to-phenotype links under varied environmental influences.</p>
<p>Further dissecting their findings, the team observed a pronounced positive correlation between the size of the training dataset and the accuracy of genomic predictions, which notably plateaued when the number of genotypes approached approximately 4,000. This saturation effect suggests diminishing returns beyond a critical dataset scale, highlighting the complexity of capturing all relevant variability using solely genotype information. Nevertheless, improvements continued marginally with larger data sizes, reaffirming the value of extensive genotype-environment trials in refining predictive accuracy.</p>
<p>Recognizing that genetic variation is only one piece of the puzzle, Prof. Dr. Jochen Reif and colleagues emphasized the importance of expanding environmental diversity within the dataset. Incorporating broader multi-location and multi-year trial data introduces vital context for environment-dependent trait expression, potentially breaking through the observed accuracy ceiling. This insight anchors their current initiative, the “Drive” project, launched in November 2024 and supported by the German Federal Ministry of Education and Research (BMBF), which aims to harness big data paradigms to revolutionize breeding research at scale.</p>
<p>Beyond the immediate improvements in predictive precision, the study provides a conceptual blueprint for dismantling entrenched data barriers within the agricultural sector. By assuming responsibility as a neutral academic trustee, the IPK team demonstrated that proprietary breeding data can be ethically shared and integrated without infringing on commercial interests. This model offers a promising route to collectively leverage data assets to accelerate genetic gain, ultimately fostering sustainable crop enhancement strategies vital for global food security.</p>
<p>The technical sophistication employed in this research reflects broader trends in computational plant biology where advanced machine learning tools are beginning to reshape how complex genotype-phenotype relationships are elucidated. Neural networks, with their adaptability to non-linear dynamics and capacity to exploit subtle epistatic interactions, represent a formidable toolkit for next-generation breeding pipelines. Their performance, however, is heavily contingent on algorithmic fine-tuning and the availability of large, well-curated training datasets encompassing both genetic markers and diverse environmental variables.</p>
<p>Moreover, the team’s approach underscores the challenges of integrating multi-source data with variable quality and completeness. Imputation of missing SNP variants, data normalization, and phenotype standardization require robust bioinformatics workflows to avoid propagating errors that could bias model outputs. The IPK group&#8217;s success in this regard highlights the critical role of data science expertise in complementing breeding and genomics to unlock meaningful biological insights from complex datasets.</p>
<p>Looking forward, the potential applications of this work extend far beyond wheat breeding. Similar frameworks could be adapted for other staple crops with complex trait architectures affected by multi-environment interactions. By fostering collaborative data sharing and harnessing state-of-the-art deep learning techniques, plant scientists can accelerate the development of climate-resilient, high-yielding crop varieties. This synergy between computational innovation and agricultural practice exemplifies the future of precision breeding in the age of big data.</p>
<p>In conclusion, the IPK-led study represents a significant milestone in genomic prediction research, showcasing how breaking down data silos and integrating large-scale, heterogeneous datasets can substantially advance predictive accuracy through deep learning. The initiative not only provides practical insights for improving wheat breeding programs but also sets the stage for harnessing big data’s full potential in plant science. As the “Drive” project progresses, the agricultural research community keenly anticipates further breakthroughs that blend computational prowess with biological wisdom to sustainably feed the world.</p>
<p>&#8212;</p>
<p><strong>Subject of Research</strong>: Genomic prediction in wheat breeding utilizing deep learning and data integration across companies.</p>
<p><strong>Article Title</strong>: Breaking down data silos across companies to train genome-wide predictions: A feasibility study in wheat</p>
<p><strong>News Publication Date</strong>: 20-Apr-2025</p>
<p><strong>Web References</strong>: http://dx.doi.org/10.1111/pbi.70095</p>
<p><strong>Keywords</strong>: deep learning, genomic prediction, wheat breeding, neural networks, data integration, SNP imputation, genotype-environment interaction, big data, plant phenotyping, agricultural data sharing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">44338</post-id>	</item>
	</channel>
</rss>
