<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>genomic prediction in crop breeding &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/genomic-prediction-in-crop-breeding/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 12 Sep 2026 19:14:45 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>genomic prediction in crop breeding &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Genomic tools promise faster sugarcane breeding, major review finds</title>
		<link>https://scienmag.com/genomic-tools-promise-faster-sugarcane-breeding-major-review-finds/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 19:14:45 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[advances in sugarcane breeding technology]]></category>
		<category><![CDATA[allele dosage]]></category>
		<category><![CDATA[breeding pipeline]]></category>
		<category><![CDATA[challenges of polyploidy in sugarcane genetics]]></category>
		<category><![CDATA[computational mate allocation in sugarcane]]></category>
		<category><![CDATA[cross prediction]]></category>
		<category><![CDATA[early-stage breeding decision optimization]]></category>
		<category><![CDATA[genetic complexity of sugarcane genome]]></category>
		<category><![CDATA[genetic diversity in sugarcane hybrids]]></category>
		<category><![CDATA[genetic gain]]></category>
		<category><![CDATA[genetic improvement strategies for sugarcane]]></category>
		<category><![CDATA[genomic prediction in crop breeding]]></category>
		<category><![CDATA[genomic selection]]></category>
		<category><![CDATA[genotype by environment interaction]]></category>
		<category><![CDATA[impact of genomic tools on sugarcane productivity]]></category>
		<category><![CDATA[mate allocation]]></category>
		<category><![CDATA[mixed models]]></category>
		<category><![CDATA[non-additive effects]]></category>
		<category><![CDATA[Polyploidy]]></category>
		<category><![CDATA[role of genomics in accelerating sugarcane genetic gain]]></category>
		<category><![CDATA[Saccharum]]></category>
		<category><![CDATA[sugarcane]]></category>
		<category><![CDATA[sugarcane genomics]]></category>
		<category><![CDATA[sustainable bioenergy crop development]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197752</guid>

					<description><![CDATA[A major review argues that moving genomic prediction and computational mate allocation to the earliest stages of sugarcane breeding could dramatically accelerate genetic gain in this polyploid crop.]]></description>
										<content:encoded><![CDATA[<p>Sugarcane is one of the world&#8217;s most important crops, underpinning global sugar supplies, bioethanol production and, increasingly, renewable biomass for materials and energy. Yet the crop remains notoriously difficult to improve genetically. A comprehensive new review published in Theoretical and Applied Genetics synthesises decades of statistical and genomic research to explain why sugarcane breeding advances so slowly, and maps out a strategy for accelerating genetic gain using genomic prediction and computational mate allocation. The authors, led by Andrew Rigby of Sugar Research Australia and colleagues including Ben Hayes and Lee Hickey at the University of Queensland, argue that the biggest untapped opportunities lie not in the final stages of clonal testing, where genomic selection is already entering routine use, but in the earliest decisions a breeding programme makes: which crosses to make, which families to advance, and which parents to recycle.</p>
<p>The central obstacle is sugarcane&#8217;s extraordinary genome. Modern cultivars are highly polyploid and frequently aneuploid, typically carrying on the order of 100 to 130 chromosomes, with local copy number varying between roughly six and fourteen copies per homoeologous group. They descend from interspecific hybridisation between the domesticated Saccharum officinarum and the wild S. spontaneum, a history of so-called nobilisation that delivered landmark disease-resistant varieties but left modern germplasm with a narrow founder base, extensive linkage disequilibrium and strong non-additive genetic effects. Recent work, including the highly contiguous polyploid reference genome of the cultivar R570 published in Nature in 2024, has resolved individual haplotypes across approximately 12 chromosome copies, opening the door to far more precise marker development and, potentially, to attributing epistatic interactions to specific genomic regions rather than anonymous marker pairs.</p>
<p>These biological realities collide with an unusually long and structured breeding pipeline. In the Australian system operated by Sugar Research Australia, controlled crosses produce true seed that enters progeny assessment trials, where entire families are ranked and parents selected backwards on the basis of family plot means. Surviving individuals advance to clonal assessment trials and then to multi-environment final assessment trials spanning plant and repeated ratoon crops. The whole journey from cross to commercial release spans a decade or more. Crucially, the review notes, information currently passes between these stages almost entirely through selection decisions: the same parents appear as contributors to family means, as clonal entries and as check clones, yet their breeding values are rarely refined by a model that considers all of these contributions simultaneously.</p>
<p>The statistical challenges of the early stages are substantial and, the authors argue, often underappreciated. Family plots in progeny assessment trials contain multiple unique seedlings from the same full-sib family, so repeated plots are samples from the same cross population rather than true biological replicates. Treating them as replicates of a single genetic entity distorts variance partitioning and inflates error variance, undermining heritability estimates and selection accuracy. Appropriate models must scale Mendelian sampling variance for family plot means. Beyond this, single-row plots are highly sensitive to spatial heterogeneity and interplot competition, which has a heritable component that can be modelled as indirect genetic effects; ignoring it risks selecting genotypes that are aggressive competitors in trials but underperform in commercial single-variety stands. Crop cycle, year and environment are also frequently confounded, complicating genotype-by-environment modelling that relies heavily on factor analytic mixed models.</p>
<p>Genomic selection, which predicts genetic merit from genome-wide markers, has matured rapidly in sugarcane since early proof-of-concept studies in 2013. GBLUP and RR-BLUP models remain robust baselines for the highly polygenic traits that dominate commercial value, including tonnes of cane per hectare, commercial cane sugar and fibre percentage. More complex approaches, including Bayesian variable-selection models, kernel methods and machine learning, have delivered incremental and trait-specific gains. Modelling non-additive effects through dominance and epistatic relationship matrices has improved clonal prediction in some settings, but the review cautions that in populations with strong relatedness and extended linkage disequilibrium, additive relationship matrices can partially absorb non-additive variation, blurring the distinction between the total genetic value relevant to clonal selection and the additive breeding value that drives long-term response in parents.</p>
<p>Allele dosage is a critical unresolved complication. Most operational genomic prediction in sugarcane has relied on pseudo-diploid encoding of single-dose markers, which discards copy-number information. Dosage-aware relationship matrices, formalised for autotetraploid potato and extended to higher ploidies, can dramatically outperform diploidised encodings in simulations with many multi-dose heterozygotes and strong dominance, yet show little advantage in populations dominated by simplex markers, which is precisely the situation in many current sugarcane panels. Continuous genotype representations based on normalised array intensities have delivered modest accuracy gains of roughly five to seven percent for commercial cane sugar and fibre, though not for cane yield. Aneuploidy adds a further layer, because nominal ploidy does not define local ploidy at any given marker, meaning that standard dosage models may be misspecified across a substantial fraction of loci.</p>
<p>The review&#8217;s most forward-looking contribution is its treatment of genomic cross prediction and mate allocation. Because crossing capacity, field space and flowering synchrony all limit how many parental combinations can be tested, the choice of crosses shapes the genetic variance entering the pipeline. Sugarcane breeders have long relied on empirically proven crosses, implicitly capturing parental merit and favourable non-additive interactions, but at the cost of reduced exploration of new combinations and heightened inbreeding risk. Predicting the expected mean and within-family genetic variance of untested crosses, and allocating matings under constraints on relatedness, flowering compatibility and seed inventory, could allow programmes to balance short-term gain against long-term diversity. Tools such as AlphaMate and SimpleMating already exist, but the authors stress that full polyploid-aware implementation remains a research priority, hindered by unresolved phasing at approximately twelvefold dosage and by variance predictions that are consistently less accurate than mean predictions.</p>
<p>Lessons from maize, wheat, cassava and potato transfer only partially. Hybrid crops with defined heterotic groups, fixed ploidy and inbred parents present a fundamentally different prediction problem from a clonally propagated, outbred polyploid in which individual-level non-additive effects are not fixed in the product. The usefulness criterion combining cross mean and within-cross variance, validated in diploid and tetraploid systems, may underperform in sugarcane without polyploid-specific adaptation. Nonetheless, training-population design principles hold: representing key founders broadly across many families outperforms deep sampling of a few, and validation schemes should mirror how prediction will actually be used, whether ranking related candidates within a cycle or forecasting performance in future environments.</p>
<p>The authors conclude with a practical research agenda. Family plots should be modelled as means of sampled full sibs rather than replicated genotypes. Environment definitions must distinguish region and year to capture stability under climate variability. Competition should be modelled explicitly in multi-entrial frameworks. Most ambitiously, progeny, clonal and final assessment trial records should be integrated within stage-integrated mixed-model analyses that combine pedigree and genomic relationships, allowing uncertainty to propagate coherently from family evaluation through clonal testing to parent recycling and cross design. Such integration, the review argues, is not merely a statistical refinement but a fundamental shift in breeding strategy, one that could deliver faster, more robust and more sustainable genetic improvement for a crop on which global food and energy systems increasingly depend.</p>
<p><strong>Subject of Research:</strong> Statistical and genomic strategies to accelerate sugarcane genetic improvement, from family trials to genomic mate allocation</p>
<p><strong>Article Title:</strong> From family trials to genomic mate allocation: statistical and genomic strategies to accelerate sugarcane genetic improvement</p>
<p><strong>Article References:</strong> Rigby, A., Atkin, F., Hayes, B., Hickey, L., &amp; Yadav, S. (2026). From family trials to genomic mate allocation: statistical and genomic strategies to accelerate sugarcane genetic improvement. <em>Theoretical and Applied Genetics, 139</em>(10), Article 266. <a href="https://doi.org/10.1007/s00122-026-05373-9" rel="noopener noreferrer">https://doi.org/10.1007/s00122-026-05373-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00122-026-05373-9" rel="noopener noreferrer">10.1007/s00122-026-05373-9</a></p>
<p><strong>Keywords:</strong> sugarcane, genomic selection, polyploidy, allele dosage, mate allocation, cross prediction, genetic gain, mixed models, genotype-by-environment interaction, Saccharum, breeding pipeline, non-additive effects</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197752</post-id>	</item>
		<item>
		<title>GWAS and genomic prediction reveal genes controlling tomato fruit weight</title>
		<link>https://scienmag.com/gwas-and-genomic-prediction-reveal-genes-controlling-tomato-fruit-weight/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Mon, 07 Sep 2026 09:25:52 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[Advances in tomato quantitative trait analysis]]></category>
		<category><![CDATA[Curated tomato accessions for genetic study]]></category>
		<category><![CDATA[Evolutionary fingerprints in tomato genome]]></category>
		<category><![CDATA[Genetic architecture of tomato size]]></category>
		<category><![CDATA[Genetic architecture of tomato yield traits]]></category>
		<category><![CDATA[Genetic markers for tomato breeding]]></category>
		<category><![CDATA[Genome-wide association studies in tomatoes]]></category>
		<category><![CDATA[genomic prediction in crop breeding]]></category>
		<category><![CDATA[Genomic prediction of crop traits]]></category>
		<category><![CDATA[Identification of fruit size genes in tomato]]></category>
		<category><![CDATA[Identification of genes controlling tomato size]]></category>
		<category><![CDATA[Impact of domestication on tomato genome]]></category>
		<category><![CDATA[Influence of DNA on tomato fruit size]]></category>
		<category><![CDATA[Role of DNA in tomato size determination]]></category>
		<category><![CDATA[Role of GWAS in crop improvement]]></category>
		<category><![CDATA[Tomato accessions and genetic diversity]]></category>
		<category><![CDATA[Tomato breeding for fruit weight]]></category>
		<category><![CDATA[Tomato domestication and evolution]]></category>
		<category><![CDATA[Tomato fruit weight genetics]]></category>
		<category><![CDATA[Wild and cultivated tomato genetic diversity]]></category>
		<category><![CDATA[Wild versus cultivated tomato genetics]]></category>
		<guid isPermaLink="false">https://scienmag.com/gwas-and-genomic-prediction-reveal-genes-controlling-tomato-fruit-weight/</guid>

					<description><![CDATA[In the race to feed a warming planet, scientists have long sought to decode the genetic instructions that make some crops flourish while others falter. Now, an international team of researchers has taken a major step forward in that quest, mapping the intricate genetic architecture that governs fruit weight in tomato—one of the world&#8217;s most [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In the race to feed a warming planet, scientists have long sought to decode the genetic instructions that make some crops flourish while others falter. Now, an international team of researchers has taken a major step forward in that quest, mapping the intricate genetic architecture that governs fruit weight in tomato—one of the world&#8217;s most economically important vegetable crops. The study, published in the journal Theoretical and Applied Genetics, reveals that the size of the tomato on your kitchen counter is overwhelmingly dictated by its DNA rather than by the environment in which it grew, and that the evolutionary journey from wild berry to beefsteak has left detectable fingerprints throughout the tomato genome.</p>
<p>The research, led by scientists affiliated with multiple institutions across the United States and Europe, harnessed the power of the Varitome population—a curated collection of 166 tomato accessions that spans the full arc of tomato domestication. This panel includes 28 accessions of the wild progenitor Solanum pimpinellifolium (SP), 117 accessions of the semi-domesticated Solanum lycopersicum var. cerasiforme (SLC), and 21 accessions of the fully cultivated Solanum lycopersicum var. lycopersicum (SLL). By studying these three groups side by side, the team was able to trace how thousands of years of human selection have sculpted the genetic landscape of a crop that now ranks among the most valuable horticultural commodities on Earth.</p>
<p>The team grew all accessions across four dramatically different environments: Antalya in Türkiye, Valencia in Spain, and Florida and Georgia in the United States. These sites represent contrasting agroecological zones, providing a rigorous test of how stable the genetic control of fruit weight truly is. The plants were arranged in randomized complete block designs, and fruit weight data were collected and normalized using a base-10 logarithmic transformation to stabilize variance. The researchers then applied a multi-environment linear mixed model, implemented through restricted maximum likelihood, to partition the observed phenotypic variance into genetic, environmental, genotype-by-environment interaction, and residual components.</p>
<p>What they found was striking. Genotype alone explained more than 92 percent of the phenotypic variance in fruit weight, while environment and genotype-by-environment interaction contributed only marginal effects. The broad-sense heritability estimate reached 0.98, and the overall model fit was exceptionally high with an R-squared of 0.96. Within individual replicated environments, heritability estimates exceeded 0.97 in both Florida and Georgia. These numbers confirm that tomato fruit weight is governed primarily by stable additive genetic effects rather than by environmental plasticity—a finding that has immediate practical implications for breeders attempting to predict fruit weight across different growing regions and seasons.</p>
<p>But the researchers did not stop at confirming what many suspected. Their next move was to dissect the molecular machinery underlying this strong genetic control, and here they deployed an approach that goes well beyond the standard toolbox of modern plant genetics. Rather than relying solely on single nucleotide polymorphisms—the workhorse markers of most genome-wide association studies—the team analyzed three distinct classes of genomic variation simultaneously: SNPs, insertions and deletions (INDELs), and structural variants (SVs). SNPs represent single-letter changes in the DNA sequence, while INDELs involve the insertion or deletion of one to several hundred base pairs. Structural variants are even larger genomic rearrangements—deletions, duplications, inversions, and transpositions that can span thousands of base pairs and alter gene dosage, chromatin structure, or regulatory sequences.</p>
<p>This multi-variant strategy proved powerful. Through genome-wide association analyses conducted across all four environments using both the TASSEL and BLINK statistical frameworks with false discovery rate correction, the team identified 15 significant SNP loci, 10 INDEL loci, and 10 SV loci associated with fruit weight. Crucially, several of these associations co-localized with genes already known to govern tomato fruit size—including fw2.2/CELL NUMBER REGULATOR, fw3.2/KLUH, fw11.3/CELL SIZE REGULATOR, lc/WUSCHEL, and fas/CLAVATA3. These genes operate through diverse biological mechanisms, from regulating cell division during fruit development to controlling the organization of the floral meristem, the structure that determines how many seed-producing compartments a fruit will contain.</p>
<p>Among all the associations detected, chromosome 5 stood out as particularly notable. A 390-base-pair deletion at position 45,551,023 was detected in all four environments, while an INDEL at position 55,361,233 appeared in three of the four environments. The consistency of these signals across geographically and climatically distinct trial sites suggests they represent stable, improvement-associated candidate loci that could be validated for use in marker-assisted selection programs. For breeders, identifying such robust markers is invaluable because it means the markers can be trusted regardless of where or when the crop is grown.</p>
<p>The study also revealed a fascinating evolutionary pattern. When the researchers tracked allele frequencies across the three domestication groups—from wild SP through semi-domesticated SLC to fully cultivated SLL—they observed clear and consistent shifts at multiple loci. Some alleles showed progressive frequency changes consistent with selection during domestication and subsequent crop improvement, while others were detected almost exclusively in cultivated germplasm, suggesting more recent origins tied to modern breeding programs. This trajectory is consistent with prior genomic work showing that domesticated SLL tomatoes retain only about 22.6 percent of the standing genetic diversity found in wild SP populations, while semi-domesticated SLC accessions retain approximately 53.8 percent. Severe bottlenecks during domestication and targeted selection for larger fruits have narrowed allelic diversity precisely in the genomic regions that control fruit weight, creating a paradox for breeders: the genes that matter most are also those with the least remaining variation in elite germplasm.</p>
<p>To translate these genetic insights into practical breeding tools, the team built five distinct genomic prediction models. Model 1 used only SNPs, Model 2 only INDELs, Model 3 only SVs, Model 4 combined all three marker types additively, and Model 5 incorporated genotype-by-environment interaction kernels on top of the combined markers. Each model was tested using four complementary cross-validation scenarios that simulate different breeding situations: CV1 predicts tested genotypes in tested environments, CV2 predicts untested genotypes in tested environments, CV0 predicts tested genotypes in untested environments, and CV00—the most demanding scenario—predicts untested genotypes in untested environments. Each scenario was repeated 100 times with random 70-30 training-testing splits to ensure robust estimates.</p>
<p>The prediction algorithms themselves were also diverse, spanning four analytical frameworks: Bayesian Genomic Linear Regression (BGLR), Partial Least Squares regression (PLS), Random Forest (RF), and Deep Learning (DL) neural networks. BGLR used reproducing kernel Hilbert space regression with 5,000 MCMC iterations, PLS employed kernel-based latent variable decomposition through the SKM package, Random Forest used 500 regression trees, and the deep learning model used a fully connected neural network with ReLU activation, L2 regularization, dropout, and the Adam optimizer implemented through TensorFlow and Keras.</p>
<p>When the dust settled from thousands of model runs, two patterns emerged clearly. First, the choice of validation scenario and prediction algorithm mattered far more than the type of genetic marker used. Partial Least Squares and BGLR consistently outperformed Random Forest and Deep Learning models, particularly under the most stringent cross-validation scenarios involving untested genotypes and untested environments. Second, while differences among marker classes were generally modest, INDEL-based models frequently achieved the highest prediction accuracies under the most challenging scenarios. This is a noteworthy finding because INDELs, especially those falling within coding sequences or promoter regions, are more likely than SNPs to have direct functional consequences on gene expression or protein structure. Their predictive power suggests they may be capturing signals that SNPs alone cannot fully resolve.</p>
<p>The integration of all three variant types into a single model—Model 4—improved biological interpretation and enabled more comprehensive identification of domestication- and improvement-associated loci than any single marker class could achieve alone. The researchers also estimated genomic heritability separately for each marker class using Bayesian models with marker-specific genomic relationship matrices, finding that each class captured a meaningful but partially overlapping slice of the genetic variance underlying fruit weight.</p>
<p>Principal component analysis of the SNP data painted a vivid picture of the domestication gradient. The first principal component, explaining 20.12 percent of total genetic variation, cleanly separated wild SP accessions from both SLC and SLL groups. The second component, accounting for an additional 8.65 percent, captured diversity within species, particularly among the genetically diverse wild accessions. SLL accessions clustered tightly near the center of the plot, reflecting the severe reduction in genetic diversity that accompanied domestication and modern breeding, while SLC occupied an intermediate position, consistent with its role as a transitional population.</p>
<p>Cross-environment correlation analyses added another layer of biological insight. Fruit weight showed strong correlations across environments for SLC and SLL accessions, with Pearson correlation coefficients ranging from 0.76 to 0.99, indicating that genotype rankings remain stable across different growing conditions. In contrast, wild SP accessions exhibited low or non-significant correlations ranging from 0.37 to 0.72, revealing that their fruit weight is highly environment-dependent. Finlay-Wilkinson regression confirmed this pattern: cultivated SLL accessions displayed nearly parallel reaction norms with shallow slopes, indicating strong environmental buffering, while wild SP accessions showed steep and highly variable slopes, demonstrating pronounced environmental sensitivity.</p>
<p>The implications of this work extend well beyond the tomato field. By demonstrating that a multi-variant genomic framework—incorporating SNPs, INDELs, and SVs alongside genotype-by-environment interaction kernels—can simultaneously dissect genetic architecture and deliver accurate predictions across environments and genetic backgrounds, the study provides a template that could be applied to other crops shaped by similar domestication histories. For a world that will need to produce substantially more food on limited land under increasingly unpredictable climatic conditions, the ability to predict complex traits like fruit weight with precision, using genomic data rather than laborious field phenotyping, represents a meaningful acceleration of the breeding cycle. The humble tomato, it turns out, has much to teach us about the future of agriculture.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Genetic architecture and genomic prediction of tomato fruit weight across the domestication continuum using multi-variant GWAS (SNPs, INDELs, and structural variants)</p>
<p><strong>Article Title:</strong> Multi-variant GWAS and genomic prediction dissect the genetic architecture underlying tomato fruit weight</p>
<p><strong>Article References:</strong> Topcu, Y., Adak, A., Kayikci, H. C., Aydin, S., Yildiz, K., Ramos, A., Tieman, D. M., Visa, S., van der Knaap, E., &amp; Sapkota, M. (2026). Multi-variant GWAS and genomic prediction dissect the genetic architecture underlying tomato fruit weight. <em>Theoretical and Applied Genetics, 139</em>(9), Article 259. <a href="https://doi.org/10.1007/s00122-026-05326-2" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s00122-026-05326-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00122-026-05326-2" target="_blank" rel="noopener noreferrer">10.1007/s00122-026-05326-2</a></p>
<p><strong>Keywords:</strong> tomato, fruit weight, GWAS, genomic prediction, SNPs, INDELs, structural variants, domestication, Solanum pimpinellifolium, Solanum lycopersicum, genotype-by-environment interaction, heritability, Varitome</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">189330</post-id>	</item>
		<item>
		<title>GWAS-informed ensembles boost genomic prediction of vitamin A carotenoids in cassava</title>
		<link>https://scienmag.com/gwas-informed-ensembles-boost-genomic-prediction-of-vitamin-a-carotenoids-in-cassava/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sun, 30 Aug 2026 08:21:45 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[accelerated breeding cycles]]></category>
		<category><![CDATA[advanced computational tools in agriculture]]></category>
		<category><![CDATA[biofortification for public health]]></category>
		<category><![CDATA[biofortified cassava varieties]]></category>
		<category><![CDATA[Cassava biofortification]]></category>
		<category><![CDATA[computational approaches to plant biofortification]]></category>
		<category><![CDATA[crop genetic diversity]]></category>
		<category><![CDATA[enhancing crop nutritional traits through genomics]]></category>
		<category><![CDATA[genetic markers for carotenoid content]]></category>
		<category><![CDATA[genome-wide association studies in crop breeding]]></category>
		<category><![CDATA[genomic estimated breeding values]]></category>
		<category><![CDATA[genomic prediction in crop breeding]]></category>
		<category><![CDATA[genomic prediction of vitamin A content]]></category>
		<category><![CDATA[GWAS and machine learning in plant breeding]]></category>
		<category><![CDATA[GWAS-informed machine learning models]]></category>
		<category><![CDATA[improving cassava nutritional quality]]></category>
		<category><![CDATA[improving nutrient content in staple crops]]></category>
		<category><![CDATA[plant genetics and genomics]]></category>
		<category><![CDATA[public health impact of vitamin A-rich crops]]></category>
		<category><![CDATA[rapid cassava breeding techniques]]></category>
		<category><![CDATA[sub-Saharan Africa staple crop improvement]]></category>
		<category><![CDATA[sub-Saharan Africa staple crops]]></category>
		<category><![CDATA[vitamin A carotenoid content in cassava]]></category>
		<guid isPermaLink="false">https://scienmag.com/gwas-informed-ensembles-boost-genomic-prediction-of-vitamin-a-carotenoids-in-cassava/</guid>

					<description><![CDATA[Cassava is a lifeline for millions of people across sub-Saharan Africa, but it is an imperfect lifeline. The starchy roots that anchor diets in Kenya, Uganda, Nigeria and beyond are notoriously poor sources of micronutrients, and vitamin A deficiency remains rampant in communities that depend on the crop as a staple. Biofortification—breeding cassava varieties whose [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Cassava is a lifeline for millions of people across sub-Saharan Africa, but it is an imperfect lifeline. The starchy roots that anchor diets in Kenya, Uganda, Nigeria and beyond are notoriously poor sources of micronutrients, and vitamin A deficiency remains rampant in communities that depend on the crop as a staple. Biofortification—breeding cassava varieties whose roots are rich in pro-vitamin A carotenoids—has long been championed as a public health solution. Now, a team of researchers working with Kenyan and Ugandan germplasm reports a substantial leap forward in the computational machinery of that effort, showing that a smarter combination of genome-wide association studies and machine learning can predict the nutritional quality of cassava roots with unprecedented accuracy.</p>
<p>The core obstacle is time. Cassava takes roughly twelve months to mature, which means breeders who wait to measure carotenoid content in harvested roots must commit years to each selection cycle. Genomic prediction offers a way around this bottleneck. The idea is to build statistical models that link tens of thousands of DNA markers to measured traits, then use those models to calculate genomic estimated breeding values—predictions of an individual&#8217;s additive genetic worth as a parent. Breeders can then select the most promising progenitors at the seedling stage, months before a single root is dug from the ground, compressing the breeding cycle and accelerating genetic gain.</p>
<p>In a study conducted at the Kenya Agricultural and Livestock Research Organization (KALRO) in Kakamega, researchers assembled a panel of 93 pro-vitamin A cassava genotypes derived from accessions held by Uganda&#8217;s National Crop Resource Research Institute. The plants were established as seedlings in January 2023, and in January 2024 two representative roots from each genotype were harvested under low-light evening conditions—a necessary precaution, because carotenoids degrade rapidly when exposed to light. Samples were rushed to the University of Nairobi within seven hours, where high-performance liquid chromatography quantified beta-carotene content in freshly harvested tissue. The measured values ranged from 0.04 to 1.03 mg per 100 g of root, with a mean of 0.52 mg/100 g. Notably, 68 percent of the population fell below the Institute of Medicine&#8217;s recommended benchmark of 6 micrograms of beta-carotene per gram, underscoring how much work remains for breeders.</p>
<p>The genotyping side of the study relied on DArTseq genotyping-by-sequencing against the Manihot esculenta 671_v8.0 reference genome, yielding 54,877 single nucleotide polymorphisms. After quality control filtering for call rate, minor allele frequency and heterozygosity using the snpReady package in R, 37,419 high-quality SNPs remained for downstream analysis. Five-fold cross-validation, in which 80 percent of the data trained each model and the remaining 20 percent served as an independent test, was used to evaluate prediction performance across all models.</p>
<p>The researchers&#8217; central innovation lay in how they identified trait-associated markers to feed into their prediction models. Conventional single-locus GWAS approaches, while widely used, suffer from limited statistical power in small populations—a persistent constraint in crop breeding programs. The team instead turned to multi-locus random marker effect GWAS (mr-GWAS), implemented through four algorithms in the mrMLM platform: mrMLM, FASTmrMLM, pLARmEB and ISIS EM-BLASSO. These methods treat SNP effects as random, shrinking estimates toward zero and applying a modified Bonferroni significance threshold of 0.05 divided by the effective number of markers, which corrects for multiple testing without the punishing conservatism of standard corrections. The payoff was clear: while the BLINK fixed-effect GWAS model detected only two significant SNPs for beta-carotene (on chromosomes 09 and 14) and one for flesh color on chromosome 01, the mr-GWAS pipeline uncovered five significant SNPs distributed across chromosomes 01, 03, 04, 14 and 18.</p>
<p>When these markers were incorporated as fixed-effect covariates into classical Bayesian genomic prediction models—Bayesian Ridge Regression, Bayesian Lasso, BayesA, BayesB and BayesC—the improvement was dramatic. Naive models struggled badly: prediction correlations on the validation set ranged from just 0.14 to 0.16 for beta-carotene. Adding the two BLINK SNPs lifted accuracy to between 0.41 and 0.43. But when all five mr-GWAS-derived SNPs were included, Bayesian Lasso reached a prediction ability of r = 0.81, with the remaining models clustered between 0.75 and 0.78. The pattern held for root flesh color as well, where a single chromosome 01 SNP boosted validation correlations from roughly 0.39–0.48 to 0.61–0.63. The lesson is that markers spanning more of the genome are more likely to be linked to the causal genes governing carotenoid accumulation, capturing a fuller share of the trait&#8217;s genetic architecture.</p>
<p>The study did not stop with parametric models, however. Traditional genomic prediction assumes purely additive gene action, yet real traits are shaped by dominance and epistasis—non-additive interactions that linear models systematically miss. To capture these effects, the researchers built non-parametric models using five machine learning algorithms: Random Forest, Support Vector Machine with radial kernel, Neural Networks, Extreme Gradient Boosting (XGBoost) and K-Nearest Neighbors, along with a stacked ensemble that used Random Forest as the meta-learner. Initially, these models performed modestly, with the ensemble topping out at r = 0.44 and the support vector machine limping along at r = 0.04.</p>
<p>The transformation came through feature selection. Using the Boruta algorithm, the team distilled 37,419 SNPs down to just 13 of the most informative predictors—a 99.97 percent reduction in dimensionality. With this streamlined marker set, performance soared: Random Forest and the ensemble model both achieved r = 0.79, while XGBoost and K-Nearest Neighbors reached r = 0.73. This result carries a practical message for computational breeders everywhere—curating the predictor set matters as much as choosing the algorithm, since removing noise variables sharpens model focus, reduces computational burden and prevents overfitting.</p>
<p>The distinction between the two modeling families maps directly onto breeding strategy. Genomic estimated breeding values from the parametric models reflect additive effects, the only genetic component parents reliably transmit to offspring, making them the right tool for early progenitor selection. Non-parametric machine learning models, by contrast, estimate total genomic value—additive and non-additive combined—which better predicts how a clone will actually perform as a finished variety. Together, the two approaches let breeders pick parents early and pick varieties accurately, shortening every stage of the pipeline from crossing to release.</p>
<p>The authors acknowledge limitations, including the modest population size of 93 genotypes and the fact that GWAS was performed on the full dataset before cross-validation, which may have introduced some optimistic bias into reported accuracies. They recommend larger datasets and nested cross-validation designs in future work. Even so, the results represent a meaningful advance for a crop that has long lagged behind maize and wheat in genomics-assisted breeding. For the millions of families whose primary food source falls short on vitamin A, faster breeding is not an abstraction—it is the difference between deficiency and health arriving sooner rather than later. With genomic prediction accuracies approaching r = 0.8 now achievable even in small breeding populations, cassava biofortification may finally be able to move at the speed the problem demands.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Genomic prediction of pro-vitamin A carotenoid (beta-carotene) content in cassava roots using GWAS-informed parametric Bayesian models and machine learning ensemble approaches</p>
<p><strong>Article Title:</strong> Multi-locus random marker effects GWAS, ensemble models and variable selection improve genomic prediction for pro-vitamin A carotenoids in cassava</p>
<p><strong>Article References:</strong> Abincha, W., Dzidzienyo, D. K., Tongoona, P., Ofori, K., Owor, B.-E., Kayondo, I. S., Ozimati, A., Mwale, S. E., &amp; Kivuva, B. M. (2026). Multi-locus random marker effects GWAS, ensemble models and variable selection improve genomic prediction for pro-vitamin A carotenoids in cassava. <em>Heliyon, 12</em>(14), Article e45360. <a href="https://doi.org/10.1016/j.heliyon.2026.e45360" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/j.heliyon.2026.e45360</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.heliyon.2026.e45360" target="_blank" rel="noopener noreferrer">10.1016/j.heliyon.2026.e45360</a></p>
<p><strong>Keywords:</strong> cassava biofortification, genomic prediction, pro-vitamin A carotenoids, beta-carotene, multi-locus GWAS, machine learning, Random Forest, ensemble models, variable selection, vitamin A deficiency, SNP markers, plant breeding</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">185362</post-id>	</item>
		<item>
		<title>United We Grow: Innovative Data Method Boosts Accuracy of Plant Predictions</title>
		<link>https://scienmag.com/united-we-grow-innovative-data-method-boosts-accuracy-of-plant-predictions/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Tue, 13 May 2025 16:32:23 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[academic data trusteeship in agriculture]]></category>
		<category><![CDATA[deep learning in agriculture]]></category>
		<category><![CDATA[enhancing crop yield predictions]]></category>
		<category><![CDATA[gene-by-environment interactions]]></category>
		<category><![CDATA[genomic prediction in crop breeding]]></category>
		<category><![CDATA[high-dimensional genetic data analysis]]></category>
		<category><![CDATA[innovative agricultural research methodologies]]></category>
		<category><![CDATA[integration of diverse agricultural datasets]]></category>
		<category><![CDATA[non-linear transformations in genetics]]></category>
		<category><![CDATA[overcoming data silos in research]]></category>
		<category><![CDATA[phenotypic traits prediction]]></category>
		<category><![CDATA[wheat breeding programs collaboration]]></category>
		<guid isPermaLink="false">https://scienmag.com/united-we-grow-innovative-data-method-boosts-accuracy-of-plant-predictions/</guid>

					<description><![CDATA[In recent years, the field of genomic prediction has undergone a transformative evolution, largely driven by advances in deep learning methodologies. Unlike traditional statistical models, which often rely on linear assumptions and pre-defined relationships, deep learning harnesses the power of flexible, non-linear transformations to capture complex patterns embedded within high-dimensional genetic data. This paradigm shift [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In recent years, the field of genomic prediction has undergone a transformative evolution, largely driven by advances in deep learning methodologies. Unlike traditional statistical models, which often rely on linear assumptions and pre-defined relationships, deep learning harnesses the power of flexible, non-linear transformations to capture complex patterns embedded within high-dimensional genetic data. This paradigm shift is particularly pertinent in crop breeding, where phenotypic traits such as yield, plant height, and heading date are influenced by intricate gene-by-environment interactions that conventional models struggle to accommodate effectively.</p>
<p>At the forefront of this cutting-edge research is a team from the Leibniz Institute of Plant Genetics and Crop Plant Research (IPK), who have undertaken a pioneering effort to integrate vast datasets spanning multiple wheat breeding programs. Acting as an academic data trustee, the IPK team successfully amalgamated information from four distinct breeding companies alongside extensive trial data accrued over a twelve-year period from various public-private partnerships. This unprecedented dataset collectively encompasses genotypic and phenotypic records from nearly 9,500 wheat genotypes evaluated across 168 diverse environmental conditions, providing a comprehensive foundation for genomic prediction endeavors.</p>
<p>One of the most formidable challenges the researchers faced was overcoming the notorious problem of data silos—isolated pools of proprietary company data that hamper large-scale analyses. By meticulously harmonizing heterogeneous phenotypic measurements and genotype-by-sequencing information, the team employed sophisticated data cleaning, standardization protocols, and imputation techniques to address missing single nucleotide polymorphism (SNP) markers. This careful curatorial process enabled the creation of a unified dataset amenable to advanced computational modelling, thereby facilitating cross-company collaboration without compromising data integrity or confidentiality.</p>
<p>Leveraging this rich resource, the researchers conducted an extensive comparative study between classical genomic prediction algorithms and modern deep learning frameworks based on artificial neural networks. Neural networks excel at discerning intricate, hierarchical patterns in structured datasets by iteratively adjusting internal parameters through backpropagation during model training. Crucially, the analyses demonstrated that by combining diverse test series flexibly, predictions could be substantially improved, reflecting a higher resolution understanding of genotype-to-phenotype links under varied environmental influences.</p>
<p>Further dissecting their findings, the team observed a pronounced positive correlation between the size of the training dataset and the accuracy of genomic predictions, which notably plateaued when the number of genotypes approached approximately 4,000. This saturation effect suggests diminishing returns beyond a critical dataset scale, highlighting the complexity of capturing all relevant variability using solely genotype information. Nevertheless, improvements continued marginally with larger data sizes, reaffirming the value of extensive genotype-environment trials in refining predictive accuracy.</p>
<p>Recognizing that genetic variation is only one piece of the puzzle, Prof. Dr. Jochen Reif and colleagues emphasized the importance of expanding environmental diversity within the dataset. Incorporating broader multi-location and multi-year trial data introduces vital context for environment-dependent trait expression, potentially breaking through the observed accuracy ceiling. This insight anchors their current initiative, the “Drive” project, launched in November 2024 and supported by the German Federal Ministry of Education and Research (BMBF), which aims to harness big data paradigms to revolutionize breeding research at scale.</p>
<p>Beyond the immediate improvements in predictive precision, the study provides a conceptual blueprint for dismantling entrenched data barriers within the agricultural sector. By assuming responsibility as a neutral academic trustee, the IPK team demonstrated that proprietary breeding data can be ethically shared and integrated without infringing on commercial interests. This model offers a promising route to collectively leverage data assets to accelerate genetic gain, ultimately fostering sustainable crop enhancement strategies vital for global food security.</p>
<p>The technical sophistication employed in this research reflects broader trends in computational plant biology where advanced machine learning tools are beginning to reshape how complex genotype-phenotype relationships are elucidated. Neural networks, with their adaptability to non-linear dynamics and capacity to exploit subtle epistatic interactions, represent a formidable toolkit for next-generation breeding pipelines. Their performance, however, is heavily contingent on algorithmic fine-tuning and the availability of large, well-curated training datasets encompassing both genetic markers and diverse environmental variables.</p>
<p>Moreover, the team’s approach underscores the challenges of integrating multi-source data with variable quality and completeness. Imputation of missing SNP variants, data normalization, and phenotype standardization require robust bioinformatics workflows to avoid propagating errors that could bias model outputs. The IPK group&#8217;s success in this regard highlights the critical role of data science expertise in complementing breeding and genomics to unlock meaningful biological insights from complex datasets.</p>
<p>Looking forward, the potential applications of this work extend far beyond wheat breeding. Similar frameworks could be adapted for other staple crops with complex trait architectures affected by multi-environment interactions. By fostering collaborative data sharing and harnessing state-of-the-art deep learning techniques, plant scientists can accelerate the development of climate-resilient, high-yielding crop varieties. This synergy between computational innovation and agricultural practice exemplifies the future of precision breeding in the age of big data.</p>
<p>In conclusion, the IPK-led study represents a significant milestone in genomic prediction research, showcasing how breaking down data silos and integrating large-scale, heterogeneous datasets can substantially advance predictive accuracy through deep learning. The initiative not only provides practical insights for improving wheat breeding programs but also sets the stage for harnessing big data’s full potential in plant science. As the “Drive” project progresses, the agricultural research community keenly anticipates further breakthroughs that blend computational prowess with biological wisdom to sustainably feed the world.</p>
<p>&#8212;</p>
<p><strong>Subject of Research</strong>: Genomic prediction in wheat breeding utilizing deep learning and data integration across companies.</p>
<p><strong>Article Title</strong>: Breaking down data silos across companies to train genome-wide predictions: A feasibility study in wheat</p>
<p><strong>News Publication Date</strong>: 20-Apr-2025</p>
<p><strong>Web References</strong>: http://dx.doi.org/10.1111/pbi.70095</p>
<p><strong>Keywords</strong>: deep learning, genomic prediction, wheat breeding, neural networks, data integration, SNP imputation, genotype-environment interaction, big data, plant phenotyping, agricultural data sharing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">44338</post-id>	</item>
	</channel>
</rss>
