<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>crop genetic diversity &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/crop-genetic-diversity/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 30 Aug 2026 08:21:48 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>crop genetic diversity &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>GWAS-informed ensembles boost genomic prediction of vitamin A carotenoids in cassava</title>
		<link>https://scienmag.com/gwas-informed-ensembles-boost-genomic-prediction-of-vitamin-a-carotenoids-in-cassava/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Sun, 30 Aug 2026 08:21:45 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[accelerated breeding cycles]]></category>
		<category><![CDATA[advanced computational tools in agriculture]]></category>
		<category><![CDATA[biofortification for public health]]></category>
		<category><![CDATA[biofortified cassava varieties]]></category>
		<category><![CDATA[Cassava biofortification]]></category>
		<category><![CDATA[computational approaches to plant biofortification]]></category>
		<category><![CDATA[crop genetic diversity]]></category>
		<category><![CDATA[enhancing crop nutritional traits through genomics]]></category>
		<category><![CDATA[genetic markers for carotenoid content]]></category>
		<category><![CDATA[genome-wide association studies in crop breeding]]></category>
		<category><![CDATA[genomic estimated breeding values]]></category>
		<category><![CDATA[genomic prediction in crop breeding]]></category>
		<category><![CDATA[genomic prediction of vitamin A content]]></category>
		<category><![CDATA[GWAS and machine learning in plant breeding]]></category>
		<category><![CDATA[GWAS-informed machine learning models]]></category>
		<category><![CDATA[improving cassava nutritional quality]]></category>
		<category><![CDATA[improving nutrient content in staple crops]]></category>
		<category><![CDATA[plant genetics and genomics]]></category>
		<category><![CDATA[public health impact of vitamin A-rich crops]]></category>
		<category><![CDATA[rapid cassava breeding techniques]]></category>
		<category><![CDATA[sub-Saharan Africa staple crop improvement]]></category>
		<category><![CDATA[sub-Saharan Africa staple crops]]></category>
		<category><![CDATA[vitamin A carotenoid content in cassava]]></category>
		<guid isPermaLink="false">https://scienmag.com/gwas-informed-ensembles-boost-genomic-prediction-of-vitamin-a-carotenoids-in-cassava/</guid>

					<description><![CDATA[Cassava is a lifeline for millions of people across sub-Saharan Africa, but it is an imperfect lifeline. The starchy roots that anchor diets in Kenya, Uganda, Nigeria and beyond are notoriously poor sources of micronutrients, and vitamin A deficiency remains rampant in communities that depend on the crop as a staple. Biofortification—breeding cassava varieties whose [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Cassava is a lifeline for millions of people across sub-Saharan Africa, but it is an imperfect lifeline. The starchy roots that anchor diets in Kenya, Uganda, Nigeria and beyond are notoriously poor sources of micronutrients, and vitamin A deficiency remains rampant in communities that depend on the crop as a staple. Biofortification—breeding cassava varieties whose roots are rich in pro-vitamin A carotenoids—has long been championed as a public health solution. Now, a team of researchers working with Kenyan and Ugandan germplasm reports a substantial leap forward in the computational machinery of that effort, showing that a smarter combination of genome-wide association studies and machine learning can predict the nutritional quality of cassava roots with unprecedented accuracy.</p>
<p>The core obstacle is time. Cassava takes roughly twelve months to mature, which means breeders who wait to measure carotenoid content in harvested roots must commit years to each selection cycle. Genomic prediction offers a way around this bottleneck. The idea is to build statistical models that link tens of thousands of DNA markers to measured traits, then use those models to calculate genomic estimated breeding values—predictions of an individual&#8217;s additive genetic worth as a parent. Breeders can then select the most promising progenitors at the seedling stage, months before a single root is dug from the ground, compressing the breeding cycle and accelerating genetic gain.</p>
<p>In a study conducted at the Kenya Agricultural and Livestock Research Organization (KALRO) in Kakamega, researchers assembled a panel of 93 pro-vitamin A cassava genotypes derived from accessions held by Uganda&#8217;s National Crop Resource Research Institute. The plants were established as seedlings in January 2023, and in January 2024 two representative roots from each genotype were harvested under low-light evening conditions—a necessary precaution, because carotenoids degrade rapidly when exposed to light. Samples were rushed to the University of Nairobi within seven hours, where high-performance liquid chromatography quantified beta-carotene content in freshly harvested tissue. The measured values ranged from 0.04 to 1.03 mg per 100 g of root, with a mean of 0.52 mg/100 g. Notably, 68 percent of the population fell below the Institute of Medicine&#8217;s recommended benchmark of 6 micrograms of beta-carotene per gram, underscoring how much work remains for breeders.</p>
<p>The genotyping side of the study relied on DArTseq genotyping-by-sequencing against the Manihot esculenta 671_v8.0 reference genome, yielding 54,877 single nucleotide polymorphisms. After quality control filtering for call rate, minor allele frequency and heterozygosity using the snpReady package in R, 37,419 high-quality SNPs remained for downstream analysis. Five-fold cross-validation, in which 80 percent of the data trained each model and the remaining 20 percent served as an independent test, was used to evaluate prediction performance across all models.</p>
<p>The researchers&#8217; central innovation lay in how they identified trait-associated markers to feed into their prediction models. Conventional single-locus GWAS approaches, while widely used, suffer from limited statistical power in small populations—a persistent constraint in crop breeding programs. The team instead turned to multi-locus random marker effect GWAS (mr-GWAS), implemented through four algorithms in the mrMLM platform: mrMLM, FASTmrMLM, pLARmEB and ISIS EM-BLASSO. These methods treat SNP effects as random, shrinking estimates toward zero and applying a modified Bonferroni significance threshold of 0.05 divided by the effective number of markers, which corrects for multiple testing without the punishing conservatism of standard corrections. The payoff was clear: while the BLINK fixed-effect GWAS model detected only two significant SNPs for beta-carotene (on chromosomes 09 and 14) and one for flesh color on chromosome 01, the mr-GWAS pipeline uncovered five significant SNPs distributed across chromosomes 01, 03, 04, 14 and 18.</p>
<p>When these markers were incorporated as fixed-effect covariates into classical Bayesian genomic prediction models—Bayesian Ridge Regression, Bayesian Lasso, BayesA, BayesB and BayesC—the improvement was dramatic. Naive models struggled badly: prediction correlations on the validation set ranged from just 0.14 to 0.16 for beta-carotene. Adding the two BLINK SNPs lifted accuracy to between 0.41 and 0.43. But when all five mr-GWAS-derived SNPs were included, Bayesian Lasso reached a prediction ability of r = 0.81, with the remaining models clustered between 0.75 and 0.78. The pattern held for root flesh color as well, where a single chromosome 01 SNP boosted validation correlations from roughly 0.39–0.48 to 0.61–0.63. The lesson is that markers spanning more of the genome are more likely to be linked to the causal genes governing carotenoid accumulation, capturing a fuller share of the trait&#8217;s genetic architecture.</p>
<p>The study did not stop with parametric models, however. Traditional genomic prediction assumes purely additive gene action, yet real traits are shaped by dominance and epistasis—non-additive interactions that linear models systematically miss. To capture these effects, the researchers built non-parametric models using five machine learning algorithms: Random Forest, Support Vector Machine with radial kernel, Neural Networks, Extreme Gradient Boosting (XGBoost) and K-Nearest Neighbors, along with a stacked ensemble that used Random Forest as the meta-learner. Initially, these models performed modestly, with the ensemble topping out at r = 0.44 and the support vector machine limping along at r = 0.04.</p>
<p>The transformation came through feature selection. Using the Boruta algorithm, the team distilled 37,419 SNPs down to just 13 of the most informative predictors—a 99.97 percent reduction in dimensionality. With this streamlined marker set, performance soared: Random Forest and the ensemble model both achieved r = 0.79, while XGBoost and K-Nearest Neighbors reached r = 0.73. This result carries a practical message for computational breeders everywhere—curating the predictor set matters as much as choosing the algorithm, since removing noise variables sharpens model focus, reduces computational burden and prevents overfitting.</p>
<p>The distinction between the two modeling families maps directly onto breeding strategy. Genomic estimated breeding values from the parametric models reflect additive effects, the only genetic component parents reliably transmit to offspring, making them the right tool for early progenitor selection. Non-parametric machine learning models, by contrast, estimate total genomic value—additive and non-additive combined—which better predicts how a clone will actually perform as a finished variety. Together, the two approaches let breeders pick parents early and pick varieties accurately, shortening every stage of the pipeline from crossing to release.</p>
<p>The authors acknowledge limitations, including the modest population size of 93 genotypes and the fact that GWAS was performed on the full dataset before cross-validation, which may have introduced some optimistic bias into reported accuracies. They recommend larger datasets and nested cross-validation designs in future work. Even so, the results represent a meaningful advance for a crop that has long lagged behind maize and wheat in genomics-assisted breeding. For the millions of families whose primary food source falls short on vitamin A, faster breeding is not an abstraction—it is the difference between deficiency and health arriving sooner rather than later. With genomic prediction accuracies approaching r = 0.8 now achievable even in small breeding populations, cassava biofortification may finally be able to move at the speed the problem demands.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Genomic prediction of pro-vitamin A carotenoid (beta-carotene) content in cassava roots using GWAS-informed parametric Bayesian models and machine learning ensemble approaches</p>
<p><strong>Article Title:</strong> Multi-locus random marker effects GWAS, ensemble models and variable selection improve genomic prediction for pro-vitamin A carotenoids in cassava</p>
<p><strong>Article References:</strong> Abincha, W., Dzidzienyo, D. K., Tongoona, P., Ofori, K., Owor, B.-E., Kayondo, I. S., Ozimati, A., Mwale, S. E., &amp; Kivuva, B. M. (2026). Multi-locus random marker effects GWAS, ensemble models and variable selection improve genomic prediction for pro-vitamin A carotenoids in cassava. <em>Heliyon, 12</em>(14), Article e45360. <a href="https://doi.org/10.1016/j.heliyon.2026.e45360" target="_blank" rel="noopener noreferrer">https://doi.org/10.1016/j.heliyon.2026.e45360</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.heliyon.2026.e45360" target="_blank" rel="noopener noreferrer">10.1016/j.heliyon.2026.e45360</a></p>
<p><strong>Keywords:</strong> cassava biofortification, genomic prediction, pro-vitamin A carotenoids, beta-carotene, multi-locus GWAS, machine learning, Random Forest, ensemble models, variable selection, vitamin A deficiency, SNP markers, plant breeding</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">185362</post-id>	</item>
		<item>
		<title>New Microhaplotype Databases and Tools Reveal Crop Genetic Diversity</title>
		<link>https://scienmag.com/new-microhaplotype-databases-and-tools-reveal-crop-genetic-diversity/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Wed, 26 Aug 2026 20:12:26 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[crop genetic diversity]]></category>
		<category><![CDATA[crop genome analysis tools]]></category>
		<category><![CDATA[crop genotyping methods]]></category>
		<category><![CDATA[DNA variation in crops]]></category>
		<category><![CDATA[duplicated crop genomes]]></category>
		<category><![CDATA[genetic mapping in agriculture]]></category>
		<category><![CDATA[high heterozygosity in crops]]></category>
		<category><![CDATA[microhaplotypes]]></category>
		<category><![CDATA[multi-variant DNA signatures]]></category>
		<category><![CDATA[no-code genetic analysis software]]></category>
		<category><![CDATA[plant breeding genetic markers]]></category>
		<category><![CDATA[plant population genetics databases]]></category>
		<guid isPermaLink="false">https://scienmag.com/new-microhaplotype-databases-and-tools-reveal-crop-genetic-diversity/</guid>

					<description><![CDATA[Plant breeders have gained a new way to see genetic variation that standard DNA markers can miss. An international research team has assembled standardized databases of “microhaplotypes”—short stretches of DNA containing several tightly linked variants—for eight important crops. The resources cover alfalfa, blueberry, cranberry, cucumber, pecan, potato, strawberry and sweetpotato, bringing together more than 60,000 [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Plant breeders have gained a new way to see genetic variation that standard DNA markers can miss. An international research team has assembled standardized databases of “microhaplotypes”—short stretches of DNA containing several tightly linked variants—for eight important crops. The resources cover alfalfa, blueberry, cranberry, cucumber, pecan, potato, strawberry and sweetpotato, bringing together more than 60,000 samples and tens of thousands of distinct sequence variants. The researchers say the framework can improve population analysis, genetic mapping and breeding decisions, particularly in crops with duplicated genomes, high heterozygosity or multiple chromosome sets. Their study, published in Theoretical and Applied Genetics, also introduces HapApp, a no-code software tool designed to help breeders add newly discovered variants to the growing databases without writing scripts.</p>
<p>Most modern crop-genotyping systems rely on single-nucleotide polymorphisms, or SNPs. A SNP records a single DNA-letter difference and usually has two possible states, making it a relatively simple biallelic marker. SNPs are abundant, inexpensive to measure and easy to analyze, but that simplicity can become a weakness in complex crops. A short genomic region may contain several nearby variants that are inherited together. Treating each position as an isolated, two-option marker throws away the information contained in the precise combination of variants. A microhaplotype preserves that combination as a single, multiallelic signature. Because the variants lie within a short sequence, they are read together and are generally assumed to remain linked, allowing researchers to determine which changes occur on the same chromosome copy. The result can be a marker with three, ten or even dozens of distinguishable allelic forms rather than only two.</p>
<p>The team built its resources from DArTag targeted-genotyping data. Unlike random reduced-representation methods such as genotyping-by-sequencing, which sample a different subset of the genome from one project to another, DArTag uses designed primers to amplify the same target regions repeatedly. Sequencing reads typically span 50 to 250 base pairs and include the chosen variant as well as neighboring polymorphisms. The researchers focused on the sequence portions of about 54 or 81 base pairs, depending on the panel design. Each DArTag report contains read counts for several sequence classes, including known reference and alternative alleles and newly observed “RefMatch” or “AltMatch” sequences that resemble those alleles but contain additional changes in the flanking DNA. These unanticipated combinations are precisely where microhaplotypes can reveal diversity that a conventional SNP call would overlook.</p>
<p>A central problem was that the same sequence could appear under different identifiers in separate projects. Without a stable name, a breeder could not reliably compare a marker analyzed in one laboratory with the apparently similar marker generated elsewhere. The researchers therefore created a species-agnostic pipeline that maps markers to chromosome or scaffold coordinates and assigns fixed identifiers such as “Chr01_000589632.” The workflow aligns amplicon sequences to 180-to-300-base-pair regions of a reference genome using BLAST, checks strand orientation and removes ambiguity in the original reports. A sequence is linked to an existing database identity only when it matches the full reference entry at 100 percent identity and coverage. A sequence that fails to match exactly but reaches at least 90 percent identity across at least 90 percent of the comparable region is treated as a candidate new variant. Sequences below that threshold are discarded as likely technical artifacts or off-target products.</p>
<p>The database-building effort also preserves a distinction between true alleles and sequences generated from duplicated genomic regions. In polyploid crops, closely related chromosome copies can be difficult to separate, and primers may amplify paralogous loci—related but nonallelic regions—alongside the intended target. The team retained these signals but annotated them rather than silently deleting them, because the same sequences may recur in future experiments. That choice creates a comprehensive catalog while allowing downstream analyses to exclude questionable markers. In a biparental population, for example, biological segregation limits the number of genuine alleles expected at a locus. Linkage software can also identify markers with abnormal inheritance or poor recombination behavior. This layered strategy lets the global database remain inclusive while giving breeders tools to isolate orthologous, biologically interpretable variation.</p>
<p>The eight crop databases vary dramatically in size and genetic richness. The alfalfa resource contains 35,259 microhaplotypes from 3,000 loci and more than 15,600 samples gathered across 25 breeding projects. Its average marker carries about 12 alleles, and the database has approached a plateau near 35,000 sequences, suggesting that the current panel has captured much of the diversity present in the sampled US breeding populations. The blueberry database contains 28,653 sequences from roughly 8,930 samples and 3,000 loci, with an average of about ten alleles per locus. Potato contributes 35,506 sequences from 3,913 loci and more than 3,100 samples, while strawberry contains 38,227 sequences from 5,000 loci and 1,880 samples. The strawberry panel targets all 28 chromosomes of its octoploid genome, including its four subgenomes, and half of its loci contain six or fewer alleles.</p>
<p>The remaining databases illustrate how sampling and genome biology shape the apparent amount of diversity. The cranberry resource contains 14,380 alleles from 3,050 loci and 4,146 samples. Cucumber has 9,823 microhaplotypes from 3,059 loci and 8,272 samples, but a mean of only about three alleles per marker; the authors caution that this may reflect the narrow breeding material sampled rather than a universally low level of cucumber diversity. Pecan contains 26,073 alleles from 3,100 loci and 6,768 samples. Sweetpotato, a globally distributed hexaploid crop with six copies of each chromosome set, contains 37,593 sequences from 3,120 loci and 9,212 samples. Its data span breeding programs in North and South America, Africa, Asia and the Caribbean, allowing the database to capture geographically structured variation rather than diversity from a single breeding population.</p>
<p>Two case studies tested whether microhaplotypes actually improve genetic analyses rather than simply generating larger databases. In pecan, the researchers constructed linkage maps from 188 offspring using three data types: microhaplotypes, only the target SNPs, and all SNPs detected within the amplified regions. The microhaplotype data retained more markers that were informative about recombination in both parents. By contrast, the SNP datasets were dominated by markers informative in only one parent, making it difficult to connect the two parental maps. When the researchers ordered markers using multidimensional scaling of genetic distances, the SNP-based maps placed some markers far from their expected physical positions. Microhaplotypes produced fewer gaps and more stable ordering because the linked variants supplied additional allele combinations and made parental phase easier to determine directly from sequencing reads.</p>
<p>The second test used 4,087 matched sweetpotato samples from seven international projects. After quality filtering, the researchers compared 27,220 microhaplotypes at 2,772 loci with 2,534 target SNPs and 20,935 SNPs extracted from the same regions. Principal-component analysis produced tighter, more clearly separated clusters with microhaplotypes, especially for the Taiwanese population. The first principal component explained 15.72 percent of the variation with microhaplotypes, compared with 10.5 percent for target SNPs and 5 percent for all SNPs. Discriminant analysis of principal components showed that the first linear discriminant axis captured 86.3 percent of between-group variance with microhaplotypes, versus about 60 percent with either SNP dataset. All three approaches ultimately achieved mean classification accuracy above 98 percent, but microhaplotypes remained near a 99 percent success plateau as more principal components were included, suggesting greater stability in a high-dimensional analysis.</p>
<p>The researchers emphasize that the databases are not a universal census of crop diversity. Marker panels were designed mainly in genic regions, so they preferentially sample potentially functional portions of the genome rather than random DNA. Database sizes are also influenced by panel length, genome complexity, ploidy, sample number and the breeding populations that collaborators were able to share. Rare-variant counts are especially sensitive to sample size: in sweetpotato, similarly sized US and Taiwanese cohorts contained vastly different numbers of private microhaplotypes, likely reflecting differences in breeding history and gene-pool structure rather than sampling alone. Even so, standardized records could help genebanks detect duplicated accessions, uncover mislabeled material, identify gaps in collections and assemble representative core sets for pre-breeding. The sequences and scripts are being released under FAIR data principles, with the databases available through Zenodo and the software through GitHub.</p>
<p>HapApp translates the computational workflow into a point-and-click interface. Users upload a DArTag report, select the crop and panel characteristics, and receive a filtered report in which every accepted microhaplotype has a stable identity, along with an updated FASTA sequence file and, when necessary, a new database version. The authors are also developing HapSearch, a planned platform for finding alleles by crop, locus, sequence similarity or project keyword and for identifying germplasm associated with unusual variants. At present, the system is built around DArTag’s proprietary MADC reports, so other targeted platforms such as GT-seq, AgriSeq and FlexSeq would require format-conversion pipelines. Still, the researchers argue that the underlying principle is portable: preserve linked sequence information, give recurring variants durable names and use multiallelic data to make breeding genomes easier to read. In crops facing climate stress, disease and changing production demands, that extra resolution could help breeders find useful genetic variation before it disappears into the noise of a two-allele marker system.</p>
<p><strong>Subject of Research:</strong> Standardized microhaplotype databases and genetic-diversity analysis for eight crop species</p>
<p><strong>Article Title:</strong> Standardized microhaplotype databases and frameworks for assessing and mining crop genetic diversity</p>
<p><strong>Article References:</strong> Zhao, D., Lin, M., Taniguti, C. H. et al. “Standardized microhaplotype databases and frameworks for assessing and mining crop genetic diversity.” <em>Theoretical and Applied Genetics</em> 139, 244 (2026). <a href="https://doi.org/10.1007/s00122-026-05340-4">Original research article</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> 10.1007/s00122-026-05340-4</p>
<p><strong>Keywords:</strong> microhaplotypes, crop genetics, plant breeding, genetic diversity, polyploid crops, DArTag genotyping, linkage mapping, population structure, HapApp</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">182475</post-id>	</item>
	</channel>
</rss>
