<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>statistical rigor in genomic models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/statistical-rigor-in-genomic-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 10:36:38 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>statistical rigor in genomic models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Non-linear kernels boost soybean genomic prediction while keeping the math breeders trust</title>
		<link>https://scienmag.com/non-linear-kernels-boost-soybean-genomic-prediction-while-keeping-the-math-breeders-trust/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 10:36:38 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[advances in plant genomic prediction methods]]></category>
		<category><![CDATA[biofuel crop breeding technology]]></category>
		<category><![CDATA[epistasis]]></category>
		<category><![CDATA[GBLUP]]></category>
		<category><![CDATA[genomic prediction]]></category>
		<category><![CDATA[genomic selection for soybean yield]]></category>
		<category><![CDATA[genotype by environment interaction]]></category>
		<category><![CDATA[heritability]]></category>
		<category><![CDATA[kernel methods]]></category>
		<category><![CDATA[machine learning in plant breeding]]></category>
		<category><![CDATA[multi-environment soybean trait prediction]]></category>
		<category><![CDATA[multi-environment trials]]></category>
		<category><![CDATA[non-linear kernel algorithms in agriculture]]></category>
		<category><![CDATA[non-linear kernel methods]]></category>
		<category><![CDATA[non-linear vs linear models in crop prediction]]></category>
		<category><![CDATA[plant breeding]]></category>
		<category><![CDATA[predictive accuracy in plant breeding]]></category>
		<category><![CDATA[protein and oil content prediction]]></category>
		<category><![CDATA[Random Forest]]></category>
		<category><![CDATA[soybean]]></category>
		<category><![CDATA[soybean genomic prediction]]></category>
		<category><![CDATA[SoyNAM population]]></category>
		<category><![CDATA[statistical rigor in genomic models]]></category>
		<category><![CDATA[variance components]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227199</guid>

					<description><![CDATA[A new study shows that non-linear kernel methods consistently match or outperform the standard linear model for predicting soybean yield, protein, and oil across multiple environments while preserving the statistical framework breeders rely on.]]></description>
										<content:encoded><![CDATA[<p>Soybean is one of the most valuable crops on Earth, feeding people, livestock, and the growing biofuel industry with its protein- and oil-rich seeds. Yet predicting which breeding line will perform best, and where, remains one of the hardest problems in plant science. A new study published in Theoretical and Applied Genetics shows that a family of mathematical tools borrowed from machine learning can sharpen those predictions without sacrificing the statistical rigor that breeders depend on. The work, led by Wanessa Alves Lima Paiva of the Federal University of Viçosa in Brazil together with colleagues at the University of Florida, demonstrates that non-linear kernel methods consistently match or beat the industry-standard linear model for forecasting soybean grain yield, protein content, and oil content across multiple environments.</p>
<p>For more than a decade, the workhorse of genomic prediction has been a method called Genomic Best Linear Unbiased Prediction, or GBLUP. The approach treats thousands of DNA markers as a single genomic relationship matrix, essentially a map of how genetically similar each breeding line is to every other line, and uses that similarity to estimate the breeding value of untested plants. It is elegant, fast, and interpretable, but it has a fundamental limitation: it is linear. Complex traits such as yield are not governed by genes acting independently. Multiple quantitative trait loci interact with one another through epistatic effects, and the same genotype can excel in one field and disappoint in another, a phenomenon known as genotype-by-environment interaction. When the underlying biology bends and curves in these ways, a straight-line model can leave predictive signal on the table.</p>
<p>The research team tackled this problem by replacing the linear relationship matrix in the mixed-model framework with four alternative kernels: Laplacian, Gaussian, Bessel, and Polynomial. Each kernel is a mathematical function that converts the raw marker profiles of two genotypes into a single similarity score, but they differ in how that similarity decays with genetic distance. The Gaussian kernel, for example, uses a squared Euclidean distance and produces a smooth, bell-shaped decay of similarity, while the Laplacian kernel relies on a Manhattan distance and decays more sharply. Through the so-called kernel trick, these functions implicitly project the marker data into a higher-dimensional space where non-linear relationships between markers, including epistatic interactions of multiple orders, can be captured without ever explicitly estimating every pairwise interaction. Crucially, because the kernels simply replace the covariance structure inside a Bayesian mixed model, the researchers could still estimate variance components and heritability, the quantities that make quantitative genetics useful for breeding decisions.</p>
<p>The evidence base was substantial. The team analyzed the SoyNAM population, a publicly available resource of roughly 5,400 recombinant inbred lines derived from 40 biparental families that each share a common elite parent. From this resource they used a subset of 1,379 genotypes evaluated across six environments in Iowa, Illinois, Indiana, and Nebraska during 2011 and 2012, yielding 8,274 phenotypic records per trait. After filtering, 4,325 single nucleotide polymorphisms remained as predictors. Because successive generations of selfing make these lines nearly homozygous, dominance effects are effectively eliminated, leaving additive and additive-by-additive epistatic effects as the main sources of genetic variation, an ideal setting for testing whether non-linear kernels can detect interaction signal that linear models miss.</p>
<p>Testing was carried out under three cross-validation schemes that mirror real breeding decisions. CV1 asks whether a model can predict genotypes never observed in environments where training data exist, the classic scenario when new lines are developed. CV2 addresses a sparser situation in which genotypes have been tested in only some locations. CV0 is the hardest test of all: predicting how genotypes perform in environments entirely absent from the training set. Predictive accuracy was quantified as the Pearson correlation between observed and predicted values, with results averaged across five independent data splits for the more variable schemes.</p>
<p>The results revealed a striking trait-specific pattern. Grain yield proved the most stubborn target. Under a main-effects-only model, the genetic component explained on average just 17.9 percent of phenotypic variance for yield, with error dominating at 82.1 percent. When genotype-by-environment interaction terms were added, the picture flipped dramatically: the interaction component became the largest contributor at 55.6 percent, and predictive accuracy under CV1 rose from a range of roughly 0.20 to 0.55 up to approximately 0.35 to 0.70. Seed composition traits told a different story. Protein content averaged 32.6 to 34.8 percent across environments and oil content 18.8 to 20.4 percent, and both traits showed much stronger genetic control, with heritability estimates around 0.45 for protein and 0.55 for oil. For these more stable traits, adding interaction terms brought only modest gains, and oil content was where the non-linear kernels showed their clearest advantage over GBLUP, outperforming it across all validation schemes.</p>
<p>The hardest test, CV0, produced one of the study&#8217;s most instructive findings. When the target environment was completely unobserved, adding explicit interaction terms provided little benefit, because interaction effects for that environment cannot be learned without phenotypic observations there. Predictions instead leaned on the main genetic effects estimated from other sites. Even so, the non-linear kernels generally outperformed GBLUP under both modeling frameworks in this scenario, suggesting that their richer similarity structure helps transfer information to novel environments. Certain environments proved consistently difficult regardless of method: Illinois 2012 was a persistent weak spot for yield prediction, while Illinois 2011 and Indiana 2012 repeatedly dragged down accuracy for protein and oil, reinforcing the principle that training populations must represent the range of target environments.</p>
<p>Hyperparameter tuning emerged as a decisive but tractable part of the workflow. For the Gaussian kernel, accuracy peaked at the smallest bandwidth tested, sigma equal to 0.0001, while the Laplacian kernel was remarkably insensitive to bandwidth between 0.0001 and 0.01, making it easier to optimize in practice. The Bessel kernel was highly sensitive to its scale parameter, with sigma of 0.1 producing a marked collapse in performance, whereas its degree and order parameters mattered little. The Polynomial kernel showed an almost flat optimization landscape, with degree 2 sufficient to capture the underlying structure. The authors note that these small optimal bandwidths indicate performance was maximized when similarity was concentrated among closely related genotypes, emphasizing local rather than global genetic relationships.</p>
<p>To stress-test the approach under known epistasis, the team turned to a simulated dataset of 1,000 individuals with 4,010 markers, in which six traits were generated under epistatic architectures ranging from 8 to 480 quantitative trait loci and heritabilities from 0.70 down to 0.30. Here the hierarchy of methods depended on genetic architecture. For traits controlled by only a handful of loci, Random Forest, a tree-based machine learning method with 500 decision trees, clearly won, because tree splits excel at capturing strong marker effects and threshold-like responses. But as the number of loci grew and the signal became distributed across the genome, the kernels, particularly Laplacian and Gaussian, became competitive or superior, and under the most polygenic settings all methods converged as broad genomic similarity captured most of the predictable signal. Notably, the Laplacian kernel delivered consistently high accuracy across every architecture tested, positioning it as a promising alternative to the widely adopted Gaussian kernel.</p>
<p>Perhaps the study&#8217;s most consequential message is about what breeders gain by staying inside the mixed-model framework. Random Forest was competitive only under the CV0 scheme for protein and oil, and it consistently trailed the other methods for yield. More importantly, machine learning models are primarily predictive black boxes: they do not directly yield variance components, heritability estimates, or the other inferential quantities that guide selection decisions. The non-linear kernel models, by contrast, occupy a practical middle ground, combining the flexibility to capture moderate non-linearities and interaction patterns with the interpretability of quantitative genetics. With no single kernel dominating across all environments, the authors advise that kernel choice should be guided by the characteristics of each dataset rather than by a universal favorite. As breeding programs worldwide race to develop soybeans that yield more under increasingly volatile climates, this work suggests that upgrading the covariance structure, rather than abandoning statistical inference for pure machine learning, may be the smartest path to more accurate and more accountable genomic prediction.</p>
<p><strong>Subject of Research:</strong> Non-linear kernel methods for genomic prediction of soybean yield and seed quality traits in multi-environment trials</p>
<p><strong>Article Title:</strong> Non-linear kernel methods for genomic prediction of soybean yield and quality in multi-environment trials</p>
<p><strong>Article References:</strong> Paiva, W. A. L., da Costa, W. G., Machado, L. P., Nascimento, A. C. C., Jarquin, D., &amp; Nascimento, M. (2026). Non-linear kernel methods for genomic prediction of soybean yield and quality in multi-environment trials. <em>Theoretical and Applied Genetics, 139</em>(10), Article 277. <a href="https://doi.org/10.1007/s00122-026-05387-3" rel="noopener noreferrer">https://doi.org/10.1007/s00122-026-05387-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00122-026-05387-3" rel="noopener noreferrer">10.1007/s00122-026-05387-3</a></p>
<p><strong>Keywords:</strong> genomic prediction, soybean, kernel methods, GBLUP, genotype-by-environment interaction, epistasis, SoyNAM population, Random Forest, variance components, heritability, multi-environment trials, plant breeding</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227199</post-id>	</item>
	</channel>
</rss>
