<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>influence of variant annotation on genetic study outcomes &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/influence-of-variant-annotation-on-genetic-study-outcomes/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 23:44:01 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>influence of variant annotation on genetic study outcomes &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Variant Scoring Tools Disagree Wildly in Gene-Based Genetic Association Studies</title>
		<link>https://scienmag.com/ai-variant-scoring-tools-disagree-wildly-in-gene-based-genetic-association-studies/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 23:44:01 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AlphaMissense]]></category>
		<category><![CDATA[CADD]]></category>
		<category><![CDATA[challenges in rare variant burden testing and disease gene discovery]]></category>
		<category><![CDATA[discrepancies in variant pathogenicity classification]]></category>
		<category><![CDATA[ESM-1b]]></category>
		<category><![CDATA[evaluation of leading genetic variant annotation algorithms]]></category>
		<category><![CDATA[gene-based genetic association analysis using UK Biobank data]]></category>
		<category><![CDATA[gene-level association testing]]></category>
		<category><![CDATA[genetic association study]]></category>
		<category><![CDATA[Genetic variant annotation tools comparison]]></category>
		<category><![CDATA[genomics]]></category>
		<category><![CDATA[GPN-MSA]]></category>
		<category><![CDATA[impact of variant scoring methods on rare variant association studies]]></category>
		<category><![CDATA[implications of tool choice in human genetics research]]></category>
		<category><![CDATA[influence of variant annotation on genetic study outcomes]]></category>
		<category><![CDATA[large-scale genomic analysis of quantitative traits]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in gene variant pathogenicity prediction]]></category>
		<category><![CDATA[optimal transport]]></category>
		<category><![CDATA[rare variants]]></category>
		<category><![CDATA[significance]]></category>
		<category><![CDATA[UK Biobank]]></category>
		<category><![CDATA[variability in DNA variant interpretation across scoring tools]]></category>
		<category><![CDATA[variant annotation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=229643</guid>

					<description><![CDATA[A large-scale UK Biobank study shows that five leading machine learning variant annotation methods produce markedly different calibration, power, and gene discovery outcomes in gene-level association testing.]]></description>
										<content:encoded><![CDATA[<p>A sweeping new analysis of half a million genomes has revealed that the machine learning tools used to judge which DNA variants are dangerous can produce strikingly different scientific discoveries depending on which tool a researcher picks. The study, published in BMC Genomics by a team at Genentech, systematically compared five leading variant annotation methods as the foundation for gene-level association testing across 14 quantitative traits measured in up to 350,377 UK Biobank participants. The findings challenge a quiet assumption embedded in much of modern rare variant genetics: that the choice of pathogenicity scorer is a technical detail rather than a decisive factor shaping what a study finds.</p>
<p>Rare variant association testing is one of the most powerful strategies in human genetics for connecting genes to disease and to measurable traits such as blood pressure, lipid levels, and kidney function. Because individual rare variants are too uncommon to test one at a time with statistical confidence, researchers aggregate the variants within a gene and ask whether the burden of rare, presumably damaging variants differs between people with high and low trait values. The catch is that the word &#8220;presumably&#8221; hides an enormous amount of machinery. With tens of thousands of variants per gene and no experimental evidence for most of them, analysts rely on computational annotation methods to decide which variants to count, which to weight heavily, and which to ignore.</p>
<p>That is where the new study intervenes. The team evaluated annotations from five methods spanning three generations of technology: CADD versions 1.6 and 1.7, a long-standing supervised model trained on evolutionary conservation and genomic features; AlphaMissense, a deep learning system from Google DeepMind that predicts the pathogenicity of every possible missense change in the human proteome; ESM-1b, a protein language model that scores variants based on learned patterns of amino acid sequences; and GPN-MSA, a genomic pretrained model that incorporates multiple sequence alignments across species. These five scorers were plugged into four primary gene-based statistical tests and six annotation-level aggregation tests, creating a large grid of method combinations evaluated against the same trait data.</p>
<p>To compare the combinations fairly, the researchers developed a novel evaluation framework based on optimal transport, a mathematical technique originally developed to compare probability distributions and solve problems in logistics and image processing. Applied here, optimal transport allowed the team to quantify how well each annotation method partitioned genetic signal across labeled sets of variants, producing relative measurements of two distinct properties: calibration, meaning whether the test&#8217;s statistical behavior matches its theoretical expectations, and power, meaning its ability to detect true associations. Separating these two qualities matters because a test can appear successful simply by being miscalibrated in a way that generates more hits, a trap that has complicated earlier benchmarking efforts in the field.</p>
<p>The results were, by the authors&#8217; own framing, markedly divergent. Tests built on CADD labels achieved the highest signal separation, meaning they did the best job of concentrating genuine genetic signal into the variants flagged as damaging. At the opposite end of the spectrum, tests using AlphaMissense labels showed the lowest calibration, indicating that the statistical behavior of gene-based tests fed with AlphaMissense scores departed most from theoretical expectations, even though AlphaMissense is widely regarded as one of the most accurate single-variant pathogenicity classifiers in clinical interpretation contexts. The lesson is uncomfortable but important: a scorer that excels at classifying individual variants as pathogenic or benign does not necessarily excel at the aggregate, gene-level task of boosting association power.</p>
<p>Perhaps the most visually striking result concerned evolutionary constraint. When the researchers examined which genes were flagged as significant by each annotation method, they found that hits from tests using GPN-MSA labels were strongly enriched among genes with high evolutionary constraint, reaching up to a 5.8-fold enrichment, compared with only 2.1- to 2.8-fold enrichment using the other methods. Evolutionary constraint measures how intolerant a gene is to mutation over millions of years of mammalian evolution, and constrained genes are known to be enriched for genuine disease biology. The GPN-MSA result suggests that a scorer grounded in cross-species sequence alignments directs statistical power toward exactly the genes where rare variant signals are most likely to be real.</p>
<p>Crucially, these differences in variant prioritization measurably shifted the patterns of genetic discovery. Different annotation methods did not merely reshuffle the rankings of the same genes; they changed which genes crossed the threshold of significance at all, meaning that a laboratory&#8217;s choice of scorer can determine whether a biologically important gene is reported or missed entirely. Yet the study also found reassuring signs of statistical health across the board: the tests showed low genomic inflation, the standard diagnostic for spurious associations driven by confounding, and similar replication rates when findings were tested in independent data. In other words, all the method combinations were behaving statistically responsibly; they were simply looking for different things.</p>
<p>The authors distill their findings into three conclusions with direct practical consequences for the field. First, no single combination of annotation method and statistical test is currently optimal for gene-level association testing, so there is no defensible default that a study can adopt without justification. Second, results from different methods and tests can and should be aggregated to increase study power, since the methods appear to capture partially overlapping but non-identical slices of the underlying biology, and combining them recovers signal that any one approach loses. Third, integrating the information content of variant annotations and gene-level annotations remains a key open problem, a gap that the authors identify as the central methodological challenge facing rare variant genetics.</p>
<p>The scale of the underlying data lends the conclusions considerable weight. The analysis drew on the UK Biobank under application number 44257 and additionally used data from the National Institutes of Health&#8217;s All of Us Research Program, with the authors acknowledging the hundreds of thousands of participants whose genomes and health measurements made the comparison possible. All authors were employees of Genentech at the time the work was performed, and several hold Roche stock, a disclosure that situates the study within an industrial research environment where rare variant discovery pipelines feed directly into drug target identification. For pharmaceutical companies, a method that systematically misses constrained genes is not an abstract statistical concern; it is a missed therapeutic hypothesis.</p>
<p>For the broader community, the study arrives at a moment when machine learning scorers are proliferating faster than the field can benchmark them. AlphaMissense, ESM-1b, and GPN-MSA represent a new wave of models built on deep learning architectures originally developed for natural language, and each embodies different assumptions about what makes a variant harmful: supervised training on known disease mutations, learned protein sequence grammar, or evolutionary conservation across species. The new results demonstrate that these assumptions translate into genuinely different discovery profiles at scale, not merely different numbers on a leaderboard. Researchers planning gene-based association studies should treat annotation choice as a first-order experimental design decision, ideally running multiple scorers and aggregating results, and method developers should recognize that the gene-level aggregation setting, not just single-variant classification accuracy, is where their tools will ultimately be judged. The full study is openly available under a Creative Commons license, with the DOI 10.1186/s12864-026-13379-2.</p>
<p><strong>Subject of Research:</strong> Comparative performance of machine learning variant annotation methods in gene-level rare variant association testing</p>
<p><strong>Article Title:</strong> Markedly divergent performance of variant annotation methods for gene-level association testing</p>
<p><strong>Article References:</strong> Aguirre, M., Irudayanathan, F. J., Crow, M., Hejase, H. A., Menon, V. K., Pendergrass, R. K., McCarthy, M. I., &amp; Fletez-Brant, K. (2026). Markedly divergent performance of variant annotation methods for gene-level association testing. <em>BMC Genomics</em>. <a href="https://doi.org/10.1186/s12864-026-13379-2" rel="noopener noreferrer">https://doi.org/10.1186/s12864-026-13379-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12864-026-13379-2" rel="noopener noreferrer">10.1186/s12864-026-13379-2</a></p>
<p><strong>Keywords:</strong> variant annotation, gene-level association testing, rare variants, UK Biobank, CADD, AlphaMissense, ESM-1b, GPN-MSA, optimal transport, genetic association study, machine learning, genomics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">229643</post-id>	</item>
	</channel>
</rss>
