Polygenic risk scores, which aggregate the tiny effects of hundreds or thousands of genetic variants into a single estimate of an individual’s predisposition to a trait or disease, have long promised a new era of personalized medicine. Yet that promise has been shadowed by a persistent inequity: the scores perform best in populations of European ancestry because the genome-wide association studies used to train them have overwhelmingly enrolled people of European descent. A new study published in Nature Genetics by Kristin Tsuo, Ying Wang, Alicia R. Martin and colleagues now offers one of the most detailed empirical maps to date of how diversity and scale in biobank data actually translate into better prediction, and the answer turns out to be more nuanced than simply pooling every dataset together.
The research team drew on 245,388 whole-genome sequences from the All of Us Research Program, one of the largest and most ancestrally diverse biomedical data resources ever assembled in the United States, and combined them with data from the UK Biobank, the dominant resource in European-ancestry genomics. From these combined resources, the investigators developed multiancestry polygenic risk scores for 32 traits and diseases, ranging from quantitative blood biomarkers to common chronic conditions. Their central question was deceptively simple: when building a risk score for a given population, is it better to use the largest possible pooled dataset, or to match the training data to the ancestry of the people being predicted?
The answer, the study shows, depends on context. For many traits, the sheer diversity of the All of Us cohort improved prediction accuracy, particularly for participants from populations that have historically been under-represented in genetic research, including people of African, Hispanic and Indigenous American ancestry. This finding matters because the performance gap between ancestry groups has been a major obstacle to equitable clinical deployment of polygenic scores. When the training data include more individuals who share ancestry segments with the target population, the score captures more of the variants and effect sizes relevant to that population, and prediction improves.
But the study also delivered a cautionary result. Maximizing sample size by meta-analyzing the All of Us and UK Biobank data was not universally optimal. For traits with lower polygenicity, meaning traits influenced by a smaller number of variants with larger effects, training exclusively on All of Us data performed best for participants of African ancestry. The explanation lies in ancestry-enriched genetic effects: variants or effect sizes that are specific to, or much more common in, particular ancestry groups. When such effects are averaged together with data from a very different population, the pooled score can dilute the very signals that matter most for the target group. In other words, more data is not always better data if the additional data come from a population whose genetic architecture differs meaningfully from the people being predicted.
To quantify these effects, the researchers examined how individual-level prediction accuracy decayed as a function of ancestry divergence between the discovery genome-wide association study and the target individual. Consistent with theoretical expectations and prior work, accuracy declined roughly linearly with increasing genetic distance from the discovery population. However, this decay was substantially attenuated when the training data themselves were multiancestry. A score trained across multiple ancestry groups retained more of its predictive power as individuals became more distant from any single discovery population, suggesting that multiancestry training acts as a form of insurance against the portability problem that has plagued single-ancestry scores.
Methodologically, the study was rigorous and comprehensive. The team used REGENIE for genome-wide association analyses, METAL for meta-analysis, and the Bayesian shrinkage frameworks PRS-CS and PRS-CSx for score construction, the latter designed specifically for cross-population prediction. They estimated SNP-based heritability and polygenicity with SBayesS and GCTB, quantified heritability with LD Score Regression, and assessed cross-ancestry genetic correlations with Popcorn. This combination allowed them to connect observed differences in score performance to underlying properties of trait architecture, such as heritability, polygenicity and the degree to which effect sizes are shared or divergent across ancestries.
One of the most striking illustrations of ancestry-enriched effects came from blood panel traits. For several hematological measures, multiancestry meta-analysis scores showed improved accuracy at the individual level in African ancestry participants, driven by variants that are far more common in African-ancestry populations than in European-ancestry populations. A classic example in human genetics is the regulatory variant in the Duffy antigen receptor for chemokines gene, which strongly influences neutrophil counts and is nearly fixed in many African-ancestry populations but rare elsewhere. Variants of this kind are invisible to predominantly European discovery studies, yet they can carry substantial predictive weight. Including diverse populations in discovery is therefore not merely a matter of fairness; it uncovers biology that would otherwise be missed entirely.
The findings arrive at a moment when polygenic risk scores are beginning to move into clinical settings, with several chronic disease scores already being evaluated for implementation in diverse US populations. Earlier work, including a widely cited 2019 analysis by Martin and colleagues, warned that clinical use of current scores could exacerbate health disparities because their accuracy is so uneven across ancestry groups. The new study provides a practical roadmap for mitigating that risk. It suggests that institutions building clinical scores should evaluate multiple training strategies per trait, considering the trait’s genetic architecture, the ancestry composition of the intended population, and the availability of ancestrally matched discovery data, rather than defaulting to the largest available meta-analysis.
The study also underscores the strategic value of programs like All of Us, which was designed from the outset to reflect the diversity of the United States, with more than half of its participants coming from under-represented racial and ethnic backgrounds. The results demonstrate concretely that this design choice pays scientific dividends: the diversity of the cohort is not just an ethical feature but a source of predictive power that a homogeneous biobank of equivalent size could not replicate. At the same time, the authors note that individual prediction accuracy still declines with ancestry divergence, and that even the best multiancestry strategies leave gaps for populations that remain thinly sampled, such as Indigenous and Pacific Islander groups. Continued investment in globally representative genomic resources remains essential.
Looking forward, the study’s framework, with its systematic comparison of single-ancestry, meta-analyzed and cross-population score construction across dozens of traits, offers a template for future work. The authors have released their PRS-CS and PRS-CSx weights, analysis code and figure-generation scripts via Zenodo, enabling other researchers to reproduce and extend the findings. As polygenic prediction edges closer to routine clinical use, the message of this research is clear: equitable genetic risk prediction will not emerge automatically from bigger data. It will require deliberate attention to who is included in discovery studies, how their genetic architecture differs across traits, and which training strategy genuinely serves each population best. Diversity, in genomics as in so many domains, is not just a matter of representation but of scientific and clinical performance.
Subject of Research: Multiancestry polygenic risk score development and evaluation using the All of Us Research Program and UK Biobank
Article Title: All of Us diversity and scale yield context-dependent improvements in polygenic prediction
Article References: Tsuo, K., Shi, Z., Ge, T., Mandla, R., Hou, K., Ding, Y., Pasaniuc, B., Wang, Y., & Martin, A. R. (2026). All of Us diversity and scale yield context-dependent improvements in polygenic prediction. Nature Genetics. https://doi.org/10.1038/s41588-026-02734-4
Image Credits: AI Generated
DOI: 10.1038/s41588-026-02734-4
Keywords: polygenic risk scores, All of Us Research Program, UK Biobank, genetic diversity, genome-wide association studies, ancestry-enriched effects, health disparities, precision medicine, multiancestry meta-analysis, genetic architecture, Nature Genetics, biobanks
Cite Scienmag News
Juliet Wilcox. (September 22, 2026). Diverse Biobank Data Sharpen Genetic Risk Prediction Where It Matters Most. Scienmag. https://scienmag.com/diverse-biobank-data-sharpen-genetic-risk-prediction-where-it-matters-most/
Juliet Wilcox. "Diverse Biobank Data Sharpen Genetic Risk Prediction Where It Matters Most." Scienmag, 22 September 2026, https://scienmag.com/diverse-biobank-data-sharpen-genetic-risk-prediction-where-it-matters-most/. Accessed 23 September 2026.
Juliet Wilcox. "Diverse Biobank Data Sharpen Genetic Risk Prediction Where It Matters Most." Scienmag. September 22, 2026. https://scienmag.com/diverse-biobank-data-sharpen-genetic-risk-prediction-where-it-matters-most/








