A team of statisticians and geneticists has unveiled a powerful new computational method that promises to untangle one of the most stubborn puzzles in modern genetics: why a single genetic variant often influences many different traits at once. The technique, called genetic factor analysis, or GFA, is described in a study published in Nature Genetics by Jean Morrison of the University of Michigan, Jason Willwerscheid of Providence College, Dhajanae Sylvertooth, Xin He and Matthew Stephens of the University of Chicago, and colleagues. By mining the enormous troves of summary data produced by genome-wide association studies, the method identifies shared polygenic factors, essentially hidden axes of genetic influence, that run beneath the surface of many seemingly unrelated human traits.
Pleiotropy, the phenomenon in which one gene or genetic variant affects multiple characteristics, has long been both a blessing and a headache for researchers. On one hand, shared genetic associations provide valuable clues about the biological pathways that connect diseases; a variant that raises both cholesterol and coronary artery disease risk, for example, points toward lipid metabolism as a common mechanism. On the other hand, detecting and interpreting these shared patterns is statistically treacherous. Genome-wide association studies, known as GWAS, typically analyze one trait at a time, and the summary statistics they produce are noisy, correlated with one another, and often derived from studies whose participants overlap substantially. Existing multivariate methods struggle with these complications, frequently requiring analysts to guess how many underlying factors exist or to assume, unrealistically, that those factors are statistically independent of one another.
GFA tackles these problems head-on with a Bayesian modeling framework rooted in empirical Bayes matrix factorization, an approach previously developed by Wang and Stephens. Conceptually, the method treats the matrix of genetic effect estimates across many traits as a composite signal that can be decomposed into a small number of pleiotropic factors, each representing a pattern of cross-trait genetic associations, plus trait-specific residual effects. Crucially, the model does not force the factors to be orthogonal, meaning it can capture overlapping biological processes that simpler dimension-reduction techniques such as principal component analysis would blur together. The method also automatically determines how many factors are needed to explain the data, sparing researchers an arbitrary and potentially consequential choice, and it explicitly models the sampling correlations that arise when GWAS datasets share participants.
To demonstrate the power of the approach, the researchers applied GFA to a clinically urgent question: the shared genetic architecture of coronary artery disease and type 2 diabetes, two of the leading causes of death worldwide, along with 22 common risk factors that include lipid measures, blood pressure traits, body size measurements and markers of inflammation. The analysis partitioned the heritability of these traits into 15 distinct pleiotropic components. Each component tells a different story. Some factors capture broad metabolic influences that push risk of both diseases in the same direction, while others reveal more nuanced patterns, such as effects that raise one risk factor while lowering another, or influences that act on diabetes but not on heart disease. This kind of decomposition transforms a wall of association statistics into a structured map of biological processes, allowing researchers to see which combinations of traits travel together genetically and which diverge.
The value of that map extends beyond description. When genetic risk for a disease can be split into interpretable components, each component becomes a hypothesis about mechanism. A factor that loads heavily on inflammation markers and on coronary artery disease, for instance, suggests that inflammatory pathways contribute to cardiovascular risk independently of cholesterol. A factor dominated by adiposity traits with effects on both diseases points to shared metabolic consequences of body fat distribution. Researchers can then prioritize these pathways for experimental follow-up, design better polygenic risk scores that reflect distinct biological axes rather than a single blended score, and identify subtypes of disease that may respond differently to treatment. The approach echoes and extends earlier soft-clustering strategies used to classify type 2 diabetes genetic loci, but does so in a fully probabilistic framework that quantifies uncertainty at every step.
The team then turned to a second, very different application: the composition of blood cells. Dozens of traits describe the relative abundance of different immune cell types in circulating blood, from red cell characteristics to the proportions of various lymphocyte and myeloid populations, and these traits are strongly genetically correlated with one another and with common diseases. Applying GFA to this battery of phenotypes yielded a biologically meaningful decomposition in which individual factors corresponded to coherent groups of related cell types. Rather than treating each blood trait as an isolated variable, the factors capture the underlying developmental and regulatory programs that shape the immune system’s cellular makeup, providing a compressed and interpretable representation of immune genetics.
This blood cell decomposition was not merely an exercise in description. The researchers used the estimated factors as instruments in multivariable Mendelian randomization, a technique that uses genetic variants as natural experiments to probe whether an exposure causally influences a disease outcome. Multivariable Mendelian randomization is notoriously sensitive to weak instruments and to correlations among the exposures being tested, and the tightly correlated blood cell traits make the analysis especially fragile. By replacing the raw traits with GFA factors, the researchers increased the precision of the causal estimates and obtained cleaner, more interpretable results about which immune cell characteristics influence disease risk.
Perhaps the most striking technical lesson from the study concerns a problem that is easy to overlook: sample overlap. Large GWAS consortia frequently recruit from the same population biobanks, so the summary statistics for supposedly different traits may be computed in many of the same individuals. This induces correlations among the estimation errors that, if ignored, can masquerade as genuine shared genetic signal or distort factor estimates in ways that render them biologically meaningless. The authors found that accounting for overlapping samples was critical to obtaining interpretable results in their Mendelian randomization application. GFA builds this correction directly into its model, estimating the error correlations from the data rather than assuming them away, which distinguishes it from many earlier approaches to cross-trait analysis.
The practical accessibility of the method is another notable feature. GFA runs on GWAS summary statistics rather than individual-level genotype data, which means it can be applied to the hundreds of publicly available GWAS results without requiring access to protected participant records. The method is implemented in an open-source R package, and the authors have deposited all the code and data needed to replicate every analysis in the paper through public repositories, an unusually thorough commitment to reproducibility. The work was supported in part by the National Human Genome Research Institute and the National Institute of Allergy and Infectious Diseases.
As biobanks swell to millions of participants and GWAS results accumulate for thousands of phenotypes, methods like GFA address a pressing need: turning scattered single-trait associations into a coherent picture of how genetic variation shapes human biology. By automatically identifying the number of shared factors, allowing those factors to overlap, and correcting for the messy realities of shared study populations, the method offers a rigorous statistical foundation for the emerging field of phenome-wide genetics. If the patterns it uncovers hold up across broader collections of traits and ancestries, the hidden architecture of pleiotropy may soon become far less hidden, and the biological threads connecting heart disease, diabetes, immune traits and beyond will be that much easier to trace.
Subject of Research: A statistical method for identifying shared polygenic factors of genetic pleiotropy across human traits using GWAS summary statistics
Article Title: Genetic factor analysis for characterizing phenome-wide patterns of genetic pleiotropy
Article References: Morrison, J., Willwerscheid, J., Sylvertooth, D., He, X., & Stephens, M. (2026). Genetic factor analysis for characterizing phenome-wide patterns of genetic pleiotropy. Nature Genetics. https://doi.org/10.1038/s41588-026-02753-1
Image Credits: AI Generated
DOI: 10.1038/s41588-026-02753-1
Keywords: genetic pleiotropy, genetic factor analysis, GWAS, polygenic factors, coronary artery disease, type 2 diabetes, Mendelian randomization, blood cell traits, Bayesian statistics, matrix factorization, heritability, Nature Genetics
Cite Scienmag News
Juliet Wilcox. (September 23, 2026). New Genetic Factor Analysis Method Reveals Hidden Shared Roots of Disease. Scienmag. https://scienmag.com/new-genetic-factor-analysis-method-reveals-hidden-shared-roots-of-disease/
Juliet Wilcox. "New Genetic Factor Analysis Method Reveals Hidden Shared Roots of Disease." Scienmag, 23 September 2026, https://scienmag.com/new-genetic-factor-analysis-method-reveals-hidden-shared-roots-of-disease/. Accessed 23 September 2026.
Juliet Wilcox. "New Genetic Factor Analysis Method Reveals Hidden Shared Roots of Disease." Scienmag. September 23, 2026. https://scienmag.com/new-genetic-factor-analysis-method-reveals-hidden-shared-roots-of-disease/

