Human genomes are deeply personal objects, and the laws and consent agreements that govern them increasingly prevent researchers from sharing the raw data on which modern genetics depends. One widely adopted workaround is the artificial genome: a synthetic DNA sequence that reproduces the statistical fingerprints of a real population without belonging to any actual person. Artificial genomes can be used to benchmark methods, test evolutionary hypotheses, and build reference panels for filling in missing genetic information, all while sidestepping data-sharing restrictions. The problem is that generating convincing artificial genomes has proven remarkably difficult. A team of computer scientists and geneticists led by Prateek Anand and Sriram Sankararaman at the University of California, Los Angeles, together with colleagues at the National University of Singapore, Stanford, Harvard Medical School, and Cornell, now reports a new deep generative model that appears to clear the central hurdles at once. Their method, called Genetic Probabilistic Circuits, or GPC, is described in a study published in PLOS Genetics.
The core challenge any model of genetic variation must confront is linkage disequilibrium, the nonrandom association of variants across the genome. Neighboring SNPs, or single nucleotide polymorphisms, tend to be inherited together, but so do many distant ones, and these long-range correlations carry much of the structure that makes a population recognizable in the data. Classical simulators built on the coalescent model capture these patterns by tracing observed variation back to latent genealogies shaped by demographic history, mutation, and recombination. That approach is expressive but computationally punishing, which is why practical tools such as msprime rely on Markovian approximations. A second tradition, descending from the product-of-approximate-conditionals framework of Li and Stephens, dispenses with explicit genealogies and fits the data distribution directly, naturally yielding hidden Markov models. HMMs have been enormously successful in haplotype phasing, genotype imputation, and ancestry inference, but their chain structure forces information between distant SNPs to pass through every intermediate hidden variable, weakening long-range dependencies at each step.
Deep learning promised a way out, and a wave of generative adversarial networks, variational autoencoders, restricted Boltzmann machines, and most recently diffusion models has been applied to genetic data. These models can produce artificial genomes that look plausible in principal component analyses and allele frequency summaries. Yet the UCLA-led team argues they carry structural liabilities for genomics. Generative adversarial networks do not define a probability distribution over the data at all, making likelihood-based inference impossible. Restricted Boltzmann machines define one, but evaluating it requires computing a partition function over an exponentially large space of configurations. Variational autoencoders expose only a lower bound on the marginal likelihood. Diffusion models can in principle compute exact likelihoods but scale poorly to dense SNP data and have not been evaluated on imputation. Crucially, none of these approaches supports efficient estimation of conditional probabilities, the mathematical operation at the heart of genotype imputation, and none offers an objective way to judge whether training has converged.
GPC’s answer is architectural. The model builds on hidden Chow-Liu trees, a class of latent variable models in which every observed SNP is paired with a hidden variable, and the hidden variables are wired together not in a chain but in a tree learned from the data itself. The Chow-Liu algorithm, dating to 1968, computes pairwise mutual information between all variables and connects them in the maximum-weight spanning tree, placing strongly correlated SNPs adjacent to one another regardless of their positions along the chromosome. Where a hidden Markov model would route the dependence between two distant but correlated SNPs through every latent variable in between, losing correlation at each hop, the learned tree carries it over a single edge. In the tree learned from a 10,000-SNP region of chromosome 15 in the 1000 Genomes dataset, the researchers found 4,599 leaf nodes, 397 nodes with degree greater than two, one hub connected to 161 other SNPs, and 18.3 percent of edges spanning more than 1,000 positions apart, a topology fundamentally unlike the unbranched chain of an HMM.
Expressiveness alone would be useless if the model were computationally intractable, and this is where probabilistic circuits enter. A probabilistic circuit represents a joint distribution as a directed acyclic graph of input, sum, and product nodes, with product nodes encoding factorizations and sum nodes encoding weighted mixtures. When the graph satisfies two structural properties, smoothness and decomposability, any marginal probability can be computed exactly in time linear in the circuit’s size. Conditional probabilities, the quantities needed for imputation, follow as simple ratios of two marginal queries, each evaluated in a single feedforward pass. Sampling artificial genomes proceeds by ancestral traversal of the circuit, also linear in size. By compiling hidden Chow-Liu trees into this circuit representation and training them with expectation-maximization on GPUs using the PyJuice package, the team trained models with more than 88 million parameters, with each EM epoch taking under two seconds on the 1000 Genomes data and full training completing in roughly two to six hours on a single NVIDIA RTX A5000 graphics card.
The practical payoff shows up most clearly in genotype imputation, the task of inferring variants a person was not directly genotyped, which underlies much of genome-wide association science. The researchers evaluated GPC across three settings using the 1000 Genomes Project, the UK Biobank, and a high-coverage 1000 Genomes release. In the general setting, where models trained on diverse ancestries impute into similarly diverse test genomes, GPC used for direct conditional imputation achieved a 27.5 percent average improvement in the squared correlation between imputed and true genotypes over the next best method, a restricted Boltzmann machine, and a striking 174 percent improvement for low-frequency variants with minor allele frequency below one percent. Direct imputation through the model itself also beat GPC’s own artificial genomes used as reference panels for the standard tool Impute5, by about 15.5 percent overall, apparently because it targets the imputation objective directly rather than injecting the noise of an intermediate simulation step.
The gains were largest precisely where existing infrastructure is weakest: populations underrepresented in public reference panels. Because large public datasets are overwhelmingly of European ancestry, imputation into African, admixed, or other non-European populations suffers from distributional mismatch. In population-specific experiments, GPC’s direct imputation improved on the next best deep generative method by 33 percent on average, and by 279 percent for low-frequency variants. Remarkably, in the 1000 Genomes experiments GPC trained on ancestry-matched private data outperformed Impute5 running on real European reference genomes, achieving a 12.3 percent average improvement in squared correlation and a 42.1 percent improvement for rare variants. In a realistic array-based scenario, imputing 12,551 SNPs missing from the HumanOmni5Exome genotyping array, GPC outperformed the European-panel Impute5 benchmark by 96.5 percent on average and by more than twelvefold for low-frequency variants. The entire held-out array imputation took about one second in a single conditional query.
Privacy, the original motivation for artificial genomes, also fared better under GPC. Using the nearest-neighbor adversarial accuracy metric, which scores values near 0.5 as the ideal balance between utility and privacy, GPC came closest to the ideal across nearly all splits of both datasets. The failure modes of the competitors were instructive. Restricted Boltzmann machines produced synthetic haplotypes that each sat closer to some individual real genome than to any other synthetic one, meaning samples clustered around training individuals, a memorization pattern that could allow synthetic genomes to be traced back to specific people. The Wasserstein generative adversarial network showed the opposite pathology, with artificial genomes occupying a region largely disjoint from the real data, sacrificing utility. The hidden Markov model failed in the most extreme form, producing genomes perfectly separable from real ones in both directions, achieving maximal privacy by failing to model the data at all. GPC’s artificial genomes were not perfectly indistinguishable, and the authors caution that improved performance on this metric does not guarantee protection against membership inference or attribute inference attacks.
Limitations remain, and the authors are candid about them. GPC currently operates on single genomic regions of roughly 10,000 to 15,000 SNPs, bounded by the memory cost of constructing the Chow-Liu tree, which grows with the square of the SNP count; reaching genome-wide scale will require hierarchical or distributed approaches. The present implementation handles haploid rather than diploid genomes, and like all generative models GPC inherits biases from its training data, leaving fairness across diverse cohorts an open problem. Still, the study marks a notable convergence of two research traditions that have mostly run in parallel: the tractable probabilistic models beloved of population geneticists and the expressive deep architectures driving modern machine learning. By learning the dependence structure of the genome directly from mutual information and then compiling it into a form where exact inference is cheap, GPC suggests that the trade-off between realism, computational tractability, and privacy in synthetic genomic data may be less inevitable than it once appeared. Code and experiments are publicly available as the field moves toward equitable genomic tools for populations long left out of reference panels.
Subject of Research: A tractable deep generative model, Genetic Probabilistic Circuits, for modeling human genetic variation data
Article Title: GPC: An expressive and tractable deep generative model for genetic variation data
Article References: Anand, P., Liu, A., Dang, M., Fu, B., Wei, X., Van den Broeck, G., & Sankararaman, S. (2026). GPC: An expressive and tractable deep generative model for genetic variation data. PLOS Genetics, 22(10), e1012321. https://doi.org/10.1371/journal.pgen.1012321
Image Credits: AI Generated
DOI: 10.1371/journal.pgen.1012321
Keywords: genetic probabilistic circuits, artificial genomes, genotype imputation, linkage disequilibrium, probabilistic circuits, hidden Chow-Liu trees, population genetics, deep generative models, genomic privacy, 1000 Genomes Project, UK Biobank, underrepresented populations
Cite Scienmag News
Juliet Wilcox. (October 8, 2026). New AI Model Learns the Hidden Architecture of Human Genetic Variation. Scienmag. https://scienmag.com/new-ai-model-learns-the-hidden-architecture-of-human-genetic-variation/
Juliet Wilcox. "New AI Model Learns the Hidden Architecture of Human Genetic Variation." Scienmag, 8 October 2026, https://scienmag.com/new-ai-model-learns-the-hidden-architecture-of-human-genetic-variation/. Accessed 8 October 2026.
Juliet Wilcox. "New AI Model Learns the Hidden Architecture of Human Genetic Variation." Scienmag. October 8, 2026. https://scienmag.com/new-ai-model-learns-the-hidden-architecture-of-human-genetic-variation/

