A team of forensic geneticists in China has shown that a panel of roughly 2,000 single nucleotide polymorphisms, or SNPs, paired with machine learning algorithms can reliably identify family relationships out to third-degree relatives and distinguish fine-scale genetic subpopulations across East Asia. The study, published in the International Journal of Legal Medicine, was led by Xiaolian Wu and colleagues at the Guangzhou Key Laboratory of Forensic Multi-Omics for Precision Identification at Southern Medical University, working with Bofeng Zhu of Southern Medical University and Shanxi Medical University. Their findings suggest that moderate-density SNP panels, which sit between the small marker sets used in traditional forensic testing and the massive arrays used in genomic research, may offer an efficient sweet spot for both kinship analysis and ancestry inference in casework.
Forensic kinship testing has long relied on short tandem repeats, the repetitive DNA sequences that underpin standard DNA profiling. STRs are excellent for identifying individuals and close relatives such as parents and children, but their statistical power fades rapidly for more distant relationships. As forensic geneticists increasingly turn to sequencing technologies and dense SNP data, a central question has emerged: how many SNPs are actually needed to resolve a given degree of relatedness, and can computational methods extract more information from a moderate number of markers than classical likelihood approaches alone? The new study addresses that question directly, testing a panel of nearly 2,000 SNPs that had previously shown strong forensic value in Chinese Han populations but had not yet been validated in other groups.
The researchers focused their kinship analysis on the Chinese Yunnan Zhuang group, a Tai-Kadai-speaking ethnic minority from Yunnan province in southwestern China. Their goal was to determine whether the moderate-density SNP panel could distinguish first-degree relatives, which include parent-child and full sibling pairs, second-degree relatives such as grandparents, grandchildren, half siblings, and avuncular pairs, third-degree relatives such as first cousins, and unrelated individuals. To do this, they computed two complementary measures for every pair of individuals: the logarithm of the likelihood ratio, a classical forensic statistic that compares the probability of the observed genotypes under a kinship hypothesis versus an unrelated hypothesis, and the cumulative identity-by-state score, which tallies how many alleles two people share at each marker simply by counting matching states.
The results were striking for close relatives. The density curves of the Log10 likelihood ratio and the cumulative identity-by-state values for first- and second-degree kinships were completely separated from those of unrelated individuals, meaning that with these markers, no close relative would be misclassified as unrelated or vice versa. Third-degree kinship proved harder, as expected, because first cousins share on average only about 12.5 percent of their DNA and the overlap with unrelated pairs becomes substantial. When the team set Log10 likelihood ratio thresholds at minus 4 and 4 to separate third-degree relatives from unrelated individuals, the system achieved a power of 86.40 percent with an error rate of zero. In practical terms, the panel can confidently exclude unrelated pairs and correctly flag a large majority of cousin-level relationships without producing false kinship calls, a critical property in forensic and disaster victim identification contexts where a false positive can have serious consequences.
Classical likelihood statistics, however, were only half of the story. The team also trained machine learning classifiers to sort pairs of individuals into four categories: first-, second-, third-degree relatives, and unrelated pairs. Using the Log10 likelihood ratio and cumulative identity-by-state values as input features, they compared several algorithms and found that the K-Nearest Neighbor model performed best, achieving an F1 score of 0.9781 across all four kinship classes. The F1 score, which balances precision and recall on a scale from zero to one, indicates that the model made very few false assignments across the full spectrum of relationship types. This result demonstrates that even a simple, interpretable machine learning method can capture patterns in SNP sharing statistics that fixed thresholds miss, effectively learning the boundary regions where likelihood ratios for different relationship classes overlap.
Beyond kinship, the study tackled a second forensic challenge: biogeographic ancestry inference at fine scale. The researchers integrated genetic data from 135 populations spanning nine geographic regions, drawn from public datasets, and conducted comprehensive population genetic analyses. The Yunnan Zhuang group showed the closest genetic affinity with geographically adjacent populations, particularly the Dai, Miao, and Tujia minorities of southern China and the Kinh population of Vietnam in Southeast Asia. This pattern is consistent with the known population history of the region, in which Tai-Kadai, Hmong-Mien, and related language groups share deep ancestral connections across southern China and mainland Southeast Asia.
Principal component analysis, phylogenetic tree construction, and ADMIXTURE-based ancestry modeling all revealed significant genetic structure among East Asian subpopulations, with especially clear separation between minorities from southern and northern China. Principal component analysis projects individuals into a low-dimensional space defined by the axes of greatest genetic variation, allowing clusters corresponding to population groups to emerge visually. Phylogenetic trees summarize the branching relationships among populations based on genetic distance, while ADMIXTURE models each individual’s genome as a mixture of ancestral components. The fact that all three approaches, applied to a panel of only about 2,000 SNPs, recovered this structure suggests the marker set captures enough ancestry-informative variation to serve as a practical tool for subpopulation discrimination, not just a research-grade dataset.
To push the discrimination further, the team applied multinomial LASSO regression, a feature-selection method that shrinks the coefficients of less informative markers to zero, effectively screening the full panel down to the most ancestry-informative subset. This process yielded 688 SNPs capable of distinguishing five East Asian subpopulations. Using these selected markers, the researchers built two classification models based on different strategies. A partial least squares-discriminant analysis model organized around a hierarchical classification approach, which first splits populations into broad groups and then subdivides them, and an Elastic Net model using a flat multi-classification approach, which assigns each sample directly to one of the five groups, both achieved overall accuracies above 0.9. The convergence of two methodologically different pipelines on similarly high accuracy strengthens the conclusion that the 688-SNP subset carries genuine, robust ancestry signal rather than artifacts of any single algorithm.
The implications for forensic practice are considerable. A moderate-density panel of this size can be genotyped efficiently with targeted sequencing or microarray platforms, at a fraction of the cost and data volume of genome-wide arrays, yet it delivers kinship resolution approaching that of much denser datasets for relationships out to second degree, and usable resolution for third degree when combined with machine learning classification. At the same time, the panel’s ability to distinguish southern from northern Chinese minorities and to place an unknown sample within East Asian substructure could help investigators narrow the geographic origin of unidentified remains or refine investigative leads. The study also adds to a growing literature on machine learning in forensic genetics, where algorithms are increasingly used to squeeze additional inferential power from standard marker sets rather than simply expanding the number of loci typed.
The authors note that the SNP panel had previously demonstrated strong forensic application value in Chinese Han populations, and the present validation in the Yunnan Zhuang group extends its evidence base to another major Chinese ethnic group. The work was funded by the Guangdong Provincial Science and Technology Program and the National Natural Science Foundation of China, and the underlying data are available from the corresponding author upon reasonable request, subject to privacy and ethical restrictions. As forensic laboratories worldwide weigh the transition from STR-centric workflows to SNP-based and sequencing-based platforms, studies like this one provide a practical benchmark: roughly 2,000 well-chosen SNPs, analyzed with a combination of classical likelihood statistics and accessible machine learning classifiers, appear sufficient to resolve close family relationships and to map fine-scale ancestry across one of the most genetically structured regions of the world.
Subject of Research: Forensic kinship identification and East Asian population discrimination using moderate-density SNPs and machine learning
Article Title: Moderate-density SNPs combined with machine learning method driven kinship identification and East Asian subpopulations discrimination analysis
Article References: Wu, X., Liu, Q., Luo, L., Lan, Q., & Zhu, B. (2026). Moderate-density SNPs combined with machine learning method driven kinship identification and East Asian subpopulations discrimination analysis. International Journal of Legal Medicine. https://doi.org/10.1007/s00414-026-04011-9
Image Credits: AI Generated
DOI: 10.1007/s00414-026-04011-9
Keywords: forensic genetics, SNP panel, kinship identification, machine learning, K-Nearest Neighbor, likelihood ratio, ancestry inference, East Asian populations, Yunnan Zhuang, population genetics, LASSO, International Journal of Legal Medicine
Cite Scienmag News
Teresa Odom. (September 26, 2026). Machine Learning Meets Forensic DNA: SNPs That Identify Relatives and Separate East Asian Populations. Scienmag. https://scienmag.com/machine-learning-meets-forensic-dna-snps-that-identify-relatives-and-separate-east-asian-populations/
Teresa Odom. "Machine Learning Meets Forensic DNA: SNPs That Identify Relatives and Separate East Asian Populations." Scienmag, 26 September 2026, https://scienmag.com/machine-learning-meets-forensic-dna-snps-that-identify-relatives-and-separate-east-asian-populations/. Accessed 26 September 2026.
Teresa Odom. "Machine Learning Meets Forensic DNA: SNPs That Identify Relatives and Separate East Asian Populations." Scienmag. September 26, 2026. https://scienmag.com/machine-learning-meets-forensic-dna-snps-that-identify-relatives-and-separate-east-asian-populations/

