<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>combined match probability &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/combined-match-probability/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 06 Oct 2026 06:30:29 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>combined match probability &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>How Big Must a DNA Database Be? Massive Argentine Study Redefines Forensic Sampling Rules</title>
		<link>https://scienmag.com/how-big-must-a-dna-database-be-massive-argentine-study-redefines-forensic-sampling-rules/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Tue, 06 Oct 2026 06:30:29 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[allele frequencies]]></category>
		<category><![CDATA[Argentine genetic diversity study]]></category>
		<category><![CDATA[Argentine population]]></category>
		<category><![CDATA[autosomal STR analysis]]></category>
		<category><![CDATA[combined match probability]]></category>
		<category><![CDATA[DNA database statistical foundations]]></category>
		<category><![CDATA[DNA profiling]]></category>
		<category><![CDATA[forensic DNA database size]]></category>
		<category><![CDATA[forensic DNA evidence reliability]]></category>
		<category><![CDATA[forensic genetics]]></category>
		<category><![CDATA[forensic sampling rules revision]]></category>
		<category><![CDATA[Genetic diversity]]></category>
		<category><![CDATA[genetic profile frequency estimation]]></category>
		<category><![CDATA[impact of database size on forensic accuracy]]></category>
		<category><![CDATA[implications for criminal justice]]></category>
		<category><![CDATA[International Journal of Legal Medicine]]></category>
		<category><![CDATA[large-scale genetic data collection]]></category>
		<category><![CDATA[population database]]></category>
		<category><![CDATA[population database sampling]]></category>
		<category><![CDATA[population genetics in forensic science]]></category>
		<category><![CDATA[power of exclusion]]></category>
		<category><![CDATA[rare alleles]]></category>
		<category><![CDATA[sample size]]></category>
		<category><![CDATA[STR loci]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=240482</guid>

					<description><![CDATA[A study of 8,237 Argentines genotyped at 22 autosomal STR loci shows that forensic parameters remain stable at modest sample sizes while capturing rare alleles requires far larger datasets.]]></description>
										<content:encoded><![CDATA[<p>Every courtroom conviction built on DNA evidence rests on a quiet statistical foundation: a population database that tells investigators how likely it is that a random person shares a particular genetic profile. For decades, forensic laboratories have assembled these databases with relatively modest sample sizes, often a few hundred individuals, assuming that such numbers adequately capture the genetic diversity of the populations they serve. A new study from Argentina, based on the largest autosomal STR dataset ever compiled for that country, now shows that this assumption deserves far more scrutiny than it has traditionally received.</p>
<p>The research, published in the International Journal of Legal Medicine, analyzed 8,237 unrelated Argentine individuals genotyped at 22 autosomal short tandem repeat loci, the repetitive DNA sequences that form the backbone of forensic identification worldwide. Rather than simply reporting allele frequencies for this enormous cohort, the team led by Antonella Belén Penacino and José Alonso Aguilar-Velázquez asked a more fundamental question: how does the size of the sample shape what the database appears to contain? To answer it, they generated 1,000 random resampling replicates for cohort sizes ranging from 500 individuals up to the full 8,237, effectively simulating thousands of alternative databases of different sizes drawn from the same population.</p>
<p>The results reveal a striking asymmetry between two properties of forensic databases that are often conflated. Across the 22 loci, the full dataset contained 344 distinct alleles, and nearly half of them, 169 alleles or 49.1 percent, occurred at frequencies below 1 percent. These rare alleles are the hidden tail of human genetic diversity, and it turns out they are extraordinarily expensive to capture. Overall allele recovery climbed quickly with sample size: a 500-person database captured 74.9 percent of the allelic diversity observed in the full cohort, a 3,000-person database reached 90.7 percent, and by 7,000 individuals the figure stood at 98.5 percent. Rare-allele recovery followed a much shallower trajectory, reaching only 81.0 percent at 3,000 individuals and still just 90.4 percent at 5,000.</p>
<p>This divergence matters because the two quantities serve different forensic purposes. The combined forensic parameters that courts care most about, the combined match probability and the combined power of exclusion, remained remarkably stable across all sampling scenarios in the study. The combined match probability showed minimal variation regardless of how many individuals were sampled, and the combined power of exclusion stayed consistently high throughout. In practical terms, even a 500-person database produced paternity and identification statistics that looked very similar to those derived from the full cohort of more than 8,000. For many routine applications, the headline numbers of forensic genetics appear almost indifferent to sample size.</p>
<p>Yet the stability of those combined parameters conceals what is happening at the level of individual alleles. When a DNA profile from a crime scene contains an allele that has never been observed in the reference database, laboratories must apply a minimum allele frequency, a floor value that prevents the match probability from being calculated as zero. The accuracy of that floor, and of the frequency estimates for genuinely rare alleles, depends directly on whether the database has sampled enough people to encounter them. A database that has recovered only 81 percent of the rare alleles present in its population is systematically underestimating the diversity that real casework will encounter.</p>
<p>The locus-specific analyses added another layer of nuance. Saturation dynamics, the point at which additional sampling stops revealing new alleles, were strongly associated with the proportion of rare alleles at each locus. Highly polymorphic markers such as FGA and D21S11, which harbor many low-frequency variants, required substantially larger sample sizes to achieve complete allele representation than less variable loci. This means that a single sample-size criterion applied uniformly across a multiplex panel is inherently misleading: some loci saturate quickly while others continue yielding new alleles thousands of samples later. The authors&#8217; finding suggests that adequacy assessments should be conducted locus by locus, weighted by the intended application of the database.</p>
<p>The Argentine context makes the study particularly significant. Argentina&#8217;s population reflects complex admixture among Indigenous American, European, and other ancestral contributions, and earlier work by some of the same research community, including studies of urban Argentine populations and regional databases from Patagonia and the central provinces, has documented meaningful genetic structure across the country. Building a reference database at this scale, with informed consent from all participants, provides the forensic community with a resource whose allele frequency estimates carry far narrower confidence intervals than the small regional datasets that have historically been used. It also offers a benchmark against which the sufficiency of smaller national databases elsewhere can be judged.</p>
<p>The methodological approach draws on concepts borrowed from ecology, where rarefaction and species-accumulation curves have long been used to estimate how much of a community&#8217;s biodiversity a survey has captured. The study&#8217;s citation of foundational work on individual-based rarefaction by Robert Colwell and colleagues signals this intellectual lineage: alleles at a forensic locus are treated much like species in an ecosystem, and the resampling replicates function as accumulation curves revealing how discovery slows as sampling proceeds. The team&#8217;s analytical toolkit included standard population genetics software such as Arlequin and the STRAF online platform for forensic STR evaluation, alongside the R statistical environment for the resampling analyses.</p>
<p>The implications for forensic practice are direct. International guidelines, including the revised recommendations for publishing genetic population data issued by leading forensic geneticists in 2017, have grappled with how large a population sample must be, and recent work by other groups has begun questioning conventional sampling guidance for highly polymorphic STR loci. The Argentine study provides the strongest empirical answer yet: the answer depends on the question being asked. If the goal is stable combined forensic parameters for routine match probability and paternity calculations, moderate sample sizes perform adequately. If the goal is comprehensive representation of allelic diversity, particularly the rare alleles that populate nearly half of the allelic spectrum, then even several thousand individuals may not suffice, and databases aiming at full allele recovery should plan for substantially larger sampling efforts.</p>
<p>Perhaps the most enduring contribution of the study is conceptual. By demonstrating that allelic-diversity recovery and forensic-parameter stability are distinct properties that respond differently to sample size, the researchers have given the forensic community a framework for evaluating population databases according to their intended use rather than a one-size-fits-all threshold. As DNA phenotyping advances, as new multiplex kits expand the number of loci typed, and as courts increasingly scrutinize the statistical foundations of DNA evidence, that distinction will only grow in importance. A database that looks statistically adequate on paper may still be missing half the rare genetic variants its population carries, and knowing exactly which questions a database can and cannot answer is now an empirical matter that studies of this scale are finally equipped to resolve.</p>
<p><strong>Subject of Research:</strong> Sample size effects on allele diversity and forensic parameters of autosomal STR loci in the Argentine population</p>
<p><strong>Article Title:</strong> Sample size effects on allele diversity, rare-allele recovery, and forensic parameters of 22 autosomal STRs in a cohort of 8,237 Argentines</p>
<p><strong>Article References:</strong> Penacino, A. B., Carvajal-Pérez, C. E., Rangel-Villalobos, H., Elsztein, L. D., Puentes, P. A., Zapata, F. A., Becerra-Loaiza, D. S., Moreno-Ortiz, J. M., Penacino, G. A., &amp; Aguilar-Velázquez, J. A. (2026). Sample size effects on allele diversity, rare-allele recovery, and forensic parameters of 22 autosomal STRs in a cohort of 8,237 Argentines. <em>International Journal of Legal Medicine</em>. <a href="https://doi.org/10.1007/s00414-026-04032-4" rel="noopener noreferrer">https://doi.org/10.1007/s00414-026-04032-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00414-026-04032-4" rel="noopener noreferrer">10.1007/s00414-026-04032-4</a></p>
<p><strong>Keywords:</strong> forensic genetics, STR loci, allele frequencies, rare alleles, sample size, population database, combined match probability, power of exclusion, Argentine population, genetic diversity, DNA profiling, International Journal of Legal Medicine</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">240482</post-id>	</item>
		<item>
		<title>Rwanda Builds First National DNA Fingerprint Baseline From 815 Profiles</title>
		<link>https://scienmag.com/rwanda-builds-first-national-dna-fingerprint-baseline-from-815-profiles/</link>
		
		<dc:creator><![CDATA[Juliet Wilcox]]></dc:creator>
		<pubDate>Mon, 21 Sep 2026 00:36:06 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[allele frequencies]]></category>
		<category><![CDATA[allele frequency analysis Rwanda]]></category>
		<category><![CDATA[autosomal STR loci Rwanda forensic science]]></category>
		<category><![CDATA[combined match probability]]></category>
		<category><![CDATA[development of national DNA database Rwanda]]></category>
		<category><![CDATA[DNA evidence in Rwandan courts]]></category>
		<category><![CDATA[DNA matching probability Rwanda]]></category>
		<category><![CDATA[DNA profiling]]></category>
		<category><![CDATA[forensic DNA]]></category>
		<category><![CDATA[forensic efficiency statistics Rwanda]]></category>
		<category><![CDATA[forensic genetic reference dataset Rwanda]]></category>
		<category><![CDATA[forensic genetics]]></category>
		<category><![CDATA[forensic genetics research Rwanda]]></category>
		<category><![CDATA[genetic profiling Rwanda criminal investigations]]></category>
		<category><![CDATA[human identification]]></category>
		<category><![CDATA[International Journal of Legal Medicine]]></category>
		<category><![CDATA[kinship analysis]]></category>
		<category><![CDATA[polymerase chain reaction]]></category>
		<category><![CDATA[population genetics]]></category>
		<category><![CDATA[Rwanda]]></category>
		<category><![CDATA[Rwanda national DNA fingerprint baseline]]></category>
		<category><![CDATA[short tandem repeats]]></category>
		<category><![CDATA[short tandem repeats (STRs) in forensic science]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204700</guid>

					<description><![CDATA[Researchers have compiled allele frequencies for 23 autosomal STR loci from 815 Rwandan individuals, creating a national DNA reference dataset with an extraordinarily low combined match probability.]]></description>
										<content:encoded><![CDATA[<p>Forensic scientists in Rwanda have compiled one of the most detailed genetic reference datasets ever assembled for the country, a milestone that could transform how DNA evidence is weighed in Rwandan courts. In a study published in the International Journal of Legal Medicine, researchers led by Aimable Ndungutse of the University of Rwanda and the Rwanda Forensic Institute, together with colleagues at the University Medical Center Hamburg–Eppendorf in Germany, report allele frequencies and forensic efficiency statistics for 23 autosomal short tandem repeat loci drawn from 815 individuals across the country. The work provides the statistical backbone that investigators, prosecutors, and defense attorneys need to translate a DNA match into a meaningful statement about probability.</p>
<p>Short tandem repeats, or STRs, are stretches of DNA in which a short sequence of two to six base pairs is repeated over and over. The number of repeats at a given locus varies widely between individuals, which is precisely what makes these markers so valuable in forensic science. Because a person inherits one copy of each autosomal locus from each parent, a profile across many independent STR loci becomes a molecular fingerprint so rare that the chance of two unrelated people sharing it is vanishingly small. But that claim is only as strong as the population data behind it. Allele frequencies differ among human populations, so calculating the probability of a random match requires knowing how common each repeat variant is in the relevant population. Without local frequency data, forensic statisticians must borrow figures from other groups, introducing uncertainty that can undermine confidence in courtrooms.</p>
<p>Rwanda&#8217;s forensic DNA capability has grown rapidly in recent years, with genetic evidence now routinely used in criminal investigations, paternity disputes, and civil cases. Yet comprehensive population-specific reference data had lagged behind. The first forensic STR study in the country analyzed a relatively small cohort of unrelated individuals with a limited marker panel, offering an initial allele frequency dataset with restricted coverage. Earlier work, including a 2004 study of allele distribution among Rwandan Tutsi and a 2003 analysis of 16 STR loci in Hutu individuals, provided valuable but narrow snapshots. The new study dramatically expands that foundation, both in sample size and in the number of markers characterized.</p>
<p>The researchers took a retrospective approach, drawing on archived STR genotype data generated between 2005 and 2015. Through database sampling, they retrieved all 815 profiles from unrelated individuals that met the study&#8217;s inclusion criteria. Because the material spanned a full decade of laboratory work, the team painstakingly reviewed laboratory records to verify the extraction and quantification methods, amplification systems, capillary electrophoresis platforms, allele-calling software, and quality assurance procedures used throughout the period. This methodological audit ensured that data generated under different protocols over the years could be combined coherently into a single reference dataset.</p>
<p>The laboratory workflow itself reflects standard forensic practice of the era. DNA was extracted using the Chelex 100 method, a resin-based technique that binds metal ions and inhibiting contaminants while releasing template DNA. For casework samples, quantification followed with the Quantifiler Duo DNA Quantification kit on an ABI 7500 Real-Time PCR System, allowing technicians to confirm both the quantity of human DNA and the presence of inhibitors before amplification. The STR amplification combined the PowerPlex 16 system with PowerPlex ESI 17 Pro and PowerPlex ESX 17 kits, yielding a combined panel of 23 autosomal STR loci, including the highly discriminating SE33 marker that is standard in European forensic practice.</p>
<p>The results confirm that all 23 loci are robustly polymorphic in the Rwandan population, but the degree of variation varies considerably from marker to marker. The number of observed alleles per locus ranged from just 7 at D16S539 to a remarkable 50 at SE33, one of the most variable STR loci in the human genome. At most loci, one or two alleles predominated while the remaining variants appeared at relatively low frequencies. Among the most common were allele 16 at D3S1358, with a frequency of approximately 0.339; allele 7 at TH01, at roughly 0.378; allele 12 at D13S317, at about 0.363; allele 10 at D7S820, at around 0.406; and allele 12 at D5S818, at approximately 0.368. These patterns echo those seen in other Bantu-speaking populations of sub-Saharan Africa, consistent with Rwanda&#8217;s demographic history, while also revealing alleles rare enough elsewhere to be locally informative.</p>
<p>The headline statistic of the study is the combined match probability across the 23-locus panel: 1.7239 times 10 to the power of minus 30. In practical terms, if two profiles match at all 23 loci, the chance that a randomly selected unrelated Rwandan individual would share that same profile is roughly one in a nonillion, a number so extreme that it effectively removes any plausible ambiguity about identity for unrelated individuals. This extraordinarily low figure reflects the high informativeness of the combined panel, driven especially by hyper-variable loci such as SE33. It also means that even partial profiles recovered from degraded crime scene samples, where only a subset of loci amplifies successfully, can still carry enormous evidential weight when interpreted against the new frequency data.</p>
<p>The forensic value of the dataset extends beyond match probabilities. Allele frequencies feed into every major statistical framework used in DNA interpretation, including likelihood ratios, paternity indices, and kinship analyses. In paternity testing, for example, the strength of evidence for or against fatherhood depends on how common the child&#8217;s paternal alleles are in the population; a rare allele shared between alleged father and child is far more persuasive than a common one. Similarly, in disaster victim identification and missing persons investigations, accurate frequency estimates are essential for weighing the possibility of coincidental matches among relatives. By providing nationally distributed data, the study reduces the geographic and ethnic sampling bias that plagued earlier, more localized efforts.</p>
<p>The work also carries scientific significance beyond the courtroom. Rwanda occupies a key position in studies of East African population history, and its STR variation contributes to a broader picture of genetic diversity in sub-Saharan Africa, the region with the deepest human genetic diversity on Earth. Recent whole-genome sequencing efforts across 44 indigenous African populations have underscored how undersampled much of the continent remains in genetic databases. Expanded forensic datasets like this one, together with earlier mitochondrial DNA studies covering Côte d&#8217;Ivoire and Rwanda, help fill critical gaps that affect both forensic statistics and population genetics research. The detection of rare alleles in the Rwandan panel adds to the growing catalog of global STR diversity and improves the precision of profile probability estimates not only locally but in international databases that incorporate African frequency data.</p>
<p>For Rwanda, the immediate implications are practical. The Rwanda Forensic Institute, the Rwanda National Police, and the National Public Prosecution Authority, all partners in the research, now have a defensible, population-specific statistical foundation for DNA testimony. As DNA evidence becomes more central to the justice system, courts will increasingly demand that match statistics rest on frequencies measured in the relevant population rather than approximations from distant groups. The study, funded by the University of Rwanda and the European Union Team Europe Initiative under the Kwigira Project, also represents a model of South–North scientific collaboration, pairing Rwandan institutions with forensic specialists in Hamburg. With the expanded characterization of highly polymorphic loci and the detection of rare alleles, the authors conclude that the findings strengthen the statistical basis of forensic DNA interpretation in Rwanda and consolidate the country&#8217;s forensic genetic resources for years to come.</p>
<p><strong>Subject of Research:</strong> Allele frequencies and forensic efficiency of autosomal STR loci in the Rwandan population</p>
<p><strong>Article Title:</strong> Allele frequencies and forensic efficiency of autosomal short tandem repeat loci in the Rwandan population</p>
<p><strong>Article References:</strong> Ndungutse, A., Daba, T. M., Krebs, O., Augustin, C., &amp; Mutesa, L. (2026). Allele frequencies and forensic efficiency of autosomal short tandem repeat loci in the Rwandan population. <em>International Journal of Legal Medicine</em>. <a href="https://doi.org/10.1007/s00414-026-04022-6" rel="noopener noreferrer">https://doi.org/10.1007/s00414-026-04022-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00414-026-04022-6" rel="noopener noreferrer">10.1007/s00414-026-04022-6</a></p>
<p><strong>Keywords:</strong> forensic genetics, short tandem repeats, allele frequencies, Rwanda, DNA profiling, human identification, kinship analysis, polymerase chain reaction, combined match probability, International Journal of Legal Medicine, population genetics, forensic DNA</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204700</post-id>	</item>
	</channel>
</rss>
