DNA methylation has become one of the most intensively studied layers of human biology, a chemical annotation system written onto the genome that helps determine which genes are active in which cells and tissues. Because methylation patterns shift with age, environment, and disease, researchers have invested heavily in high-throughput platforms that can read these marks across the entire genome at low cost. The Illumina Infinium MethylationEPIC BeadChip, known simply as the EPIC array, is currently the workhorse of that effort, interrogating more than 850,000 cytosine modification sites across a wide range of genomic features, including CpG islands, open chromatin regions identified by the ENCODE project, DNase hypersensitive sites, and more than 90 percent of the CpGs covered by its predecessor, the 450K array. Yet a new analysis published in Epigenetics Communications by Zhou Zhang, Chang Zeng, and Wei Zhang of Northwestern University Feinberg School of Medicine delivers a sobering technical audit of the platform, revealing that a substantial fraction of its probes carry hidden liabilities that could distort results, particularly in studies of non-European populations.
The team set out to answer a deceptively simple question: which of the EPIC array’s 866,836 CpG probes actually behave as designed? The answer matters because epigenome-wide association studies, or EWAS, depend on the assumption that each probe reliably measures methylation at one specific genomic location. When a probe fails that assumption, the resulting signal can reflect genetic variation, off-target binding, or outright misannotation rather than genuine methylation differences. Such artifacts have plagued earlier Infinium platforms, and previous studies of the 450K and EPIC arrays had already flagged between 6 and 11 percent of all probes as problematic. The Northwestern group extended that work by systematically characterizing the entire EPIC array and, crucially, by asking how the answers change depending on which human population a study examines.
The first stage of the analysis involved mapping every probe sequence back to the human genome reference. This is not as straightforward as it sounds, because the Infinium assay relies on bisulfite conversion, a chemical treatment that converts unmethylated cytosines to uracil, which is read as thymine, while leaving methylated cytosines untouched. To represent all possible post-conversion states, the researchers generated four single-stranded versions of the genome in silico: forward and reverse strands, each in a methylated and unmethylated configuration, with all non-CpG cytosines converted to thymines. A further complication arises with Type II probes, some of which contain ambiguous R nucleotides that can represent either adenine or guanine depending on the methylation state of an internal CpG site. The team replaced every R with all possible A and G combinations, ultimately producing 3,505,864 distinct probe sequences, which they then aligned to the bisulfite-converted reference using the BLAT alignment tool.
The results of this exhaustive mapping exercise were striking. Of the 866,836 probes on the array, 91,034 failed the alignment criteria, which allowed no more than two mismatches and no more than four ambiguous bases, and required a match at the fiftieth nucleotide of each 50-base probe. In other words, roughly one in ten probes did not align reliably to the human genome reference at all. Beyond that, 21,256 probes mapped ambiguously to multiple loci, meaning their signals represent a mixture of genuine methylation at the intended site and spurious signal from elsewhere in the genome. The researchers also compared their mapping results with the original annotations supplied by Illumina and found that 448 probes were assigned to a different gene than the manufacturer’s annotation file claimed, a small but potentially misleading discrepancy for any study interpreting methylation changes in terms of nearby gene regulation.
To understand where the ambiguous probes sit in the functional landscape of the genome, the team performed an enrichment analysis using regulatory element annotations from the ENCODE Analysis Hub at the European Bioinformatics Institute, including DNase I hypersensitivity sites, transcription factor binding sites, and histone modification peaks pooled across cell lines. They compared the observed number of overlaps between ambiguous probes and each regulatory element against a null distribution generated by randomly sampling matched sets of array probes 10,000 times. The ambiguous probes generally followed the genomic distribution of all EPIC probes, but they showed a significant trend of enrichment for H3K9me3, a histone mark associated with constitutive heterochromatin and transcriptionally silent regions, with an empirical p-value below 0.001. This suggests that problematic probes are disproportionately located in compacted, repetitive portions of the genome, precisely where cross-hybridization is most likely to occur.
The second, and arguably more novel, stage of the study examined how common genetic variation within probe sequences varies across human populations. The researchers drew on variant call data from the 1000 Genomes Project, covering 26 global populations organized into five major continental groups: African, European, East Asian, South Asian, and Admixed American. For each bi-allelic SNP in each population, they calculated allele frequencies and then searched the 754,546 unambiguous EPIC probes for common SNPs, defined as those with a minor allele frequency above 0.05, within 20 bases of the interrogated CpG site. This window matters because SNPs near or within the probe sequence can alter hybridization efficiency, causing the measured fluorescence to reflect the genetic variant rather than the methylation state, a phenomenon known as methylation quantitative trait locus confounding when it propagates into association results.
The population-specific findings were pronounced. East Asian populations carried the fewest common SNP-containing probes, 54,810, followed by European populations with 68,155, South Asian with 70,068, Admixed American with 70,576, and African populations with 96,674. That African populations show the highest count is consistent with the well-established observation that genetic diversity is greatest in populations of African ancestry, a legacy of human demographic history. The practical implication is uncomfortable for the field: a probe that performs cleanly in an East Asian cohort may carry a common SNP in an African cohort, and vice versa. Any EWAS conducted in a diverse or non-European population using a single, population-agnostic probe exclusion list risks either discarding usable data or, worse, retaining probes that generate population-specific artifacts that masquerade as epigenetic differences between groups.
To address this, the authors generated a curated list of optimal CpG probes for each major global population, released as supplemental tables alongside the paper. The resource allows investigators to filter their EPIC array results according to the ancestry of their study cohort, excluding probes that are cross-hybridizing, misannotated, or polymorphic in the relevant population. The authors envision a straightforward workflow: researchers who identify differential methylation at any CpG using the EPIC array can cross-reference their hits against this resource to confirm that their findings are not driven by cross-hybridization or by common SNPs specific to the population under study. Given that the supplemental tables enumerate both the problematic probes and the population-specific SNP content of the array, the resource functions as a practical quality-control layer that can be adopted without new experiments.
The broader significance of the study extends beyond a single chip. Methylation arrays remain far more cost-effective than whole-genome bisulfite sequencing for large cohort studies, and they will continue to underpin EWAS in epidemiology, cancer research, aging studies, and environmental epigenetics for years to come. But the field has increasingly recognized that genetic variation and epigenetic variation are entangled, and that interpreting methylation differences between populations without accounting for underlying allele frequency differences can produce spurious conclusions about epigenetic contributions to disease. The Northwestern analysis provides a concrete, quantified demonstration of the scale of that problem on the most widely used current platform, showing that the interplay between probe design and population genetics is not a marginal concern but affects tens of thousands of probes.
For the research community, the message is clear: the EPIC array remains a powerful and efficient tool, but its outputs are only as reliable as the quality control applied to them. Studies of diverse human populations stand to benefit most from the new characterization, since they face the greatest risk of population-specific artifacts. As epigenetic research expands beyond historically overrepresented European cohorts, resources like this probe-level audit will be essential for ensuring that discoveries about the epigenetic roots of complex traits and diseases are genuine, reproducible, and applicable to all of humanity rather than a genetically narrow slice of it.
Subject of Research: Technical characterization of the Illumina EPIC DNA methylation array for population-diverse epigenetic research
Article Title: Characterization of the Illumina EPIC array for optimal applications in epigenetic research targeting diverse human populations
Article References: Zhang, Z., Zeng, C., & Zhang, W. (2022). Characterization of the Illumina EPIC array for optimal applications in epigenetic research targeting diverse human populations. Epigenetics Communications, 2(1), Article 7. https://doi.org/10.1186/s43682-022-00015-9
Image Credits: AI Generated
DOI: 10.1186/s43682-022-00015-9
Keywords: DNA methylation, EPIC array, Illumina, epigenomics, cross-hybridization, single nucleotide polymorphism, 1000 Genomes Project, EWAS, population genetics, probe quality control, bisulfite conversion, genomic annotation
Cite Scienmag News
Juliet Wilcox. (October 3, 2026). Hidden Flaws in a Leading DNA Methylation Chip Could Skew Studies of Diverse Populations. Scienmag. https://scienmag.com/hidden-flaws-in-a-leading-dna-methylation-chip-could-skew-studies-of-diverse-populations/
Juliet Wilcox. "Hidden Flaws in a Leading DNA Methylation Chip Could Skew Studies of Diverse Populations." Scienmag, 3 October 2026, https://scienmag.com/hidden-flaws-in-a-leading-dna-methylation-chip-could-skew-studies-of-diverse-populations/. Accessed 3 October 2026.
Juliet Wilcox. "Hidden Flaws in a Leading DNA Methylation Chip Could Skew Studies of Diverse Populations." Scienmag. October 3, 2026. https://scienmag.com/hidden-flaws-in-a-leading-dna-methylation-chip-could-skew-studies-of-diverse-populations/

