When Colombian high-school students sit the national SABER 11 examination each year, they are taking two different versions of the test: one in March, aimed mostly at students whose academic year begins in August, and one in September for the far larger cohort whose school year starts in February. The two cohorts come from strikingly different educational worlds, with the March group dominated by private schools and the September group drawn overwhelmingly from public schools, so the two groups differ substantially in ability. That makes it essential to verify that the two parallel test forms are truly comparable, and a new study in the journal Large-scale Assessments in Education shows how to do it, while revealing that some of the standard statistical wisdom about fairness testing breaks down under the extreme conditions that large-scale assessments routinely create.
The study, conducted by John Alexander Calderón and Nelson Andrés Rodríguez of Colombia’s National Institute for Educational Evaluation (Icfes) and Víctor H. Cervantes of the University of Illinois at Urbana-Champaign, examined whether items on the Mathematics test of SABER 11 functioned differently for the two testing populations. This question is central to a psychometric property called differential item functioning, or DIF, the phenomenon in which examinees of equal underlying ability have different probabilities of answering an item correctly depending on which group they belong to. If items show DIF between groups, then score comparisons between those groups may reflect bias rather than genuine differences in the measured trait, undermining the fairness of any conclusions drawn from the results.
DIF has been a concern in testing since at least the 1960s, but interest intensified after the 1984 “Golden Rule” settlement in the United States, which pushed the testing industry to distinguish statistically between real group differences and bias against particular groups. Under item response theory, the framework most large-scale assessments use for scoring, DIF is defined precisely: for an item scored correct or incorrect, a two-parameter logistic model assigns each item a discrimination parameter, describing how sharply the item separates high-ability from low-ability examinees, and a difficulty parameter, locating the ability level at which an examinee has a fifty percent chance of answering correctly. If either parameter differs across groups, the item’s characteristic curves diverge. Uniform DIF corresponds to a difference in difficulty alone, producing a constant shift between the curves; non-uniform DIF involves a difference in discrimination, so the curves cross and the group advantage varies with ability; and mixed DIF involves both.
Because SABER 11 already uses item response theory for its scaling and scoring, the researchers chose an IRT-based DIF procedure, the non-compensatory DIF (NCDIF) index from Raju’s Differential Functioning of Items and Tests framework, complemented by the widely used Mantel–Haenszel procedure. The NCDIF index quantifies the area between the two groups’ item characteristic curves, weighted by the distribution of ability in the focal group, meaning the group of special interest. This weighting has an attractive property: it emphasizes parameter differences where they matter most for the focal group’s actual scores. The statistical test on the NCDIF index is conducted through parametric bootstrap, known in the DFIT framework as item parameter replication, in which the item parameters are repeatedly re-estimated from simulated data to build the distribution of the statistic under the null hypothesis of no DIF.
Following a framework laid out by Sireci and Rios for tailoring DIF analyses to specific testing contexts, the team had to make a series of practical decisions: which detection method to use, how to define the comparison groups and their sample sizes, how to construct the matching variable on which the two groups are compared, whether to incorporate effect size measures, and at what level to analyze the results. Most of these choices could be settled from the existing literature. But two could not, because no published research had explored the performance of the NCDIF index under conditions that match SABER 11: sample sizes reaching roughly 40,000 examinees in the majority group against about 1,500 in the minority group, a ratio of up to 1:25, combined with a moderate gap in mean ability between the two populations.
To resolve this, the researchers ran a series of simulation studies in which they generated response data for test forms mirroring the real structure of SABER 11, with 40-item forms sharing an anchor of half their items, using item parameters drawn from the actual operational pool of SABER 11 mathematics items rather than from artificially clean, well-distributed parameter sets. They manipulated sample size and ratio, the impact between groups (shifting the focal group’s mean ability from zero up to plus or minus 0.8 standard deviations), the purification of the matching variable, and the use of effect size classifications, with 200 replications per condition. The purification step matters because the matching variable itself must be free of DIF contamination: the two groups’ ability scales are linked using common items, and if items with DIF are included in the linking, the resulting bias can masquerade as or mask genuine DIF. The researchers used a two-stage purification in which the linking is first performed with all items, suspect items are removed, and the linking is repeated with the purified set.
The simulation results overturned a piece of conventional advice. Previous studies had suggested that DIF analyses work best when the two groups being compared have similar sample sizes, and it had been recommended that the larger group be subsampled to match the smaller one. But in these simulations, the Type I error rates, meaning the rate of incorrectly flagging items as showing DIF when they truly do not, were no worse for the extreme 1,500-versus-40,000 condition than for equal-sized groups of 40,000 each. In fact, the error rates remained less well controlled precisely in the conditions with equal sample sizes at the 40,000 level. The contrast between the 1,500-to-1,500 and 1,500-to-40,000 conditions showed no statistically significant difference in either Type I error or power after purification. The practical conclusion was clear: there was no reason to discard data and subsample the larger group.
The second major finding concerned the interplay of impact and effect sizes. Before purification, Type I error rates ballooned in the largest samples, especially when there was a true difference in mean ability between the groups, an asymmetry that also depended on whether the focal group was favored or disadvantaged. After purification, the NCDIF index’s error rates approached nominal levels, but the inflation caused by impact remained. Crucially, applying effect size guidelines, which classify flagged items by the practical magnitude of the difference rather than statistical significance alone, drove the false positive rates to nearly zero across almost all conditions, while barely reducing power for detecting moderate or large DIF given the enormous sample sizes involved. For the NCDIF procedure, power was near-perfect for items with genuinely non-negligible DIF. An analysis of variance confirmed that nearly all interactions among the experimental factors were significant, with the largest effects tied to the interplay of sample size ratio, the number of DIF items, and the use of effect sizes.
With these design decisions settled, the team applied the full protocol to actual SABER 11 data: the Mathematics test form administered on the second date of 2018, with 39,377 examinees as the reference group, and the form from the first date of 2019, with 1,508 examinees as the focal group. After purification, the item parameter replication test flagged eight to ten of the 22 common items, and the Mantel–Haenszel procedure flagged three to four. But when the effect size classifications were applied, only a single item, item 18, was classified as showing non-negligible DIF. Its NCDIF value was 0.03096 and its Mantel–Haenszel delta was 1.6386, placing it in the “large DIF” category. Once this item was removed from the common set, the Stocking–Lord scale linking between the two forms yielded the transformation constants needed to place both groups’ abilities on a single scale, with the residual mean difference of 0.59 favoring the August-cohort group, consistent with the historical advantage of that subpopulation.
The content review of item 18 proved instructive. The item asks students to judge whether three different algebraic procedures for solving the equation (x + 2)(x + 3) = 5(x + 3) were performed correctly, with each student’s work shown step by step. The researchers found that although two of the procedures were on track toward the correct solution, none of them actually reached a final answer, and the point at which the student Nelson’s work stopped could appear incorrect to examinees from the September cohort more frequently than to those from the March cohort. For an examinee of ability 1.0 on the scale where the March group’s abilities were standardized, the probability of a correct response was roughly 0.62 for the March group but below 0.35 for the September group. The item’s characteristic curves crossed at an ability value of about −0.824, meaning the difference was small near the average of the focal group but grew rapidly with ability. Icfes’s mathematics team has since revised the item, making the three procedures more explicit and ensuring each reaches a final answer.
The study’s implications extend well beyond Colombia. The authors emphasize that detecting DIF is only the first step toward fairness, and that content review, think-aloud protocols with students, teacher interviews, and curricular analyses should follow to understand why an item functions differently. They also caution that future simulation studies of DIF statistics, whether NCDIF, Mantel–Haenszel, or any other index, should draw item parameters from realistic operational pools rather than sanitized, evenly distributed sets, since the distribution of item difficulties within a test interacts with detection behavior in ways that clean simulated pools fail to capture. In their data, the average difficulty of common items was shifted by 0.66 relative to non-common items, an asymmetry that may itself interact with detection rates. For practitioners running large-scale assessments anywhere in the world, the message is that statistical significance alone, with samples in the tens of thousands, can manufacture bias where none exists, and that effect size classification, scale purification, and realistic item parameter pools are the practical safeguards against turning fairness checks into false alarms.
Cite Scienmag News
Courtney Benton. (September 11, 2026). Studying differential item functioning in Colombia’s large scale SABER 11 tests. Scienmag. https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/
Courtney Benton. "Studying differential item functioning in Colombia’s large scale SABER 11 tests." Scienmag, 11 September 2026, https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/. Accessed 11 September 2026.
Courtney Benton. "Studying differential item functioning in Colombia’s large scale SABER 11 tests." Scienmag. September 11, 2026. https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/

