Friday, September 11, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Science Education

Studying differential item functioning in Colombia’s large scale SABER 11 tests

September 11, 2026
in Science Education
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 7 mins read
0
Studying differential item functioning in Colombia’s large scale SABER 11 tests

Studying differential item functioning in Colombia’s large scale SABER 11 tests

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

When Colombian high-school students sit the national SABER 11 examination each year, they are taking two different versions of the test: one in March, aimed mostly at students whose academic year begins in August, and one in September for the far larger cohort whose school year starts in February. The two cohorts come from strikingly different educational worlds, with the March group dominated by private schools and the September group drawn overwhelmingly from public schools, so the two groups differ substantially in ability. That makes it essential to verify that the two parallel test forms are truly comparable, and a new study in the journal Large-scale Assessments in Education shows how to do it, while revealing that some of the standard statistical wisdom about fairness testing breaks down under the extreme conditions that large-scale assessments routinely create.

The study, conducted by John Alexander Calderón and Nelson Andrés Rodríguez of Colombia’s National Institute for Educational Evaluation (Icfes) and Víctor H. Cervantes of the University of Illinois at Urbana-Champaign, examined whether items on the Mathematics test of SABER 11 functioned differently for the two testing populations. This question is central to a psychometric property called differential item functioning, or DIF, the phenomenon in which examinees of equal underlying ability have different probabilities of answering an item correctly depending on which group they belong to. If items show DIF between groups, then score comparisons between those groups may reflect bias rather than genuine differences in the measured trait, undermining the fairness of any conclusions drawn from the results.

DIF has been a concern in testing since at least the 1960s, but interest intensified after the 1984 “Golden Rule” settlement in the United States, which pushed the testing industry to distinguish statistically between real group differences and bias against particular groups. Under item response theory, the framework most large-scale assessments use for scoring, DIF is defined precisely: for an item scored correct or incorrect, a two-parameter logistic model assigns each item a discrimination parameter, describing how sharply the item separates high-ability from low-ability examinees, and a difficulty parameter, locating the ability level at which an examinee has a fifty percent chance of answering correctly. If either parameter differs across groups, the item’s characteristic curves diverge. Uniform DIF corresponds to a difference in difficulty alone, producing a constant shift between the curves; non-uniform DIF involves a difference in discrimination, so the curves cross and the group advantage varies with ability; and mixed DIF involves both.

Because SABER 11 already uses item response theory for its scaling and scoring, the researchers chose an IRT-based DIF procedure, the non-compensatory DIF (NCDIF) index from Raju’s Differential Functioning of Items and Tests framework, complemented by the widely used Mantel–Haenszel procedure. The NCDIF index quantifies the area between the two groups’ item characteristic curves, weighted by the distribution of ability in the focal group, meaning the group of special interest. This weighting has an attractive property: it emphasizes parameter differences where they matter most for the focal group’s actual scores. The statistical test on the NCDIF index is conducted through parametric bootstrap, known in the DFIT framework as item parameter replication, in which the item parameters are repeatedly re-estimated from simulated data to build the distribution of the statistic under the null hypothesis of no DIF.

Following a framework laid out by Sireci and Rios for tailoring DIF analyses to specific testing contexts, the team had to make a series of practical decisions: which detection method to use, how to define the comparison groups and their sample sizes, how to construct the matching variable on which the two groups are compared, whether to incorporate effect size measures, and at what level to analyze the results. Most of these choices could be settled from the existing literature. But two could not, because no published research had explored the performance of the NCDIF index under conditions that match SABER 11: sample sizes reaching roughly 40,000 examinees in the majority group against about 1,500 in the minority group, a ratio of up to 1:25, combined with a moderate gap in mean ability between the two populations.

To resolve this, the researchers ran a series of simulation studies in which they generated response data for test forms mirroring the real structure of SABER 11, with 40-item forms sharing an anchor of half their items, using item parameters drawn from the actual operational pool of SABER 11 mathematics items rather than from artificially clean, well-distributed parameter sets. They manipulated sample size and ratio, the impact between groups (shifting the focal group’s mean ability from zero up to plus or minus 0.8 standard deviations), the purification of the matching variable, and the use of effect size classifications, with 200 replications per condition. The purification step matters because the matching variable itself must be free of DIF contamination: the two groups’ ability scales are linked using common items, and if items with DIF are included in the linking, the resulting bias can masquerade as or mask genuine DIF. The researchers used a two-stage purification in which the linking is first performed with all items, suspect items are removed, and the linking is repeated with the purified set.

The simulation results overturned a piece of conventional advice. Previous studies had suggested that DIF analyses work best when the two groups being compared have similar sample sizes, and it had been recommended that the larger group be subsampled to match the smaller one. But in these simulations, the Type I error rates, meaning the rate of incorrectly flagging items as showing DIF when they truly do not, were no worse for the extreme 1,500-versus-40,000 condition than for equal-sized groups of 40,000 each. In fact, the error rates remained less well controlled precisely in the conditions with equal sample sizes at the 40,000 level. The contrast between the 1,500-to-1,500 and 1,500-to-40,000 conditions showed no statistically significant difference in either Type I error or power after purification. The practical conclusion was clear: there was no reason to discard data and subsample the larger group.

The second major finding concerned the interplay of impact and effect sizes. Before purification, Type I error rates ballooned in the largest samples, especially when there was a true difference in mean ability between the groups, an asymmetry that also depended on whether the focal group was favored or disadvantaged. After purification, the NCDIF index’s error rates approached nominal levels, but the inflation caused by impact remained. Crucially, applying effect size guidelines, which classify flagged items by the practical magnitude of the difference rather than statistical significance alone, drove the false positive rates to nearly zero across almost all conditions, while barely reducing power for detecting moderate or large DIF given the enormous sample sizes involved. For the NCDIF procedure, power was near-perfect for items with genuinely non-negligible DIF. An analysis of variance confirmed that nearly all interactions among the experimental factors were significant, with the largest effects tied to the interplay of sample size ratio, the number of DIF items, and the use of effect sizes.

With these design decisions settled, the team applied the full protocol to actual SABER 11 data: the Mathematics test form administered on the second date of 2018, with 39,377 examinees as the reference group, and the form from the first date of 2019, with 1,508 examinees as the focal group. After purification, the item parameter replication test flagged eight to ten of the 22 common items, and the Mantel–Haenszel procedure flagged three to four. But when the effect size classifications were applied, only a single item, item 18, was classified as showing non-negligible DIF. Its NCDIF value was 0.03096 and its Mantel–Haenszel delta was 1.6386, placing it in the “large DIF” category. Once this item was removed from the common set, the Stocking–Lord scale linking between the two forms yielded the transformation constants needed to place both groups’ abilities on a single scale, with the residual mean difference of 0.59 favoring the August-cohort group, consistent with the historical advantage of that subpopulation.

The content review of item 18 proved instructive. The item asks students to judge whether three different algebraic procedures for solving the equation (x + 2)(x + 3) = 5(x + 3) were performed correctly, with each student’s work shown step by step. The researchers found that although two of the procedures were on track toward the correct solution, none of them actually reached a final answer, and the point at which the student Nelson’s work stopped could appear incorrect to examinees from the September cohort more frequently than to those from the March cohort. For an examinee of ability 1.0 on the scale where the March group’s abilities were standardized, the probability of a correct response was roughly 0.62 for the March group but below 0.35 for the September group. The item’s characteristic curves crossed at an ability value of about −0.824, meaning the difference was small near the average of the focal group but grew rapidly with ability. Icfes’s mathematics team has since revised the item, making the three procedures more explicit and ensuring each reaches a final answer.

The study’s implications extend well beyond Colombia. The authors emphasize that detecting DIF is only the first step toward fairness, and that content review, think-aloud protocols with students, teacher interviews, and curricular analyses should follow to understand why an item functions differently. They also caution that future simulation studies of DIF statistics, whether NCDIF, Mantel–Haenszel, or any other index, should draw item parameters from realistic operational pools rather than sanitized, evenly distributed sets, since the distribution of item difficulties within a test interacts with detection behavior in ways that clean simulated pools fail to capture. In their data, the average difficulty of common items was shifted by 0.66 relative to non-common items, an asymmetry that may itself interact with detection rates. For practitioners running large-scale assessments anywhere in the world, the message is that statistical significance alone, with samples in the tens of thousands, can manufacture bias where none exists, and that effect size classification, scale purification, and realistic item parameter pools are the practical safeguards against turning fairness checks into false alarms.

Subject of Research: Differential item functioning analysis in large-scale assessments, applied to the Mathematics test of Colombia’s SABER 11 examination

Subject of Research: Science Education

Article Title: Differential item functioning analysis in large scale assessments: a case study for DIF in SABER 11

Article References: Calderón, J. A., Rodríguez, N. A., & Cervantes, V. H. (2026). Differential item functioning analysis in large scale assessments: a case study for DIF in SABER 11. Large-scale Assessments in Education, 14(1), Article 23. https://doi.org/10.1186/s40536-026-00294-x

Image Credits: AI Generated

DOI: 10.1186/s40536-026-00294-x

Keywords: differential item functioning, DIF, large scale assessments, SABER 11, item response theory, NCDIF index, Mantel–Haenszel, test fairness, effect size, scale purification, validity evidence, psychometrics

Cite Scienmag News

Courtney Benton. (September 11, 2026). Studying differential item functioning in Colombia’s large scale SABER 11 tests. Scienmag. https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/

Courtney Benton. "Studying differential item functioning in Colombia’s large scale SABER 11 tests." Scienmag, 11 September 2026, https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/. Accessed 11 September 2026.

Courtney Benton. "Studying differential item functioning in Colombia’s large scale SABER 11 tests." Scienmag. September 11, 2026. https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/

Tags: assessment comparability across cohortsColombian high school examinationsColombian high-school testingdifferential item functioningeducational equity in Colombiaeducational equity in testingeducational evaluation in Colombiafairness in standardized testingfairness testing challenges in large assessmentsimpact of socioeconomic factors on test performancelarge-scale assessments in educationnational student assessment validityprivate vs public school performancepsychometric analysis of test itemspsychometric evaluationpublic vs private school performanceSABER 11 examinationSABER 11 test analysisstatistical challenges in DIF analysisstatistical methods in DIF analysistest fairness and validitytest form comparability
Share26Tweet16
Previous Post

Sedentary behaviour change: barriers and strategies reported by individuals with type 2 diabetes during a multi-component intervention

Next Post

What Doctors Wear Shapes Patient Trust and Infection Fears in Sri Lanka

Related Posts

Medical interns share experiences running quality improvement projects in South African primary care
Science Education

Medical interns share experiences running quality improvement projects in South African primary care

September 11, 2026
Professor wins grant to study AI in early childhood education
Science Education

Professor wins grant to study AI in early childhood education

September 10, 2026
Navigating Stigma and Access: Cancer Care Experiences of Adults with Intellectual Disability
Science Education

Navigating Stigma and Access: Cancer Care Experiences of Adults with Intellectual Disability

September 10, 2026
Gamified web platform boosts active learning in corporate education
Science Education

Gamified web platform boosts active learning in corporate education

September 10, 2026
Growth mindset and motivation drive student success in math powerhouse nations
Science Education

Growth mindset and motivation drive student success in math powerhouse nations

September 10, 2026
Virtual and augmented reality transform neurosurgical training, systematic review finds
Science Education

Virtual and augmented reality transform neurosurgical training, systematic review finds

September 10, 2026
Next Post
What Doctors Wear Shapes Patient Trust and Infection Fears in Sri Lanka

What Doctors Wear Shapes Patient Trust and Infection Fears in Sri Lanka

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Groundwater Arsenic Leaves Fingerprints in DNA Repair Genes of Exposed Women
  • Marine Bacteria’s Fengycin Emerges as a Powerful Eco-Friendly Antifouling Candidate
  • What Doctors Wear Shapes Patient Trust and Infection Fears in Sri Lanka
  • Studying differential item functioning in Colombia’s large scale SABER 11 tests

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading