Friday, October 2, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Psychology & Psychiatry

Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests

October 2, 2026
in Psychology & Psychiatry
Glenn Wilkins
By Glenn Wilkins Scienmag Editorial Profile - Clinical Psychology
Reading Time: 6 mins read
0
Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests

Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests

Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Cognitive ability tests have long been the gold standard of employee selection, prized for their unmatched power to predict workplace performance. Yet for nearly a century, the same tests have shown persistent average score differences between demographic groups, casting a long shadow over their use in hiring. A sweeping systematic review published in Trends in Psychology by Stephen Cuppello, Lara D. Zibarras and Philip J. Corr of City St George’s, University of London, has now brought together the sprawling evidence on why these gaps appear, and the answer is more nuanced than either side of the long-running debate has typically admitted. After screening more than 1,100 papers and analysing 225 studies in depth, the authors conclude that while biological and environmental factors genuinely contribute to group differences, so too do a cluster of factors rooted in the design and administration of the tests themselves, most notably stereotype threat, differences in domain experience and self-confidence.

The scale of the problem the review addresses is considerable. Meta-analytic research has found roughly a one standard deviation gap in mean scores between Black and White samples in occupational testing, with a more recent analysis by Sackett and colleagues reporting an effect size of d = .79, still a large difference by psychological standards. Socioeconomic status shows a more modest association, with childhood SES explaining around five percent of adult intelligence in one major study. Gender differences, by contrast, are largely absent from overall general intelligence scores, largely because early test builders such as Binet and Terman deliberately constructed their instruments to avoid them, but they emerge clearly within specific cognitive domains: men tend to score higher on spatial and numerical tasks, women on verbal ones. These patterns matter because cognitive ability tests remain the single strongest predictor of job performance for candidates without prior experience in a role, a conclusion drawn from 85 years of selection research by Schmidt and Hunter.

To untangle the competing explanations, the researchers followed PRISMA-P systematic review protocols, searching PsycINFO, PsychArticles, Web of Science and Business Source Ultimate with deliberately broad search terms designed to capture multiple and conflicting viewpoints. From an initial haul of 1,124 papers, 316 were removed as duplicates, 541 were screened out on abstracts and 140 failed the inclusion criteria, leaving 127 papers that were fully reviewed. A further 98 studies were added through backward and forward citation searching, bringing the total to 225. The team excluded child and clinical samples, non-human research and studies of cognitive decline, focusing instead on healthy working-aged adults, and they deliberately excluded research based solely on standardised academic tests such as the SAT, whose heavy reliance on test preparation would have skewed the findings. Each paper was tagged against emerging explanatory accounts, ultimately yielding ten factors: biological differences, environmental differences, latent trait and measurement invariance, criterion validity, item bias, test-taking behaviour, anxiety, attitudes, experience and stereotype threat.

The biological evidence, drawn from 63 papers, is dominated by research on gender. Brain imaging studies have repeatedly found differences in neural activation between women and men during cognitive tasks, although these are not universal and one study found activation differences inconsistent with the cognitive domains that actually showed behavioural gaps. Hormonal accounts have attracted enormous attention: testosterone has been linked to spatial performance in both sexes, oestrogen to verbal performance, and experimental administration of testosterone has improved spatial performance in women. Several studies found women’s performance fluctuates across the menstrual cycle, with better spatial performance during the menstrual phase and better verbal performance during the luteal phase, though not all studies replicate these effects. Twin research has provided modest support for genetic contributions, and one longitudinal study found pubertal testosterone predicted adult spatial performance in men. Crucially, however, 61 of the 63 biological papers concerned gender, only two touched ethnicity and none addressed socioeconomic status, and no biological studies were conducted on high-stakes testing.

Environmental factors, examined in 15 papers, showed unambiguous support. Parental education, individual education level, income, language proficiency and generational immigration status, all of which vary by ethnicity and socioeconomic background, were related to cognitive ability scores. Strikingly, ethnic differences were significantly reduced when controlling for years of education, language proficiency and immigration status, and differences were more pronounced on verbal tasks, a pattern consistent with environmental contribution. Childhood preference for gendered spatial toys and childhood spatial play predicted adult spatial performance, and oral contraceptive use and type influenced verbal and spatial task performance. These findings carry a sober implication for employers: because environmental factors are unlikely to be mitigated through changes to testing procedures, no redesign of a test will ever fully eliminate group differences.

On the technical side of the ledger, the review found that test bias itself plays a smaller role than critics have often assumed. Across 23 papers on item bias and differential item functioning, studies generally either failed to identify biased items or found that biased items negligibly affected overall scores, even though high-quality studies with large samples confirmed that differential item functioning does exist on individual items. Interestingly, several studies found that using human figures rather than abstract shapes as item content reduced the gender gap in mental rotation performance, and gender-stereotyped item content sometimes created bias. Criterion validity studies, which draw on genuine high-stakes testing data, showed that cognitive tests frequently do not underpredict the job performance of women or ethnic minorities, and in some cases overpredict it, which argues against simple test bias as the explanation for score gaps. Yet the picture is inconsistent, with some studies finding tests less predictive of training performance for Black recruits, and the authors note that most criterion studies assume no bias in the criterion itself, typically subjective supervisor ratings.

The most heavily researched contextual factor is stereotype threat, the phenomenon whereby awareness of a negative stereotype about one’s group impairs performance on the stereotyped task. Seventy-five papers met the review’s criteria, with robust findings that threat impacts performance by gender, ethnicity and socioeconomic status, and worse effects for people holding multiple threatened identities. The two best-supported mechanisms are the depletion of working memory and executive resources and the misinterpretation of anxious arousal. Encouragingly, a range of interventions work: simply stating that a test shows no group differences, presenting a competent role model from the threatened group, teaching people about stereotype threat itself, mindfulness exercises and self-affirmation have all disrupted the effect. But the review flags a critical weakness: very few studies were conducted on high-stakes testing, and field studies such as one by Gillespie and colleagues struggled to replicate laboratory-sized effects in real occupational settings, suggesting experimental conditions may inflate the apparent magnitude of threat.

Experience and self-confidence emerged as quietly powerful factors. Multiple studies found that training, practice tests and practice items improved spatial performance more for women than men, in some cases eliminating the gender gap entirely, and playing action video games reduced the gap with women improving more than men. Practice and training also improved Black test-takers’ performance more than White test-takers’. Self-confidence mediated the relationship between gender and spatial performance, and in a particularly compelling experiment, Estes and Felker showed that manipulating confidence directly increased women’s performance on a spatial task. The authors suggest that effective test strategies may themselves be developed through domain experience, meaning that greater exposure to test content and more elaborate instructions in selection testing could reduce group differences without sacrificing validity.

The review is candid about the limitations of the evidence base it synthesised. Only 14 percent of studies met the criteria for ecological validity, meaning they were based on or closely resembled genuine high-stakes test use. Seventy-one percent relied on majority or exclusively student samples, reaching 94 percent in anxiety research and 91 percent in attitudes research, and students differ systematically from working populations in age, education, socioeconomic background and motivation. Fifty-five percent of studies used exclusively US samples, and only four percent drew participants exclusively from outside North America and Europe. Risk-of-bias appraisal using the Mixed Methods Appraisal Tool revealed widespread failure to control for confounders in non-randomised studies, and only 44 of the 225 papers examined multiple factors at once, leaving the interactions between explanations, for example between stereotype threat, anxiety and self-confidence, largely unmapped.

The practical upshot is a roadmap for fairer hiring. The authors recommend that test developers and users work to mitigate stereotype threat, address advantages gained through domain experience and exposure, and reduce differences in self-confidence, while also exploring state anxiety and continuing to screen for differential item functioning even though its overall impact is small. Because organisations lack highly valid selection methods that show no group differences, and because cognitive tests remain the best available predictor of performance, the diversity-validity dilemma will not be solved by abandoning these tests. Instead, the review suggests, progress lies in recognising that some of the gap is an artifact of the testing process itself, and that careful changes to how tests are designed, framed and administered can chip away at it, even as the deeper biological and environmental roots of group differences remind us that no test redesign will make them disappear entirely.

Subject of Research: Explanatory factors for demographic group mean score differences on cognitive ability tests in employee selection

Article Title: Factors Related to Mean Score Group Differences on Cognitive Ability Tests: A Systematic Review

Article References: Cuppello, S., Zibarras, L. D., & Corr, P. J. (2025). Factors Related to Mean Score Group Differences on Cognitive Ability Tests: A Systematic Review. Trends in Psychology. https://doi.org/10.1007/s43076-025-00504-5

Image Credits: AI Generated

DOI: 10.1007/s43076-025-00504-5

Keywords: cognitive ability tests, employee selection, stereotype threat, psychometrics, group differences, test bias, spatial ability, self-confidence, socioeconomic status, measurement invariance, test anxiety, systematic review

Cite Scienmag News

Glenn Wilkins. (October 2, 2026). Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests. Scienmag. https://scienmag.com/why-test-scores-diverge-landmark-review-maps-the-roots-of-group-gaps-on-cognitive-tests/

Glenn Wilkins. "Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests." Scienmag, 2 October 2026, https://scienmag.com/why-test-scores-diverge-landmark-review-maps-the-roots-of-group-gaps-on-cognitive-tests/. Accessed 2 October 2026.

Glenn Wilkins. "Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests." Scienmag. October 2, 2026. https://scienmag.com/why-test-scores-diverge-landmark-review-maps-the-roots-of-group-gaps-on-cognitive-tests/

Tags: Cognitive ability test score disparitiescognitive ability testsdemographic group differences in standardized testingemployee selectionenvironmental vs biological factors in cognitive testingfactors contributing to persistent test score gapsgroup differencesimpact of test design on racial score gapsimplications for fair employee selection practicesinfluence of domain experience and self-confidence on test outcomesmeasurement invariancemeta-analysis of demographic differences in standardized testsoccupational testing and racial performance gapspsychometricsroots of group gaps in cognitive assessmentsself-confidencesocioeconomic statusspatial abilitystereotype threatstereotype threat in testing environmentssystematic reviewsystematic review of cognitive test score divergencetest anxietytest bias
Share26Tweet16
Previous Post

Hair Cortisol Reveals Hidden Chronic Stress Hormone Exposure in Subtle Adrenal Disorder

Next Post

Digitalization Lowers Emissions in Emerging Economies, but Carbon Lock-In Slows the Fix

Related Posts

Psychiatry Residents Lack Confidence Reading Brain Scans, Survey Finds
Psychology & Psychiatry

Psychiatry Residents Lack Confidence Reading Brain Scans, Survey Finds

October 2, 2026
Parks Cut Kids’ Screen Time, but Sidewalks Tell a Different Story for Boys and Girls
Psychology & Psychiatry

Parks Cut Kids’ Screen Time, but Sidewalks Tell a Different Story for Boys and Girls

October 2, 2026
Why Psychologists Should Stop Treating Percentages Like Normal Data
Psychology & Psychiatry

Why Psychologists Should Stop Treating Percentages Like Normal Data

October 2, 2026
Too Many Likes, Not Too Few: Online Popularity Shifts Moral Choices
Psychology & Psychiatry

Too Many Likes, Not Too Few: Online Popularity Shifts Moral Choices

October 2, 2026
Good Deeds Backfire: How Friends’ Green Habits Quietly Erode Our Own
Psychology & Psychiatry

Good Deeds Backfire: How Friends’ Green Habits Quietly Erode Our Own

October 2, 2026
Simple Family Routines Boost Preschoolers’ Brain Skills and Ease Parental Stress
Psychology & Psychiatry

Simple Family Routines Boost Preschoolers’ Brain Skills and Ease Parental Stress

October 2, 2026
Next Post
Digitalization Lowers Emissions in Emerging Economies, but Carbon Lock-In Slows the Fix

Digitalization Lowers Emissions in Emerging Economies, but Carbon Lock-In Slows the Fix

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Two-Drug Combo Supercharges Liver Cancer Therapy in Preclinical Study
  • Parasite Adhesion Receptor Emerges as Drug Target in Schistosomiasis Study
  • Digitalization Lowers Emissions in Emerging Economies, but Carbon Lock-In Slows the Fix
  • Why Test Scores Diverge: Landmark Review Maps the Roots of Group Gaps on Cognitive Tests

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading