Monday, September 7, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Science Education

Estimating sampling variance in multilevel models under complex survey designs

September 7, 2026
in Science Education
Courtney Benton
By Courtney Benton Scienmag Editorial Profile - Science and Technology Policy
Reading Time: 6 mins read
0
Estimating sampling variance in multilevel models under complex survey designs

Estimating sampling variance in multilevel models under complex survey designs

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every year, researchers around the world download huge datasets from international education surveys such as TIMSS and PISA, fit multilevel models to the nested data, and publish conclusions about students, schools, and education systems. A new open-access study now warns that many of those analyses may rest on shaky statistical ground—not because the models are wrong, but because the standard errors behind them are being estimated incorrectly. The research, published in Large-scale Assessments in Education, provides the first systematic comparison of two competing approaches for computing sampling variances in multilevel models under complex sample designs, and it delivers both reassurance and a caution: when implemented correctly, the two approaches agree almost perfectly, but common misapplications can quietly inflate confidence intervals or, worse, bias the point estimates themselves.

The study, conducted by Xiaying Zheng of the American Institutes for Research and Shenghai Dai and Antranik Kirakosian of Washington State University, addresses a persistent practical puzzle in survey statistics. Large-scale assessments almost never use simple random sampling. Instead, they rely on multistage designs in which schools are first sorted into explicit strata—such as state, poverty level, or school type—then sorted again by “implicit” stratifying variables like geographic location or socioeconomic composition, and finally selected with probabilities proportional to their size. Students are then sampled within the chosen schools, sometimes with deliberate oversampling of particular groups, as Australia did with Indigenous students in TIMSS 2015. This architecture saves money and guarantees representation of key subgroups, but it breaks the assumptions of conventional variance estimation. Clustering inflates variances, stratification often shrinks them, and unequal selection probabilities distort everything in between.

Multilevel models have become the method of choice for analyzing such hierarchical data because they explicitly model students nested within schools and allow researchers to make inferences at both levels. But a multilevel model alone handles only the clustering; it does not automatically account for stratification or unequal weighting. To do that, analysts typically turn to the “sandwich” estimator, a design-based variance formula embedded as the default in widely used software such as Mplus and Stata. The sandwich estimator takes its name from its structure, in which a variance-covariance matrix reflecting the sample design is wedged between two copies of the inverse Fisher information matrix. Its middle layer is built by summing the outer products of log-likelihood gradient contributions from every cluster, aggregated within strata—and that is precisely where the trouble begins.

The crux of the problem, the authors explain, is identifying the correct stratification variable. Analysts might assume they should simply cross all the explicit and implicit stratifying variables from the sampling stage, but this quickly becomes intractable when there are many variables, and it collapses entirely when an implicit stratifier is continuous. Others include only the explicit strata, silently discarding the information carried by implicit sorting. The paper’s key insight is that the correct variable already exists in the data files, hiding in plain sight: it is the “variance strata” indicator created to generate replicate weights. Because schools are serpentine-sorted by implicit stratifiers and then sampled systematically, adjacent pairs of selected schools capture the effects of both explicit and implicit stratification without any manual crossing of variables. In TIMSS these variance strata are called “Jackknife sampling zones”; in PISA they appear as the “randomized final variance stratum.” Whatever the label, the authors argue, secondary analysts should hunt down this variable in the technical documentation rather than improvising from the sampling strata.

The alternative to the sandwich is replication, the approach officially recommended by most large-scale survey programs and used to produce the replicate weights shipped with TIMSS, PISA, and NAEP data. In the Jackknife Repeated Replication scheme used by TIMSS, adjacent pairs of schools form variance strata; each school is dropped in turn while its partner’s weight is doubled, generating 160 sets of replicate weights from 75 strata. PISA instead uses a balanced repeated replication variant, Fay’s method, in which one unit in each stratum has its weight inflated by a factor of 1.5 and the other deflated by 0.5, with a Hadamard matrix orchestrating the pattern across roughly 80 replicates. The sampling variance is then approximated from the spread of parameter estimates across these replicate runs. Despite its prominence in single-level analysis, replication has been widely assumed to be inapplicable to multilevel models, and no major software package—HLM, Stata, SAS, or Mplus—currently supports it for that purpose. Zheng and colleagues find no theoretical reason for this exclusion, and their work demonstrates that, with careful handling of disaggregated and rescaled weights, replication works just as well for two-level models.

To test both approaches rigorously, the team built a simulated population of 32,000 schools with realistic size distributions and generated student outcomes from a random-intercept-and-slope model with an intraclass correlation of 0.25, mirroring typical values in U.S. education. From this population they drew 5,000 samples under a two-stage design closely modeled on TIMSS: 160 schools selected with probabilities proportional to size along a sorted list, with explicit stratification at the school level and informative oversampling at the student level, in which students in the lowest quartile of a stratifying variable were sampled at twice the probability of their peers. Each sample was then analyzed under eight conditions spanning correct and incorrect implementations of the sandwich estimator, the Jackknife method, and the balanced repeated replication method.

The results were striking in their symmetry. When correctly specified, the sandwich estimator, the Jackknife method, and balanced repeated replication all produced confidence interval coverage rates essentially at the nominal 95 percent level, with ratios of average standard errors to empirical standard deviations hovering near the theoretical value of one. The practical message is liberating: researchers can confidently use either method depending on what their software and data allow. The misapplication results were equally instructive. When the sandwich estimator was run with only the explicit strata, or with no stratification information at all, standard errors were consistently overestimated, producing confidence intervals that were unnecessarily wide and statistical power that was silently squandered. In an era when null results can derail research programs, throwing away precision through misspecified strata is not a harmless error.

The second set of simulations tackled a subtler and more dangerous pitfall: ignoring informative level-1 weights. When the differential sampling of students within schools was disregarded, the point estimates themselves went astray, with substantial bias appearing in the level-2 intercept, the level-1 regression coefficient, and the residual variance, along with badly miscalibrated confidence intervals. This finding undercuts recent suggestions that level-1 weights can simply be dropped for convenience; the authors stress that such simplification is legitimate only when within-school sampling is completely random. They also confirm that level-1 weights should be rescaled—either to sum to the cluster sample size or to the effective cluster sample size—before entering the model, while level-2 weight rescaling is technically optional for two-level models but still recommended for correct pseudo-likelihood and fit indices. The good news for practitioners is asymmetry: including level-1 weights when they are uninformative does no harm, whereas excluding them when they are informative can bias results.

To show what this means in the real world, the authors turned to the Australian fourth-grade mathematics data from TIMSS 2015, comprising 6,057 students in 287 schools, and modeled the relationship between mathematics achievement and students’ confidence in mathematics. As in the simulations, conditions that omitted the informative level-1 weights—necessary here because Indigenous students were oversampled—produced noticeably different point estimates for both fixed and random effects. Conditions that used only explicit stratification (state or territory) or no stratification at all inflated the standard errors of the school-level regression coefficient and the random-intercept variance. When the sandwich estimator and the Jackknife method were both implemented correctly, their standard errors were generally comparable, though not identical, with the Jackknife yielding slightly larger standard errors for some parameters and slightly smaller for others. The authors combined results across the five plausible-value mathematics scores using Rubin’s method and made their data and code freely available through Zenodo, lowering the barrier for analysts who want to replicate the correct procedures.

The study closes with a call that is likely to resonate across the survey research community: software developers should build replication methods for multilevel models into the packages researchers actually use. Until then, the authors recommend that analysts consult each survey’s sampling and weighting documentation to locate the variance strata variable, apply disaggregated and properly rescaled level-specific weights, and align the treatment of replicate weights with the treatment of final weights. The work also acknowledges its limits—symmetric confidence intervals remain problematic for variance and covariance components whose sampling distributions are asymmetric, and the simulations captured only a slice of possible design conditions—but as a first systematic demonstration that replication is a fully valid alternative for multilevel variance estimation, the paper fills a genuine gap. For the thousands of secondary analysts who open TIMSS and PISA files each year, the guidance is simple to state: identify the Jackknife zones, respect the level-1 weights, and when the sandwich makes you hesitate, replicate.

Subject of Research: Sampling variance estimation for multilevel models under complex sample designs in large-scale surveys

Subject of Research: Science Education

Article Title: When the sandwich makes you hesitate, replicate: on sampling variance estimation of multilevel models under complex sample design

Article References: Zheng, X., Dai, S., & Kirakosian, A. (2026). When the sandwich makes you hesitate, replicate: on sampling variance estimation of multilevel models under complex sample design. Large-scale Assessments in Education, 14(1), Article 14. https://doi.org/10.1186/s40536-026-00285-y

Image Credits: AI Generated

DOI: 10.1186/s40536-026-00285-y

Keywords: complex sample design, variance estimation, multilevel model, sandwich estimator, replication method, large-scale assessments, sample weights, stratification, Jackknife, balanced repeated replication, TIMSS, standard errors

Cite Scienmag News

Courtney Benton. (September 7, 2026). Estimating sampling variance in multilevel models under complex survey designs. Scienmag. https://scienmag.com/estimating-sampling-variance-in-multilevel-models-under-complex-survey-designs/

Courtney Benton. "Estimating sampling variance in multilevel models under complex survey designs." Scienmag, 7 September 2026, https://scienmag.com/estimating-sampling-variance-in-multilevel-models-under-complex-survey-designs/. Accessed 7 September 2026.

Courtney Benton. "Estimating sampling variance in multilevel models under complex survey designs." Scienmag. September 7, 2026. https://scienmag.com/estimating-sampling-variance-in-multilevel-models-under-complex-survey-designs/

Tags: accuracy of standard error estimation in nested dataaccurate standard error estimation in nested databest practices for multilevel model variance estimationbias and confidence interval accuracy in multilevel modelingcomparison of variance estimation methodscomplex survey design analysiseffects of misapplying variance estimation techniqueseffects of stratification and clustering on variance estimatesimpact of misapplied variance estimation methodsimpact of sampling design on statistical inferenceimplications for TIMSS and PISA data analysisimplications of variance misestimation on educational policylarge-scale educational assessment analysismultilevel modeling in education surveysmultilevel models in complex survey designsmultistage sampling in education surveysmultistage sampling in large-scale assessmentsopen-access research on survey statisticssampling bias correction in PISA and TIMSS datastatistical best practices in complex survey datastatistical reliability in large-scalesurvey design considerations for educational researchsurvey sampling variance estimationsurvey sampling variance estimation in multilevel models
Share26Tweet16
Previous Post

India’s low-fertility states face limits and prospects in fertility recovery

Next Post

Local food environments drive small fish consumption in Ugandan schoolchildren

Related Posts

National survey examines faculty views on incentives and competency integration in Saudi health education
Science Education

National survey examines faculty views on incentives and competency integration in Saudi health education

September 7, 2026
FAU wins $400,000 NSF grant to boost STEM workforce readiness
Science Education

FAU wins $400,000 NSF grant to boost STEM workforce readiness

September 7, 2026
Model-based reasoning in STEM education: systematic review of literature
Science Education

Model-based reasoning in STEM education: systematic review of literature

September 6, 2026
Women bear higher economic burden of chronic disease costs in Mexico
Science Education

Women bear higher economic burden of chronic disease costs in Mexico

September 6, 2026
Teacher self efficacy and workplace motivation linked to multidimensional burnout
Science Education

Teacher self efficacy and workplace motivation linked to multidimensional burnout

September 6, 2026
Enjoyment of reading shows measurement invariance across grades, IRT study finds
Science Education

Enjoyment of reading shows measurement invariance across grades, IRT study finds

September 6, 2026
Next Post
Local food environments drive small fish consumption in Ugandan schoolchildren

Local food environments drive small fish consumption in Ugandan schoolchildren

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Local food environments drive small fish consumption in Ugandan schoolchildren
  • Estimating sampling variance in multilevel models under complex survey designs
  • India’s low-fertility states face limits and prospects in fertility recovery
  • Longer hours in Bahrain’s banks linked to mental health and productivity decline

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading