Missing data is the quiet saboteur of modern research. Whether a participant skips a questionnaire item, drops out of a longitudinal study, or a sensor fails mid-experiment, the gaps left behind can distort estimates, inflate uncertainty, and drain statistical power in ways that no amount of clever analysis can fully repair. Now, a new study published in Behavior Research Methods proposes an unexpected remedy: instead of treating data integration and missing data as separate problems, researchers can deliberately borrow cases from entirely independent studies to shore up their own incomplete datasets. The approach, developed by Po-Yi Chen of National Taiwan Normal University, reframes what it means to combine data across studies, turning integration itself into a weapon against missingness.
The core idea builds on a framework known as integrative data analysis, or IDA, which has traditionally been used to pool information from multiple studies in order to stabilize estimates and increase sample sizes. Past work on IDA has generally treated missing data as an obstacle to be overcome before datasets can be merged, a nuisance that complicates the harmonization of measures and models. Chen’s study flips that logic on its head. Rather than viewing integration as something that must survive missing data, the new research shows that integration can actively limit the damage that missing data inflicts on the analyses researchers actually care about.
The key concept is what Chen calls auxiliary cases. These are cases drawn from other independent samples or existing studies that are not part of the researcher’s target group of interest, but that carry sufficient data on the target variables to reduce the impact of missingness on the focal analyses. In other words, if a sleep researcher is studying the relationship between daytime sleepiness and cognitive performance in a clinical sample, but many participants are missing scores on some measures, cases from a completely different study that measured the same constructs can be brought in to help. These borrowed cases do not change the research question or contaminate the target population; they simply add statistical information that stabilizes the estimation process.
The mechanism for including auxiliary cases is a multiple-group model, a structural equation modeling technique in which cases from different sources are treated as belonging to different groups within a single analysis. The auxiliary cases form their own group, and the model imposes what are known as cross-group measurement invariance constraints on the measurement portion of the model. This means that the way latent constructs are measured, including factor loadings and related measurement parameters, is forced to be the same across the target group and the auxiliary group, ensuring that the borrowed cases are genuinely measuring the same things. Crucially, however, the focal structural parameters, such as the regression paths and structural relationships that constitute the actual research question, are allowed to be freely estimated for the target group. The auxiliary group thus contributes information about the measurement apparatus without dictating the substantive findings.
To test whether this strategy actually works, Chen conducted a series of simulation studies, the standard methodological tool for evaluating statistical techniques under known conditions. The simulations compared scenarios with complete data against scenarios with missing data, varying conditions that researchers commonly face. The results were striking in one direction: the benefits of including auxiliary cases were substantially larger under missing data conditions than under complete data conditions. When data were complete, adding auxiliary cases offered only modest gains, consistent with earlier findings that integrative data analysis can stabilize estimates. But when data were missing, the borrowed cases delivered meaningful improvements in both statistical power and the efficiency of the estimates, meaning narrower standard errors and a better chance of detecting true effects that would otherwise be drowned in uncertainty.
That asymmetry makes intuitive sense once the underlying statistics are considered. Modern missing data techniques, such as full information maximum likelihood estimation, extract as much information as possible from incomplete observations, but they cannot conjure information that simply is not there. Every missing value represents lost Fisher information, and the efficiency of parameter estimates degrades accordingly. Auxiliary cases inject fresh information about the joint distribution of the target variables, effectively replenishing what the missingness drained away. Under complete data, there is little to replenish, so the auxiliary cases have little to offer. Under missingness, they arrive precisely when the analysis is most starved of information.
Yet the study is equally clear about the limits of the approach, and its warnings are as important as its promises. If the cross-group invariance constraints imposed on the measurement model are improper, that is, if the measures do not actually function equivalently across the target and auxiliary samples, the inclusion of auxiliary cases can backfire, introducing additional bias into the estimates and even reducing statistical power. The problem is compounded when inter-dataset heterogeneities, genuine differences between the datasets in populations, procedures, or contexts, are not accounted for in the analyses. Under these conditions, the borrowed cases are not contributing neutral information about the measurement apparatus; they are smuggling in assumptions that are false, and the model’s estimates absorb the error.
Notably, these risks were found to be especially pronounced when full information maximum likelihood was used as the estimation method. FIML is one of the most widely recommended approaches for handling missing data in structural equation models, prized for its efficiency and its ability to use all available information without deleting cases. But the same property that makes FIML powerful, its reliance on the model’s assumptions to fill in the informational gaps, also makes it sensitive to violations. When invariance constraints are wrong or heterogeneity is ignored, FIML propagates those errors through the entire model, and the auxiliary cases become a liability rather than an asset. The lesson for practitioners is that the auxiliary-case strategy is not a plug-and-play fix; it demands careful verification that the measures are truly comparable across datasets and that between-study differences are explicitly modeled.
To demonstrate the approach on real data rather than simulated numbers, Chen provided an empirical example based on two independent experimental datasets drawn from the National Sleep Research Resource, a publicly accessible repository supported by the National Heart, Lung, and Blood Institute. The two datasets, the Apnea Positive Pressure Long-term Efficacy Study, known as APPLES, and the Best Apnea Interventions in Research study, or BestAIR, both included the Epworth Sleepiness Scale, a widely used measure of daytime sleepiness whose psychometric properties have been extensively validated. Treating one dataset as the target sample and the other as a source of auxiliary cases, the example illustrated how the multiple-group framework operates in practice, complete with the invariance testing steps needed to justify the cross-group constraints. The analysis code for both the simulations and the empirical example has been made available through the Open Science Framework, and the datasets themselves are freely available upon request from the repository, giving other researchers a concrete template to follow.
The broader significance of this work lies in how it repositions two familiar tools on the methodological landscape. Data integration has long been framed as a way to answer bigger questions by combining evidence, and missing data handling has been framed as a way to salvage what a flawed dataset can still tell us. Chen’s study shows that these framings are not merely parallel but intertwined: the act of integrating data can itself be a missing data intervention, and the choice of which outside cases to borrow becomes a design decision with direct consequences for power, bias, and efficiency. As open data repositories grow and secondary analysis becomes ever more central to the research enterprise, the pool of potential auxiliary cases will only deepen. The approach will not suit every study, since it requires overlapping measures, defensible invariance, and honest attention to heterogeneity across sources. But for researchers staring down a dataset riddled with gaps, the message is genuinely encouraging: the remedy for missing data may already exist in someone else’s study, waiting to be borrowed, constrained carefully, and put to work.
Subject of Research: Using auxiliary cases from independent samples in multiple-group models to mitigate the impact of missing data on structural equation analyses
Article Title: Including auxiliary cases to address missing data issues through multiple-group models
Article References: Chen, P.-Y. (2026). Including auxiliary cases to address missing data issues through multiple-group models. Behavior Research Methods, 58(11), Article 304. https://doi.org/10.3758/s13428-026-03180-0
Image Credits: AI Generated
DOI: 10.3758/s13428-026-03180-0
Keywords: missing data, auxiliary cases, multiple-group models, integrative data analysis, measurement invariance, full information maximum likelihood, structural equation modeling, statistical power, data integration, simulation study, Behavior Research Methods, sleep research
Cite Scienmag News
Glenn Wilkins. (September 30, 2026). Borrowed Cases, Better Statistics: New Method Fights Missing Data by Pooling Outside Samples. Scienmag. https://scienmag.com/borrowed-cases-better-statistics-new-method-fights-missing-data-by-pooling-outside-samples/
Glenn Wilkins. "Borrowed Cases, Better Statistics: New Method Fights Missing Data by Pooling Outside Samples." Scienmag, 30 September 2026, https://scienmag.com/borrowed-cases-better-statistics-new-method-fights-missing-data-by-pooling-outside-samples/. Accessed 30 September 2026.
Glenn Wilkins. "Borrowed Cases, Better Statistics: New Method Fights Missing Data by Pooling Outside Samples." Scienmag. September 30, 2026. https://scienmag.com/borrowed-cases-better-statistics-new-method-fights-missing-data-by-pooling-outside-samples/

