When schools across Germany closed their doors in the spring of 2020, roughly 80 percent of children and young people worldwide were suddenly required to learn from home, and teachers were forced to reinvent their profession overnight. The educational consequences have been measured repeatedly since then: mean achievement scores dropped, and the share of students failing to reach minimum academic standards rose to unprecedented levels. What has remained far less clear is precisely who these low-achieving students are in the aftermath of the pandemic, and whether the extraordinary circumstances of distance learning created entirely new risk profiles or simply deepened old inequalities. A new study published in the journal Large-scale Assessments in Education tackles that question with an unusually powerful analytical apparatus, applying educational data mining to one of the largest national assessments ever conducted in Germany, and its findings offer both reassurance and warning.
The research team, led by Kristoph Schumann, Karoline A. Sachse and Stefan Schipolowski of the Institute for Educational Quality Improvement at Humboldt-Universität zu Berlin, together with Rebecca Schneider of the University of Münster, analyzed data from the IQB Trends in Student Achievement Study 2021, a nationally representative assessment of fourth-graders at the end of primary school. The dataset included 26,844 students in 1,464 schools, of whom 24,511 were analyzed for reading proficiency and 24,500 for mathematics. Data collection took place between April and August 2021, after more than a year of disrupted schooling, and combined standardized achievement tests of up to 160 minutes per student, a test of basic cognitive ability, and questionnaires completed by students, parents and teachers. From a pool of roughly 150 candidate variables, the researchers distilled 103 predictors spanning demographics, socio-economic status, motivational characteristics, home learning environments, and detailed accounts of how distance learning actually functioned during the closures.
The outcome variable was deliberately stringent. Rather than treating achievement as a continuous score, the team defined low achievers in binary terms: students who failed to meet the minimum standards set out in the German National Educational Standards for primary education, the threshold considered a prerequisite for successful learning in secondary school. About 80 percent of students reached the minimum standard in reading and 82 percent in mathematics, leaving a substantial minority of roughly one in five children below the critical line in at least one domain. Proficiency estimates were based on Weighted Likelihood Estimation rather than plausible values, a technical decision the authors justify by noting that their machine learning approach was designed to surface complex, potentially nonlinear interactions between background variables that a simpler background model might obscure.
The analytical centerpiece of the study is a method called Prediction Rule Ensembles, or PRE, an innovative technique in educational data mining that balances interpretability and predictive accuracy. The approach works in three steps: a large ensemble of decision trees is generated, rules are extracted from every node of those trees, and all rules together with linear predictors are then fed into a lasso-regularized logistic regression, which prunes the ensemble down to a sparse final set of terms. Each rule functions as a dummy variable that takes the value one only when all of its conditions are met, which means rules can capture interactions, such as the joint effect of few books at home and poor class participation in distance learning, that conventional regression would struggle to express. Hyperparameters, including a maximum tree depth of two, a learning rate of 0.01 and a sample fraction of 0.75, were tuned via ten-fold cross-validation, with upsampling applied within each fold to compensate for the imbalanced outcome.
The headline results confirm that the old risks have not gone anywhere. Two of the four most important variables characterizing low achievers in both reading and mathematics were socio-economic status, operationalized by the Highest International Socio-Economic Index of Occupational Status, and cultural capital at home, measured by the number of books in the household. The other two were subject-related anxiety variables, and here the study delivered a striking surprise: anxiety in mathematics showed high importance not only for mathematics but also for reading, and its importance was comparable to that of the classic social disparity indicators, exceeding many other predictors in the models. The authors argue that anxiety should therefore not be treated as a marginal or secondary factor but as a central variable closely linked to low achievement during this period.
Class-level composition mattered as much as individual circumstances. The class-averaged socio-economic index proved more predictive than the individual student’s own score, and teacher ratings of class-level conditions were overrepresented among the most critical decision rules, underscoring, the authors say, the importance of incorporating classroom context into educational research. Further high-ranking predictors included academic self-concept, age, generational immigration status, parental occupational class, and cognitive activation in reading instruction, which mattered for mathematics as well. Gender behaved as expected: girls held an advantage in reading, while gender played no substantial role in mathematics. Not having attended pre-primary education was also associated with elevated risk.
Crucially, pandemic-specific variables did not disappear once the established factors were taken into account. Although COVID-19-related measures ranked below the structural social inequalities, they consistently appeared among the top twenty predictors in both models. The proportion of in-person teaching a class received and parents’ ratings of the technical resources available at home both contributed meaningfully to characterizing low achievers. Other distance-learning variables, such as parental reports of teacher feedback, parental ability to support their children, the quality of internet access at home, and whether instruction alternated within a single day, surfaced in the most important predictive rules. The message, according to the researchers, is that the pandemic’s fingerprint on educational risk is real, even if it did not outweigh the deeper inequalities that predate it.
Some of the study’s most interesting findings concern interactions. One rule indicated that the predicted risk of failing the reading standard increased by 0.22 on the logit scale for students with ten or fewer books at home in classes where the mathematics teacher rated the participation of all students in distance learning as rather or very poor. Pairwise partial dependence plots showed the predicted failure probability reaching 52 percent for this group, against 33 to 46 percent for all other students, with the effect of participation holding only in combination with low cultural capital. Another rule showed that high mathematics anxiety, above 2.67 on a four-point scale, raised the risk of failing the mathematics standard mainly in classes where the German teacher judged learning objectives to be poorly achieved. The authors suggest such patterns may hint at compensatory mechanisms: engaging distance learning might buffer disadvantages, while struggling classes amplify them, though they stress these hypotheses require causal testing.
Methodologically, the models performed respectably. On a held-out test set of 25 percent of the data, the reading model achieved sensitivity of 69.10 percent and specificity of 71.76 percent, while the mathematics model reached 69.87 and 78.27 percent respectively, yielding balanced accuracies of 70.43 and 74.19 percent. In repeated ten-fold cross-validation against a single decision tree, a lasso-regularized logistic regression and a random forest, the Prediction Rule Ensembles were significantly better than the single tree, statistically indistinguishable from the lasso regression, and either slightly worse than the random forest for reading, at the p < 0.01 level, or slightly, though not significantly, better for mathematics. The authors interpret this as a mixed verdict: the method delivers interpretable risk profiles at near-forest accuracy, but offers little raw performance gain over classic regression.
The team is careful about the limits of their claims. The analysis is exploratory and descriptive, not causal; the outcome distribution required upsampling, which can inflate performance metrics and heighten sensitivity to outliers; only the first of fifteen multiply imputed datasets could be used; and the data reflect a singular global crisis, limiting generalization to ordinary school years. Model instability in rule-based methods also means variable importance rankings could shift with different random seeds, though the authors report the main results remained comparable. The study was preregistered, and the researchers articulate explicit hypotheses for future confirmatory work, including the suggestion that the conventional threshold of 100 books for measuring cultural capital may be poorly suited to identifying students near the achievement floor, since cutpoints of 10 and 25 books proved more informative here.
For policymakers and educators, the practical upshot is twofold. First, the children who fell below minimum standards during the pandemic looked, in broad outline, much like the at-risk populations of previous decades: socio-economically disadvantaged, culturally less resourced, more often from immigrant families. Second, the machinery of remote schooling, internet access, technical equipment, teacher feedback, and genuine student participation, left a measurable imprint on who failed, often interacting with family background in ways that suggest targeted support could matter enormously. If future crises force schools to close again, the study argues, ensuring high-quality remote learning with the participation of all students, particularly those with the fewest resources at home, should be a priority area of focus. As mean proficiency in Germany hit a new low point in 2021, this data-driven portrait of who was left behind offers a foundation for directing limited educational resources to the students who need them most.
Cite Scienmag News
Kristina Jarvis. (September 9, 2026). Identifying low achievers in large-scale pandemic assessments using educational data mining. Scienmag. https://scienmag.com/identifying-low-achievers-in-large-scale-pandemic-assessments-using-educational-data-mining/
Kristina Jarvis. "Identifying low achievers in large-scale pandemic assessments using educational data mining." Scienmag, 9 September 2026, https://scienmag.com/identifying-low-achievers-in-large-scale-pandemic-assessments-using-educational-data-mining/. Accessed 9 September 2026.
Kristina Jarvis. "Identifying low achievers in large-scale pandemic assessments using educational data mining." Scienmag. September 9, 2026. https://scienmag.com/identifying-low-achievers-in-large-scale-pandemic-assessments-using-educational-data-mining/

