Online higher education has a quiet attrition problem. Across distance-learning institutions, dropout rates routinely fall between 30 and 50 percent, and most universities still discover who is at risk through retrospective surveys that arrive long after the student has disengaged — or already gone. A team of researchers at the University of Salamanca in Spain now argues that the fix is not a better prediction score, but a better prediction pipeline: one that forecasts risk early, explains itself in plain language, and audits itself for bias before it ever reaches an instructor’s dashboard. Their system, described in the International Journal of Data Science and Analytics, is released alongside BurnoutGuard, an open-source Moodle plugin that turns the framework into something a university could actually deploy.
The study rests on one of the largest public learning analytics resources available: the Open University Learning Analytics Dataset, or OULAD, which captures 32,593 student–course registrations and more than ten million interactions with a virtual learning environment. Of those registrations, 29,228 carry recorded platform activity and form the analytical sample. Rather than treating the dataset as a single prediction problem, the researchers split it by time. Their headline result is deliberately conservative: using only the data available during the first weeks of a course, and scoring only the students still enrolled at each horizon, the model reaches an area under the ROC curve — AUROC — of 0.718 at week four, after which performance plateaus rather than continuing to climb.
That plateau is itself a finding, and so is the way the team measured it. Many published dropout models report impressive numbers by quietly including students who have already formally withdrawn; such students are trivially separable from their peers, because their records simply stop. The Salamanca group shows that leaving those students in the risk set would inflate the apparent week-eight AUROC to 0.78 and manufacture a steady upward trend that does not actually exist. By restricting the sample to students still enrolled at each prediction horizon, they present a number that reflects what an early warning system would genuinely face in practice: a harder, noisier, and far more useful problem.
For context, the researchers also trained a model on the full course, which reaches an AUROC of 0.936. They are explicit that this figure is a complete-information reference rather than a deployment estimate. A leakage analysis attributes the gap between the early and full-course models not to a handful of suspicious, outcome-adjacent variables, but to features aggregated over the entire course presentation — information that simply does not exist in week four. The distinction matters for anyone shopping for dropout prediction tools: a vendor quoting 0.93 accuracy may be selling a system that only works once the semester is effectively over.
Architecturally, the early-horizon predictor is a stacking ensemble of tree-based models — the family that includes random forests and gradient-boosted methods such as XGBoost, LightGBM and CatBoost — designed specifically for temporally truncated data. Stacking combines the outputs of several base classifiers through a higher-level model, letting the system hedge between learners that excel on different patterns of early engagement. Tree ensembles are a pragmatic choice for tabular behavioural data of this kind: they handle mixed feature types well, tolerate missingness, and, crucially for this project, are compatible with modern explanation methods that can attribute each prediction to individual input features.
That explainability layer is built on SHAP — SHapley Additive exPlanations — a technique rooted in cooperative game theory that distributes credit for a prediction among the features that produced it. Applied globally, SHAP delivers a strikingly interpretable picture of why online students drop out. The single most important signal is the course day of the student’s last platform access, accounting for 33.6 percent of global feature importance, followed by the number of assessments submitted at 11.0 percent. Perhaps most counterintuitive for educators: prior academic grades matter very little for withdrawal. Disengagement, not weak transcripts, is the early fingerprint of risk — a student who has not logged in for ten days is a far stronger alarm than one who arrived with mediocre qualifications.
But the paper’s most methodologically consequential section may be its fairness audit. Predictive models in education have documented potential to encode bias, systematically over- or under-flagging students by gender, age or socioeconomic status. The team evaluated demographic parity, false-positive and false-negative rate parity, and per-group calibration across demographic groups. For gender and socioeconomic status the gaps are small; for age they are appreciable. And in a result that should give pause to institutions rushing to ‘de-bias’ their models, the researchers show that forcing demographic parity in the age case actually worsens equal opportunity — the guarantee that qualified students have equal chances of a positive flag regardless of group membership. Fairness criteria, they demonstrate, can conflict, and choosing among them is a policy decision, not a technical afterthought.
Timing changes the fairness picture too. When the audit is run at the deployment horizons rather than at the end of the course, two problems emerge that end-of-course evaluations would miss entirely. First, the early risk scores are miscalibrated — a score of 0.7 does not mean a 70 percent chance of dropout — until a post-hoc calibrator is fitted to the early-horizon outputs. Second, the gender gap at week four is roughly two and a half times the full-course figure, and it persists under the capacity-based operating point the team evaluated. In other words, a system that looks acceptably fair in a retrospective, full-course evaluation can be meaningfully less equitable in the first weeks, exactly when institutions would act on its warnings.
The practical deliverable is BurnoutGuard, a Moodle plugin released under the GNU GPL v3 licence, with versioned releases archived on Zenodo and all experiments run with a fixed random seed for reproducibility. The plugin implements the deployment framework end to end, including the calibration step, and — critically for adoption — explains each individual prediction in natural language. Instead of an opaque risk percentage, an instructor sees which behaviours drove the flag: recency of last access, missed assessments, declining activity. A small pilot at the authors’ institution, in which eight instructors used the tool and gave informal feedback, served as a technical proof of concept; no student data beyond what the institution already held were collected.
The broader significance of the work lies less in any single AUROC figure than in the template it offers. Early warning systems in education have been around for over a decade — Purdue’s Course Signals being the best-known example — but most published models optimise for end-of-course accuracy, ignore calibration, and treat fairness as optional. By making early prediction the primary result, auditing fairness at the horizons where interventions actually happen, and shipping the whole stack as open source, the Salamanca team has produced what amounts to a reproducible checklist for equity-aware early warning. For the millions of students now studying at a distance, the difference between a week-four alert and a semester-end post-mortem may be the difference between a phone call from a tutor and a withdrawal form. The model cannot make students stay. But it can, at last, tell institutions who needs the call — and show its reasoning when it does.
Subject of Research: An explainable and fair machine learning early warning system for predicting student dropout in online higher education
Article Title: Explainable and fair early warning system for academic risk prediction in online higher education
Article References: Sánchez-Pravos, L., Aucancela Guaman, M. A., González-Briones, A., & Chamoso, P. (2026). Explainable and fair early warning system for academic risk prediction in online higher education. International Journal of Data Science and Analytics, 22(1), Article 315. https://doi.org/10.1007/s41060-026-01318-z
Image Credits: AI Generated
DOI: 10.1007/s41060-026-01318-z
Keywords: learning analytics, student dropout prediction, explainable AI, SHAP, algorithmic fairness, early warning systems, online higher education, machine learning, Moodle, Open University Learning Analytics Dataset, calibration, educational data mining
Cite Scienmag News
Blake Davidson. (October 2, 2026). AI Flags Failing Online Students in Weeks — and Shows Its Work. Scienmag. https://scienmag.com/ai-flags-failing-online-students-in-weeks-and-shows-its-work/
Blake Davidson. "AI Flags Failing Online Students in Weeks — and Shows Its Work." Scienmag, 2 October 2026, https://scienmag.com/ai-flags-failing-online-students-in-weeks-and-shows-its-work/. Accessed 2 October 2026.
Blake Davidson. "AI Flags Failing Online Students in Weeks — and Shows Its Work." Scienmag. October 2, 2026. https://scienmag.com/ai-flags-failing-online-students-in-weeks-and-shows-its-work/

