When a student’s internet connection drops in the middle of an online examination, the learning management system records something ominous: unanswered questions, a truncated session, sometimes a blank submission. To the grading pipeline, that trace looks identical to the trace left by a student who never studied at all. The result is that an infrastructure failure quietly becomes an academic verdict, and the student who lost power during a practical exam receives the same failing score as the student who simply did not prepare. A new study published in Discover Education argues that this conversion of technical failure into academic punishment is not only unfair but measurable, and that the evidence has been sitting in the logs all along.
The research, conducted by Samer Yaghi and Aiman Ahmed AbuSamra at the University College of Applied Sciences (UCAS) in Gaza, Palestine, analysed 675 examination instances from 226 undergraduate students across five sections and seven assessments of a Web Databases course. The setting was chosen deliberately. Gaza operates under compounded infrastructural constraint: degraded and intermittent internet capacity, severe resource scarcity, and an unreliable electricity supply that forces households onto improvised alternatives. Connectivity during an exam is therefore not a background assumption but a contingent condition that can change within a single sitting. The authors treat Gaza as a stress-test case: if infrastructure-induced disruption can be identified and validated under these conditions, the same log-based approach should work in any environment where connectivity cannot be assumed.
The methodological core of the study is a severity-graded definition of disruption built from a simple quantity: the blank ratio, the proportion of questions left unanswered in an attempt. A student who omits one question for academic reasons is a different phenomenon from a student whose session collapsed mid-exam, and an unweighted indicator would conflate them. The researchers therefore graded disruption in four tiers: no unanswered questions; minor disruption below a 0.50 blank ratio, likely incidental omission; substantial disruption between 0.50 and 1.0, indicating a probable mid-session interruption; and complete disconnection, where a full-length session produced no transmitted answers at all. Throughout the analysis, substantial disruption means a blank ratio of at least 0.50, a deliberately conservative threshold that trades sensitivity for specificity so that reported effects cannot be attributed to ordinary omissions.
The most striking result comes from a within-subject recovery analysis. Seven students showed substantial disruption in a final examination and later sat a stable supplementary examination under normal connectivity. Their scores rose from a mean of 10.0 percent under disruption to 98.6 percent under stable conditions, a mean within-subject gain of 88.6 percentage points. The Wilcoxon signed-rank test returned p = 0.016, and the matched-pairs rank-biserial correlation was 1.00, meaning every single pair moved in the same direction. A bootstrap confidence interval for the mean gain, computed with 10,000 resamples, ranged from 78.0 to 97.1 points. The authors are candid about the constraints: with seven pairs, the smallest attainable two-sided p-value is 0.0156, so the evidential weight rests on the consistency and magnitude of the effect rather than on a p-value that could not have been smaller.
Because institutional policy records the higher of two scores, a sceptic might argue that students simply prepared harder for the second sitting. To address this, the researchers embedded the disrupted exam between two unimpeded measurements of the same student. Of the seven disrupted students, five had at least one earlier examination with no disruption and a recorded grade. Those five had averaged 77.8 percent before the disruption, scored 10.0 percent during it, and 99.0 percent afterwards. The trajectory is a collapse and a return, not an ascent: recovery exceeded the pre-disruption baseline by only 21.3 points in a course where the supplementary theory exam has a median score of 100 percent, a pattern consistent with returning to a ceiling rather than with a genuine learning gain. The alternative account, that these students were underperforming at 77.8 percent, then advanced to 99.0 percent while passing through 10.0 percent in between for unrelated reasons, requires considerably more auxiliary assumptions than the infrastructure explanation.
Prevalence analysis revealed a second major finding: disruption is not evenly distributed across assessment formats. Practical examinations, which require sustained file uploads and continuous interaction with a server-side environment, showed disruption rates of 33.8 percent and 34.5 percent, while theory examinations, delivered as automatically graded multiple-choice papers, showed rates between just 1.5 and 4.4 percent. Wilson confidence intervals for the smaller practical cohorts were wide but remained well separated from the theory figures, indicating the gap is not an artifact of sample size. Three mechanisms plausibly contribute: practical exams run longer, presenting a larger temporal window for connectivity loss; a failure during file transmission destroys work already completed rather than merely delaying an answer; and multiple-choice papers can tolerate brief interruptions between item submissions. The practical implication is direct, bearing on how long practical sittings should run and whether they should be segmented into independently submitted parts.
Perhaps the most behaviourally distinctive pattern was the complete-disconnection signature. Thirty-nine instances occurred in which students remained in sessions averaging 70 minutes despite transmitting zero answers, spanning 35 unique students across four examinations. Plotted against session duration, these instances form a distinct band along the zero-answer axis. The authors qualify the signature carefully: recorded durations among these cases ranged from under two minutes to more than ninety, and short zero-answer sessions are equally consistent with a candidate who entered and withdrew. The behavioural argument applies most strongly to the sustained cases, those exceeding thirty minutes, which retain the majority of the set. A robustness check across the five sections, taught by three instructors using an identical examination instrument, found no significant differences in disruption rates, arguing against attributing the pattern to any individual instructor’s administration practices.
The study also positions itself against the existing institutional machinery. At UCAS, as at many institutions, students who encounter a failure contact their instructor, a technical support unit assists with access, and a formal Incomplete request can be filed and reviewed. The authors document five structural limits of this arrangement, the most consequential being that the burden of initiating, documenting, and advocating falls entirely on the student, at exactly the points where those with the weakest infrastructure and least confidence in claiming redress are filtered out. Empirically, the two mechanisms capture partly non-overlapping populations: among eighteen final-examination instances belonging to students who later sat a supplementary exam, six carried no blank-answer signature at all. The log-based measure is therefore proposed not as a substitute for adjudication but as a systematic input to it, ensuring the question is asked in every flagged case rather than only those a student manages to raise.
The findings also expose a blind spot in mainstream fairness-aware learning analytics. Formal criteria such as equalized odds require equal true-positive and false-positive rates across protected groups, but the inequity documented here originates in an external event during the assessment itself. There is no protected group to equalise across, no classifier, and no error rate, only a measurement taken under conditions that differed between candidates sitting the same paper. The authors argue that treating the examination log as a record of technical conditions, not solely of knowledge, is a prerequisite for fair assessment in infrastructure-constrained environments, and that a subset of traces traditionally read as disengagement or academic-integrity concerns in fact records a failed channel rather than a failing candidate.
A threshold sensitivity analysis strengthens the conclusion: repeating the primary comparison across every cut-point from a blank ratio of 0.30 to 0.70 identified the same seven students with unchanged statistics, so the 0.50 convention is not load-bearing. The authors acknowledge limitations, including the single-course, single-institution scope, the absence of direct network telemetry, and the small matched-pair sample. The anonymized dataset is openly available on Zenodo under a CC BY 4.0 licence with k-anonymity of at least five, and the complete analysis code is openly released, making the pipeline reproducible on standard Moodle gradebook exports alone. Future work will pursue multi-institution validation, triangulation against independently filed disruption reports, and real-time detection. For the millions of students sitting high-stakes digital examinations under unstable connectivity, the message is stark: the inequality was always there in the data, waiting to be counted.
Subject of Research: Fairness-aware learning analytics for detecting internet-induced inequality in online examinations
Article Title: Fairness aware learning analytics for identifying internet induced inequality in online examinations
Article References: Fairness aware learning analytics for identifying internet induced inequality in online examinations. (n.d.). https://doi.org/10.1007/s44217-026-02182-6
Image Credits: AI Generated
DOI: 10.1007/s44217-026-02182-6
Keywords: learning analytics, online examinations, digital inequality, educational fairness, Moodle, connectivity, assessment, Gaza, algorithmic fairness, educational data mining, digital divide, higher education
Cite Scienmag News
Courtney Benton. (September 26, 2026). When the Internet Fails, Grades Fail: New Analytics Expose Hidden Inequality in Online Exams. Scienmag. https://scienmag.com/when-the-internet-fails-grades-fail-new-analytics-expose-hidden-inequality-in-online-exams/
Courtney Benton. "When the Internet Fails, Grades Fail: New Analytics Expose Hidden Inequality in Online Exams." Scienmag, 26 September 2026, https://scienmag.com/when-the-internet-fails-grades-fail-new-analytics-expose-hidden-inequality-in-online-exams/. Accessed 26 September 2026.
Courtney Benton. "When the Internet Fails, Grades Fail: New Analytics Expose Hidden Inequality in Online Exams." Scienmag. September 26, 2026. https://scienmag.com/when-the-internet-fails-grades-fail-new-analytics-expose-hidden-inequality-in-online-exams/

