The Great Digital Test Switch: Global Exam Scores Survived the Move From Paper to Screens — With One Startling Exception
When the Trends in International Mathematics and Science Study—TIMSS, the four-yearly census of achievement that shapes curriculum debates and league tables for millions of schoolchildren—switched from paper booklets to computer screens, an uncomfortable question traveled with it into every exam room: were we still measuring the same thing? A new analysis published August 29, 2026, in the journal Large-scale Assessments in Education offers the most reassuring answer yet. Researchers who compared students tested on paper with students tested on screens in Ireland, Azerbaijan, Bahrain, and Flanders, the Dutch-speaking region of Belgium, found no significant difference in overall mathematics or science performance in any of the four education systems. The digital switchover, in short, did not move the needle. But buried beneath the averages was one striking anomaly: in Azerbaijan, boys scored markedly higher in science when the test appeared on a screen—an effect large enough to warrant a national investigation.
Run every four years, TIMSS is among the most consequential measuring instruments in education, and until recently every participant in every country sat the same kind of test: printed booklets, pencils, and hand-processed answer sheets. That changed in 2019, when about half of the participating countries migrated to a digital version delivered on computers, laptops, or tablets, with most of the remaining systems completing the migration in the 2023 cycle. The motivations were compelling. Digital delivery opens the door to interactive items and Problem-Solving and Inquiry (PSI) tasks that paper cannot replicate, promises more colorful and dynamic formats that hold children’s attention, reduces the volume of responses requiring human scoring, and eliminates the laborious data entry of paper materials. Crucially, the digital assessment was engineered to be comparable with its paper predecessors, so that the long-term trend lines on which education policymakers depend would remain interpretable across the transition.
Engineering intent, however, is not the same as psychometric certainty. Any difference in performance caused not by what students know but by how they are asked to show it—a phenomenon measurement specialists call a mode effect—can silently corrupt a trend line. To guard against this, every country that switched to digital in 2019 also administered a parallel paper-based “bridge” study with a nationally representative sample of students. The international evidence was sobering. A pilot item-equivalence study conducted before the 2019 cycle, together with the subsequent technical report, concluded that a statistical adjustment built into the international scaling procedures could preserve the linkage between formats—yet even after that adjustment, ten countries, among them Austria, Canada, Georgia, Hong Kong, Italy, Lithuania, the Netherlands, Portugal, Qatar, and Sweden, still showed residual country-level mode effects in at least one subject or grade. By 2023 the bridge study was no longer internationally required, because paper-to-digital links had already been established. Only four jurisdictions voluntarily ran one: Ireland, Azerbaijan, Bahrain, and Flanders.
The new report, compiled by Aidan Clerkin of the Educational Research Centre in Dublin together with colleagues in Belgium, Bahrain, and Azerbaijan, collates all four national comparison studies into a single picture. The design is elegantly simple. In each country, students were randomly assigned at the school or class level to sit either the digital main study or the paper bridge study, so the two groups should, on average, possess identical underlying proficiency; any systematic difference in scores can therefore be attributed to the mode of administration itself. Flanders ran its comparison at fourth grade only; Ireland and Bahrain covered both grades; and Azerbaijan’s bridge study covered eighth grade alone. The digital samples were nationally representative and large, ranging from 4,336 to 6,174 students per grade, while the paper samples ranged from 1,128 to 1,779 students per grade, selected through equivalent procedures and completing an abridged assessment built solely from trend items carried forward from 2019. Achievement was estimated with five plausible values per subject and analyzed using the IEA’s IDB Analyzer, allowing comparisons of overall means, gender differences, content and cognitive domains, and access to devices at home.
The headline verdict is unambiguous: at the level of overall achievement, the change of medium left scores untouched, in mathematics and science alike, in all four countries. A faint directional pattern was visible—differences tended to favor the digital test among fourth graders, while eighth-grade differences tended to favor paper in Ireland and Bahrain—but almost none of these gaps proved statistically significant once measurement error was taken into account. Azerbaijan broke the pattern in the opposite direction, with its eighth graders tending to score higher on screens. In practical terms, the authors conclude, the trend measurements linking 2019 to 2023 in these systems can be treated as valid: any rise or fall in the 2023 results reflects real changes in what children know, not an artefact of the keyboard.
The exception came from Azerbaijan, and it is the study’s most eye-catching finding. There, boys who took the science assessment digitally scored 408 points on the international reporting scale, while boys who took it on paper scored 382—a 26-point gap equivalent to one quarter of a standard deviation, far from trivial by the standards of international assessment. No corresponding difference appeared among girls, and the researchers found no gender-related mode effects anywhere else: in Ireland, Flanders, and Bahrain, the scores of boys and girls were statistically indistinguishable across formats. That reassurance matters now more than usual, because the TIMSS 2023 cycle reported widening gender gaps in several countries, and researchers trying to establish whether girls are genuinely losing ground need to know that the test itself is not distorting the picture. In Azerbaijan, however, the mode of delivery appears to interact with gender in a way that demands explanation.
The researchers also dissected the results by content domain within each subject, and by cognitive domain—the progression from Knowing, through Applying, to Reasoning. Here the findings were sparse and unsystematic in three countries. In Bahrain, eighth graders tested on paper outperformed their digital peers in the mathematics domain Data and Probability; in Ireland, paper-tested eighth graders led in Number and in Physics; and Bahrain recorded a solitary difference on the eighth-grade mathematics domain of Applying. Flanders showed no significant differences in either subject, and neither did any fourth-grade comparison anywhere. Azerbaijan, once again, stood apart: students tested digitally scored significantly higher in three of the four science domains—Physics, Chemistry, and Earth Science—as well as in Geometry and Measurement. The cognitive asymmetry was more telling still: on the higher-order domains of Applying and Reasoning, Azerbaijani students performed significantly better on the digital version of both the mathematics and the science assessments, while scores on the recall-oriented Knowing domain differed little—precisely the signature one would expect if screens somehow helped students think through complex, multi-step items.
The strongest pattern in the entire study, however, had nothing to do with test mode and everything to do with what children have at home. Across Flanders, Bahrain, and Azerbaijan, students with access to a computer or tablet at home outperformed those without—most sharply when they took the test on a screen. In Flanders, digitally tested fourth graders with home computer access scored 526 in mathematics and 494 in science, against 482 and 433 for peers without such access; in Bahrain the corresponding gaps were 468 versus 450 in mathematics and 481 versus 463 in science. Azerbaijani eighth graders showed the same asymmetry on the digital assessment, while in Ireland differences were small and non-significant at both grades and in both modes. The exception that illuminates the rule came from Bahrain’s eighth grade, where computer access predicted scores even among paper testers—hinting that there, a home computer functions partly as a proxy for socioeconomic status and broader learning resources rather than as digital familiarity alone.
The authors are candid about the limits of their evidence. With dozens of comparisons, some significant results are expected by chance alone—so-called Type I errors, or false positives—and the domain-level and device analyses were explicitly exploratory, since the bridge studies were designed primarily to test overall achievement. There is also a deeper structural asymmetry: the plausible values used to score the paper samples rest exclusively on trend items carried over from 2019, whereas the digital samples’ scores also absorb responses to PSI tasks and other innovative item types available only on screen. The two estimates, in other words, are not built from identical information. Nor is equivalence guaranteed elsewhere in the assessment world: sizeable mode effects have been documented in PISA 2015, in PISA reading short-text responses, and in PIRLS 2021, and Irish test developers once recorded a “substantial mode effect” among second graders on a national test—so pronounced that only a paper version was released at that grade level.
That is why the report singles out Azerbaijan for urgent follow-up: verifying that the digital and paper samples were truly equivalent in sociodemographic terms, drilling down to item-level comparisons—particularly on digital tools such as the on-screen ruler—and even cognitive interviews to understand how Azerbaijani children engage with low-stakes tests in each format. For the other three countries, the message is quiet confidence. Policymakers can now read changes between the 2019 and 2023 cycles as substantive signals, whether they stem from pandemic aftershocks, curriculum reforms, or shifting home and school contexts, rather than as artefacts of the switch from pencil to pixel. Countries that transitioned in 2023 without running their own bridge studies enjoy no such clarity, and may struggle to disentangle genuine changes in proficiency from procedural ones. In a world where a few points on an international test can trigger budget battles and political storms, that distinction is the difference between responding to a real crisis and chasing a ghost in the machine.
Cite Scienmag News
Celia A. (August 29, 2026). Switching to digital testing changes student scores, TIMSS evidence from four regions. Scienmag. https://scienmag.com/switching-to-digital-testing-changes-student-scores-timss-evidence-from-four-regions/
Celia A. "Switching to digital testing changes student scores, TIMSS evidence from four regions." Scienmag, 29 August 2026, https://scienmag.com/switching-to-digital-testing-changes-student-scores-timss-evidence-from-four-regions/. Accessed 29 August 2026.
Celia A. "Switching to digital testing changes student scores, TIMSS evidence from four regions." Scienmag. August 29, 2026. https://scienmag.com/switching-to-digital-testing-changes-student-scores-timss-evidence-from-four-regions/

