Every three years, the release of PISA results sends shockwaves through ministries of education around the world. Headlines declare national decline or triumph, and governments frequently respond with sweeping reforms. But a new study raises an uncomfortable question: is the world’s most influential international test actually measuring the same thing that schools measure? According to research published in Large-scale Assessments in Education, the answer is a qualified no. Drawing on an unusually rich dataset that links the Swedish PISA 2022 results with national test scores and teacher-assigned grades, a team led by Linda Borger of the University of Gothenburg, together with Rolf Strietholt, Stefan Johansson, and Samuel Greiff, found that PISA behaves like a fundamentally different kind of achievement measure from the assessments schools use every day.
The study exploited a rare opportunity. When Sweden administered PISA 2022, the national agency for education recorded students’ personal identification numbers, allowing the researchers to link the test scores of 6,072 fifteen-year-olds with official register data containing their Grade 9 national test results and final subject grades in Swedish, mathematics, and science. Because PISA samples by age rather than grade, nearly 97 percent of the participants were ninth-graders, making the comparison across the three assessment types unusually clean. Linkage was achieved for 99.5 percent of the sample, and the researchers applied survey weights and adjusted for the clustering of students within 262 schools to ensure that their statistical models reflected the true population structure.
The conceptual stakes are high because the three measures were built for different purposes. PISA, run by the OECD, deliberately avoids any national curriculum. It assesses what the OECD calls literacy: the capacity of students to analyse, reason, communicate, and apply knowledge in unfamiliar, real-world situations, whether or not those situations resemble classroom content. National tests and school grades, by contrast, are explicitly curriculum-based, evaluating mastery of subject syllabuses. The stakes also differ sharply. PISA is low-stakes for individual students, who receive no feedback and face no consequences, whereas Swedish grades determine admission to upper secondary and tertiary education, and national test results feed into grading decisions. Motivation research suggests these differences can affect how hard students try, potentially depressing PISA performance relative to what students know.
The first striking finding concerned the correlations among the PISA domains themselves. PISA reading, mathematics, and science scores were intercorrelated at values between 0.81 and 0.88, an extraordinarily high degree of overlap for tests that are supposedly distinct subject assessments. By comparison, the correlations between PISA scores and the school achievement measures ranged from 0.58 to 0.74, moderately strong but clearly not indicating interchangeability. Meanwhile, the school-based measures correlated with each other at 0.62 to 0.92, with the strongest link, 0.92, between mathematics grades and mathematics national test results. Notably, PISA reading correlated almost equally with grades and tests in Swedish, mathematics, and science, around 0.61 to 0.63, suggesting that reading literacy acts as a broad competence that underpins performance across the entire curriculum.
To disentangle what drives these patterns, the researchers turned to confirmatory factor analysis, specifically a bifactor model with a multitrait-multimethod structure. In this framework, the three subjects, Swedish or reading, mathematics, and science, serve as traits, while the three assessment types, PISA, national tests, and grades, serve as methods. A general achievement factor captures variance common to all nine measures, while orthogonal specific factors capture variance unique to particular subjects or particular assessment formats. The researchers built up their models step by step: a single general factor fit poorly, adding method factors improved things only marginally, but a model combining the general factor with both method and subject factors achieved acceptable fit across all standard indices, including CFI, TLI, RMSEA, and SRMR.
The factor loadings told a clear and, for PISA, sobering story. All nine measures loaded strongly on the general achievement factor, with loadings between 0.69 and 0.96, confirming that a broad academic ability underlies performance regardless of how it is measured. But the method factors revealed an asymmetry. The PISA method factor showed substantial loadings of 0.54 to 0.62, meaning a large share of PISA variance is specific to the PISA assessment format itself. The national test method factor, in contrast, loaded below 0.20, indicating that national tests share nearly all their variance with grades and add little method-specific noise. In other words, grades and national tests behave as close relatives, while PISA stands apart.
The subject factors delivered the study’s most consequential result. Grades and national tests showed strong loadings of roughly 0.50 on the Swedish and mathematics specific factors, demonstrating that these school-based measures genuinely differentiate between what a student knows in one subject versus another. The corresponding PISA loadings were tiny, between 0.12 and 0.16. PISA, in short, barely distinguishes reading from mathematics from science. Rather than mapping subject-specific curricular learning, it appears to capture a general, literacy-oriented competence shaped by its own assessment framework. This finding echoes earlier item-level studies, including work by Pokropek and colleagues across 33 countries, which found that a general factor accounted for about 84 percent of the common variance in PISA scores, with subject-specific factors too weak to interpret reliably.
Why does PISA blur the boundaries between subjects? The authors point to several plausible mechanisms. PISA’s literacy framework emphasizes applied, real-world contexts, and its items are notoriously text-heavy: research on Swedish PISA mathematics tasks found they demand far more reading than the national mathematics tests, so reading comprehension skill can drive performance even in the mathematics domain. PISA’s plausible-value imputation procedure may also inflate inter-domain correlations, because each student’s scores are conditioned on performance in other domains and on background variables. Additionally, PISA administers items from all domains in a single session, so fatigue, motivation, and other situational factors may influence performance across domains simultaneously. Low stakes may further dampen effort; as the OECD itself notes, PISA captures what students actually do under low-consequence conditions, not their maximum potential.
The policy implications are significant. Because PISA results often trigger national reforms, the authors warn against interpreting PISA scores as direct indicators of subject-level proficiency or using them to justify curriculum-specific interventions. Scholars have cautioned that over-reliance on international assessments can narrow curricula and misdirect priorities. At the same time, the researchers are careful not to declare PISA useless: it remains a valuable, curriculum-independent benchmark for cross-national comparison and trend monitoring, immune to national grade inflation. The honest conclusion is that PISA, national tests, and grades measure overlapping but non-interchangeable constructs. Whether PISA’s broader literacy focus or the schools’ subject-specific measures better capture socially relevant learning is, the authors stress, a normative question that statistics alone cannot settle. What statistics can settle, and this study does, is that treating these metrics as equivalent would be a measurement mistake with real consequences for education policy.
Subject of Research: Comparing the measurement structure of PISA scores, national tests, and teacher-assigned grades in reading, mathematics, and science
Article Title: Crossing metrics: are PISA, national tests, and grades measuring the same?
Article References: Borger, L., Strietholt, R., Johansson, S., & Greiff, S. (2026). Crossing metrics: are PISA, national tests, and grades measuring the same?. Large-scale Assessments in Education, 14(1), Article 47. https://doi.org/10.1186/s40536-026-00321-x
Image Credits: AI Generated
DOI: 10.1186/s40536-026-00321-x
Keywords: PISA, educational assessment, psychometrics, school grades, national tests, large-scale assessment, bifactor model, multitrait-multimethod, test validity, education policy, Sweden, student achievement
Cite Scienmag News
Courtney Benton. (September 30, 2026). PISA Scores Are Not Interchangeable With Grades and National Tests, Swedish Study Finds. Scienmag. https://scienmag.com/pisa-scores-are-not-interchangeable-with-grades-and-national-tests-swedish-study-finds/
Courtney Benton. "PISA Scores Are Not Interchangeable With Grades and National Tests, Swedish Study Finds." Scienmag, 30 September 2026, https://scienmag.com/pisa-scores-are-not-interchangeable-with-grades-and-national-tests-swedish-study-finds/. Accessed 30 September 2026.
Courtney Benton. "PISA Scores Are Not Interchangeable With Grades and National Tests, Swedish Study Finds." Scienmag. September 30, 2026. https://scienmag.com/pisa-scores-are-not-interchangeable-with-grades-and-national-tests-swedish-study-finds/

