<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>PISA 2018 &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/pisa-2018/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 21:28:21 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>PISA 2018 &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Explainable AI Reveals What Really Predicts Math Achievement Across Ten Countries</title>
		<link>https://scienmag.com/explainable-ai-reveals-what-really-predicts-math-achievement-across-ten-countries/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:28:21 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[boosted tree models]]></category>
		<category><![CDATA[CatBoost]]></category>
		<category><![CDATA[comparative education]]></category>
		<category><![CDATA[comparative education methodology]]></category>
		<category><![CDATA[country-specific predictors]]></category>
		<category><![CDATA[cross-country education comparison]]></category>
		<category><![CDATA[educational data analysis]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[Explainable Artificial Intelligence]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[international education research]]></category>
		<category><![CDATA[large-scale assessment]]></category>
		<category><![CDATA[large-scale assessment analysis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in education]]></category>
		<category><![CDATA[mathematics achievement]]></category>
		<category><![CDATA[mathematics achievement prediction]]></category>
		<category><![CDATA[PISA 2018]]></category>
		<category><![CDATA[PISA student performance]]></category>
		<category><![CDATA[plausible values]]></category>
		<category><![CDATA[SHAP]]></category>
		<category><![CDATA[socioeconomic status]]></category>
		<category><![CDATA[STEM achievement factors]]></category>
		<category><![CDATA[survey weights]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202868</guid>

					<description><![CDATA[A survey-weighted explainable machine-learning analysis of PISA 2018 data from 74,235 students in ten education systems shows that books at home and parental occupational status are the most stable predictors of mathematics achievement, while other factors vary by country.]]></description>
										<content:encoded><![CDATA[<p>A new study has applied explainable artificial intelligence to one of the largest educational datasets in the world, and the results offer both a sobering confirmation and a methodological wake-up call for comparative education research. Working with mathematics achievement data from 74,235 fifteen-year-olds across ten purposively selected education systems, researchers Liu Liu of the University of Georgia and Rui Dai of Arizona State University built a survey-weighted, plausible-value-aware machine-learning workflow designed specifically for the complexities of the Programme for International Student Assessment, or PISA. Their analysis, published in Large-scale Assessments in Education, demonstrates that when the technical realities of large-scale assessment design are taken seriously, a boosted tree model consistently outperforms conventional linear benchmarks, and the predictors that matter most are strikingly stable in some respects and surprisingly country-specific in others.</p>
<p>The ten systems examined were Argentina, Chile, Chinese Taipei, Finland, Hungary, Italy, Japan, Korea, the Philippines, and the United States. The authors stress that this was a purposive comparative sample, chosen to span geographic regions, OECD membership status, and a wide range of average performance, from a weighted mean mathematics score of 352.57 in the Philippines to 531.14 in Chinese Taipei. The design supports comparisons across these ten systems but does not license claims about all countries or continents. Within the sample, 2,617 schools contributed students, with country sample sizes ranging from 4,838 in the United States to 11,975 in Argentina.</p>
<p>What distinguishes this study from much of the educational data-mining literature is its fidelity to PISA&#8217;s assessment architecture. Mathematics achievement in PISA is not a single number but a set of ten plausible values, multiple imputed estimates that represent uncertainty in each student&#8217;s latent proficiency. The researchers fitted and evaluated every model across all ten plausible values rather than relying on just the first. They also incorporated the final student weight, W_FSTUWT, in descriptive statistics, model fitting, and test-set evaluation, and used the 80 student replicate weights to estimate sampling variability. Total variance for each performance metric combined the mean replicate-weight sampling variance with the between-plausible-value variance, following the formula U-bar plus (1 + 1/M) times B, so that confidence intervals reflected both complex sampling uncertainty and achievement-scaling uncertainty. Few machine-learning studies of PISA data go to these lengths.</p>
<p>The predictor side of the analysis was equally disciplined. The authors constructed a transparent primary candidate pool of 33 predictors from PISA 2018 student questionnaire variables and OECD-derived indices, organized into six substantive domains: student demographics, family socioeconomic status and home resources, engagement and learning time, peer climate, belonging and parent support, and classroom climate. Rather than assuming all 33 variables mattered everywhere, they performed country-specific stability feature selection using only training schools. Three complementary methods were applied: mutual information, which captures general dependence; weighted elastic net regression, a regularized linear method that handles correlated predictors through combined L1 and L2 penalties; and weighted random forest ranking, which captures nonlinear and interaction-based structure. Feature rankings were repeated across ten plausible values and five internal school-level folds, yielding 50 selection runs per country. A predictor was deemed stable if selected in at least half of those runs. The resulting stable sets contained between 11 and 15 predictors, averaging 14.0 per country.</p>
<p>Four models were then compared under country-specific school-level holdout evaluation, with roughly 70 percent of students in training and 30 percent in test sets. Splitting by school rather than by student reduced leakage from students in the same school appearing on both sides. The contenders were weighted linear regression, weighted linear regression augmented with pairwise interactions, random forest, and CatBoost, a gradient-boosted tree ensemble designed for structured tabular data. Averaged across countries, CatBoost was the clear winner, achieving a mean weighted R-squared of 0.358 and a mean mean absolute error of 57.29 score points. Random forest followed with a mean weighted R-squared of 0.313 and an MAE of 59.30. The interaction-augmented linear model barely improved on the additive linear benchmark, with mean R-squared values of 0.290 versus 0.286. CatBoost posted the lowest weighted MAE in all ten systems, and its advantage over random forest was consistent but moderate: roughly 2.01 MAE points and 0.045 in R-squared on average. Its country-level R-squared ranged from 0.210 in Italy to 0.443 in Hungary.</p>
<p>The interpretive core of the study used SHAP, or Shapley additive explanations, a technique that decomposes each model prediction into additive contributions from individual features. The authors were careful to frame SHAP summaries as model-based explanations of predictive associations, not causal effects. Because CatBoost performed best, interpretation focused on its SHAP values, computed on held-out test schools and averaged across all ten plausible values to produce PV-robust stability summaries. The headline finding is that books at home was the most stable predictor of all, ranking among the top ten SHAP predictors in all ten countries on average and among the top five in 8.7 countries on average. Highest parental occupational status was nearly as stable, appearing in the top ten in 9.2 countries and the top five in 7.1. Socioeconomic status itself appeared in all ten country-specific models and ranked in the top ten in 8.3 countries on average, though its average top-five count was lower at 4.3.</p>
<p>Beyond the socioeconomic core, the picture became more heterogeneous. Grade placement entered the stable feature sets of seven countries and ranked in the top ten in all seven, but it was not selected in Chinese Taipei, Japan, or Korea. Mathematics learning time appeared in six countries and ranked in the top ten in all six, with an average top-five count of 5.5, and it ranked especially highly in Chinese Taipei, Japan, and the United States. Test effort was retained in only four systems but was consistently influential there, ranking in the top five in all four and even first in Finland, Italy, and Korea in the diagnostic visualization. Directed instruction and disciplinary climate also showed cross-national reach, ranking in the top ten in 7.0 and 5.8 countries on average respectively. The authors argue that this mix of stability and country-specificity is precisely why a single pooled importance ranking would be misleading.</p>
<p>Sensitivity analyses reinforced the robustness of these conclusions. Comparing SHAP summaries based on the first plausible value with summaries averaged across all ten showed only small differences, with maximum absolute gaps of 0.8 countries for top-ten counts, 1.2 countries for top-five counts, and 1.25 rank positions for mean rank, confirming that the main pattern was not an artifact of PV1MATH. A second sensitivity analysis removed grade placement from the stable feature sets where present and refitted the CatBoost models. Performance dropped modestly, with a mean R-squared change of -0.028 and a mean MAE increase of 1.17 score points, the largest decline occurring in Chile, where R-squared fell by 0.091 and MAE rose by 3.92 points. Crucially, the SHAP stability pattern remained broadly similar without grade: books at home, parental occupational status, learning time, socioeconomic status, test effort, directed instruction, and disciplinary climate all stayed among the most stable predictors.</p>
<p>The authors are explicit about the limits of their contribution. They do not claim to have discovered new determinants of mathematics achievement; domains such as socioeconomic resources, home literacy environments, learning time, and classroom climate are already well established. The value lies in showing how these familiar predictors behave under a rigorous, country-specific, survey-weighted, plausible-value-aware explainable machine-learning framework, and in separating three questions that are often conflated: which models predict best, which predictors are stable across systems, and which are context-specific. They also caution that the analysis is predictive rather than causal, that the ten-system sample is not statistically representative of global education, that the 33-predictor pool cannot exhaust all relevant influences such as school policies or teacher characteristics, and that cross-national questionnaire comparisons may be affected by response styles, translation, and measurement comparability.</p>
<p>The implications reach well beyond this dataset. The framework, with its train-only feature selection, school-level holdout evaluation, replicate-weight variance estimation, and all-plausible-value pooling, offers a reproducible template that could be extended to more PISA systems, to other assessments such as TIMSS and PIRLS, to school-level and system-level predictors, and to fairness analyses examining whether prediction errors or SHAP patterns differ across gender, socioeconomic, immigrant, or language groups. Repeated-cycle analyses could test whether cross-national stability patterns persist as PISA evolves. For a field increasingly drawn to black-box prediction, the study makes a compelling case that explainability and methodological rigor are not optional extras but the foundation of responsible machine learning in comparative education.</p>
<p><strong>Subject of Research:</strong> Explainable machine learning applied to predicting and interpreting mathematics achievement in PISA 2018 across ten education systems.</p>
<p><strong>Article Title:</strong> Explainable AI for predicting and interpreting mathematics achievement: a cross-national analysis of PISA 2018</p>
<p><strong>Article References:</strong> Liu, L., &amp; Dai, R. (2026). Explainable AI for predicting and interpreting mathematics achievement: a cross-national analysis of PISA 2018. <em>Large-scale Assessments in Education, 14</em>(1), Article 46. <a href="https://doi.org/10.1186/s40536-026-00320-y" rel="noopener noreferrer">https://doi.org/10.1186/s40536-026-00320-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40536-026-00320-y" rel="noopener noreferrer">10.1186/s40536-026-00320-y</a></p>
<p><strong>Keywords:</strong> PISA 2018, mathematics achievement, explainable artificial intelligence, SHAP, CatBoost, machine learning, survey weights, plausible values, large-scale assessment, socioeconomic status, comparative education, feature selection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202868</post-id>	</item>
		<item>
		<title>Who Counts as Resilient? Study Reveals How Definitions Reshape Education Rankings</title>
		<link>https://scienmag.com/who-counts-as-resilient-study-reveals-how-definitions-reshape-education-rankings/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 12:33:24 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[academic resilience]]></category>
		<category><![CDATA[cross-national comparison]]></category>
		<category><![CDATA[cross-national education comparisons]]></category>
		<category><![CDATA[Educational Equity]]></category>
		<category><![CDATA[educational inequality and resilience]]></category>
		<category><![CDATA[educational measurement]]></category>
		<category><![CDATA[Educational resilience]]></category>
		<category><![CDATA[effects on country rankings]]></category>
		<category><![CDATA[expectancy-value theory]]></category>
		<category><![CDATA[impact of resilience definitions]]></category>
		<category><![CDATA[influence of operational definitions on resilience data]]></category>
		<category><![CDATA[large-scale assessment]]></category>
		<category><![CDATA[large-scale assessment analysis]]></category>
		<category><![CDATA[measurement of disadvantaged student success]]></category>
		<category><![CDATA[methodological challenges in resilience research]]></category>
		<category><![CDATA[OECD PISA statistics]]></category>
		<category><![CDATA[operationalization]]></category>
		<category><![CDATA[PISA 2018]]></category>
		<category><![CDATA[policy implications of resilience measurement]]></category>
		<category><![CDATA[protective factors]]></category>
		<category><![CDATA[socioeconomic status]]></category>
		<category><![CDATA[student motivation]]></category>
		<category><![CDATA[thresholds]]></category>
		<category><![CDATA[variability in resilience prevalence]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194223</guid>

					<description><![CDATA[A new analysis of PISA 2018 data from 75 education systems shows that the choice of operational definition and threshold can change estimates of academic resilience nearly fivefold and reshape which countries rank as equity leaders.]]></description>
										<content:encoded><![CDATA[<p>Every few years, the OECD&#8217;s PISA results include a statistic that education ministers around the world eagerly quote: the percentage of disadvantaged students who nevertheless succeed at school. These academically resilient students are held up as proof that poverty need not determine achievement, and their numbers fuel cross-national comparisons, policy borrowing, and headlines about which school systems beat the odds. But a sweeping new analysis suggests that this celebrated statistic may be far less solid than it appears. Depending on how researchers choose to define resilience, the share of resilient students in the very same dataset can vary almost fivefold, and the countries ranked highest or lowest can shift dramatically from one definition to the next.</p>
<p>The study, published in the journal Large-scale Assessments in Education, was conducted by Markéta Žáková and Tomáš Lintner of Masaryk University in the Czech Republic. Drawing on PISA 2018 data from more than 600,000 fifteen-year-old students across 75 education systems, the pair set out to answer a deceptively simple question: does it matter which operational definition of academic resilience a researcher uses? The answer, they found, is a resounding yes, with consequences that ripple through prevalence estimates, country rankings, and conclusions about which factors help disadvantaged students beat the odds.</p>
<p>Academic resilience combines two ingredients: adversity, usually measured by socioeconomic status, and positive adaptation, usually measured by achievement. But researchers disagree about how to stitch those ingredients together. One common approach simply identifies students who land in the bottom slice of the socioeconomic distribution and the top slice of achievement, for example the poorest quarter who score in the top quarter. A second approach instead computes what each student&#8217;s achievement should be, statistically speaking, given their family background, and flags students who substantially outperform that expectation, an approach known as the residual method. A third treats resilience as a process rather than an outcome, predicting achievement as a continuous variable within the disadvantaged group and never labeling anyone resilient at all.</p>
<p>Žáková and Lintner compared these approaches systematically. They estimated a top-achiever definition, a residual definition computed within each country, and a residual definition benchmarked against an international standard, each at three different threshold levels, 20, 25, and 33 percent, spanning the range most commonly used in the published literature. All told, this produced nine binary operationalizations plus the continuous-outcome model, yielding nearly a thousand country-level analyses. Each estimate was pooled across ten plausible values for achievement and twenty multiply imputed datasets, with standard errors that fully accounted for PISA&#8217;s complex two-stage sampling design.</p>
<p>The headline finding is stark. The cross-country mean prevalence of academically resilient students ranged from 7.5 percent under the strictest definition to 35.3 percent under the most inclusive residual benchmark, a nearly fivefold difference computed from identical data. Part of this spread is arithmetic, since looser thresholds mechanically admit more students, but the deeper problem emerges when countries are ranked. Rankings were reasonably stable across thresholds within a given definition, yet they diverged sharply across definitions. The correlation between rankings produced by the top-achiever approach and the residual approaches fell as low as 0.51, and between the two residual variants, which differ only in whether the statistical expectation is local or global, it dropped to between 0.32 and 0.47. A country celebrated as an equity champion under one definition can rank unremarkably under another, and the discrepancy is itself informative about what its disadvantaged students actually do well.</p>
<p>The authors illustrate why with a thought experiment grounded in their results. A lower-performing system with a steep socioeconomic gradient may rank poorly on the top-achiever definition, because few of its disadvantaged students reach absolute excellence, yet rank highly on the within-country residual definition, because modest local expectations are easy to exceed. That same system may then sink again on the internationally benchmarked residual definition, since beating a weak local bar is not the same as meeting a global standard. The researchers argue that resilience rankings should therefore be reported under multiple definitions side by side, as complementary views of equity rather than competing estimates of a single quantity, a practice the OECD itself briefly adopted nearly a decade ago.</p>
<p>What about the protective factors that resilience research is meant to uncover? Here the news is more reassuring at the global level and more troubling at the level of individual countries. When the authors pooled effects meta-analytically across all 75 systems, three motivational constructs drawn from expectancy-value theory, students&#8217; self-perceived reading competence, their enjoyment of reading, and their attitude toward learning, were consistently and positively associated with resilience under every operationalization and threshold, and girls consistently outperformed boys. Read that result alone, and the choice of definition seems inconsequential.</p>
<p>The country-by-country picture tells a different story. When each education system was analyzed separately, as is standard in PISA-based research, findings were consistent across all nine binary specifications in only 20 percent of systems for gender, 40 percent for attitude toward learning, 69 percent for reading enjoyment, and 72 percent for self-perceived competence. Just two of the 75 systems produced fully consistent results for all four predictors. Threshold level alone flipped conclusions in somewhere between 3 and 37 percent of countries depending on the factor. Much of this inconsistency reflects statistical power, since stricter thresholds shrink samples and smaller-effect predictors, notably gender and attitude toward learning, are the least stable. But the practical consequence does not depend on the cause: a researcher studying one country under one definition could legitimately conclude that a factor matters there when a defensible alternative would say it does not.</p>
<p>The most conceptually striking result came from a complementary interaction analysis that requires no thresholds at all. By modeling whether the socioeconomic gradient in achievement flattens at higher levels of each factor, the authors tested what the term protective factor actually implies. The findings complicate comfortable assumptions: the gradient was flatter, not steeper, among students with higher self-perceived reading competence and stronger attitudes toward learning, but it was significantly steeper, not flatter, among students who enjoyed reading more. In other words, reading enjoyment, though positively associated with achievement on average, was linked to wider rather than narrower socioeconomic gaps, a pattern that no threshold-based definition could ever detect. These moderation effects also varied in direction across countries, appearing in only a minority of systems individually.</p>
<p>The authors close with four practical recommendations: match the operationalization to the research question, report sensitivity analyses at a minimum of two thresholds, present cross-country rankings under multiple definitions, and document or pre-register operationalization decisions so findings can be interpreted in light of the choices that produced them. They caution that their analysis is cross-sectional and cannot establish causation, that their predictors were individual-level only, and that PISA&#8217;s socioeconomic index has itself attracted methodological criticism. Still, the broader message is hard to escape. Academic resilience, as measured in large-scale assessments, is not a fixed quantity waiting to be counted but a construct whose observed properties depend partly on how it is defined, and both researchers and policymakers ignore that dependence at their peril.</p>
<p><strong>Subject of Research:</strong> How operationalization and threshold choices affect estimates of academic resilience and its protective factors across 75 education systems using PISA 2018 data.</p>
<p><strong>Article Title:</strong> Academic resilience in 75 education systems: how operationalization and threshold choices shape findings on protective factors</p>
<p><strong>Article References:</strong> Academic resilience in 75 education systems: how operationalization and threshold choices shape findings on protective factors. (n.d.). <a href="https://doi.org/10.1186/s40536-026-00318-6" rel="noopener noreferrer">https://doi.org/10.1186/s40536-026-00318-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40536-026-00318-6" rel="noopener noreferrer">10.1186/s40536-026-00318-6</a></p>
<p><strong>Keywords:</strong> academic resilience, PISA 2018, socioeconomic status, educational equity, large-scale assessment, protective factors, operationalization, thresholds, student motivation, expectancy-value theory, educational measurement, cross-national comparison</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194223</post-id>	</item>
	</channel>
</rss>
