<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>plausible values &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/plausible-values/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 21:28:21 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>plausible values &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Explainable AI Reveals What Really Predicts Math Achievement Across Ten Countries</title>
		<link>https://scienmag.com/explainable-ai-reveals-what-really-predicts-math-achievement-across-ten-countries/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:28:21 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[boosted tree models]]></category>
		<category><![CDATA[CatBoost]]></category>
		<category><![CDATA[comparative education]]></category>
		<category><![CDATA[comparative education methodology]]></category>
		<category><![CDATA[country-specific predictors]]></category>
		<category><![CDATA[cross-country education comparison]]></category>
		<category><![CDATA[educational data analysis]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[Explainable Artificial Intelligence]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[international education research]]></category>
		<category><![CDATA[large-scale assessment]]></category>
		<category><![CDATA[large-scale assessment analysis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in education]]></category>
		<category><![CDATA[mathematics achievement]]></category>
		<category><![CDATA[mathematics achievement prediction]]></category>
		<category><![CDATA[PISA 2018]]></category>
		<category><![CDATA[PISA student performance]]></category>
		<category><![CDATA[plausible values]]></category>
		<category><![CDATA[SHAP]]></category>
		<category><![CDATA[socioeconomic status]]></category>
		<category><![CDATA[STEM achievement factors]]></category>
		<category><![CDATA[survey weights]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202868</guid>

					<description><![CDATA[A survey-weighted explainable machine-learning analysis of PISA 2018 data from 74,235 students in ten education systems shows that books at home and parental occupational status are the most stable predictors of mathematics achievement, while other factors vary by country.]]></description>
										<content:encoded><![CDATA[<p>A new study has applied explainable artificial intelligence to one of the largest educational datasets in the world, and the results offer both a sobering confirmation and a methodological wake-up call for comparative education research. Working with mathematics achievement data from 74,235 fifteen-year-olds across ten purposively selected education systems, researchers Liu Liu of the University of Georgia and Rui Dai of Arizona State University built a survey-weighted, plausible-value-aware machine-learning workflow designed specifically for the complexities of the Programme for International Student Assessment, or PISA. Their analysis, published in Large-scale Assessments in Education, demonstrates that when the technical realities of large-scale assessment design are taken seriously, a boosted tree model consistently outperforms conventional linear benchmarks, and the predictors that matter most are strikingly stable in some respects and surprisingly country-specific in others.</p>
<p>The ten systems examined were Argentina, Chile, Chinese Taipei, Finland, Hungary, Italy, Japan, Korea, the Philippines, and the United States. The authors stress that this was a purposive comparative sample, chosen to span geographic regions, OECD membership status, and a wide range of average performance, from a weighted mean mathematics score of 352.57 in the Philippines to 531.14 in Chinese Taipei. The design supports comparisons across these ten systems but does not license claims about all countries or continents. Within the sample, 2,617 schools contributed students, with country sample sizes ranging from 4,838 in the United States to 11,975 in Argentina.</p>
<p>What distinguishes this study from much of the educational data-mining literature is its fidelity to PISA&#8217;s assessment architecture. Mathematics achievement in PISA is not a single number but a set of ten plausible values, multiple imputed estimates that represent uncertainty in each student&#8217;s latent proficiency. The researchers fitted and evaluated every model across all ten plausible values rather than relying on just the first. They also incorporated the final student weight, W_FSTUWT, in descriptive statistics, model fitting, and test-set evaluation, and used the 80 student replicate weights to estimate sampling variability. Total variance for each performance metric combined the mean replicate-weight sampling variance with the between-plausible-value variance, following the formula U-bar plus (1 + 1/M) times B, so that confidence intervals reflected both complex sampling uncertainty and achievement-scaling uncertainty. Few machine-learning studies of PISA data go to these lengths.</p>
<p>The predictor side of the analysis was equally disciplined. The authors constructed a transparent primary candidate pool of 33 predictors from PISA 2018 student questionnaire variables and OECD-derived indices, organized into six substantive domains: student demographics, family socioeconomic status and home resources, engagement and learning time, peer climate, belonging and parent support, and classroom climate. Rather than assuming all 33 variables mattered everywhere, they performed country-specific stability feature selection using only training schools. Three complementary methods were applied: mutual information, which captures general dependence; weighted elastic net regression, a regularized linear method that handles correlated predictors through combined L1 and L2 penalties; and weighted random forest ranking, which captures nonlinear and interaction-based structure. Feature rankings were repeated across ten plausible values and five internal school-level folds, yielding 50 selection runs per country. A predictor was deemed stable if selected in at least half of those runs. The resulting stable sets contained between 11 and 15 predictors, averaging 14.0 per country.</p>
<p>Four models were then compared under country-specific school-level holdout evaluation, with roughly 70 percent of students in training and 30 percent in test sets. Splitting by school rather than by student reduced leakage from students in the same school appearing on both sides. The contenders were weighted linear regression, weighted linear regression augmented with pairwise interactions, random forest, and CatBoost, a gradient-boosted tree ensemble designed for structured tabular data. Averaged across countries, CatBoost was the clear winner, achieving a mean weighted R-squared of 0.358 and a mean mean absolute error of 57.29 score points. Random forest followed with a mean weighted R-squared of 0.313 and an MAE of 59.30. The interaction-augmented linear model barely improved on the additive linear benchmark, with mean R-squared values of 0.290 versus 0.286. CatBoost posted the lowest weighted MAE in all ten systems, and its advantage over random forest was consistent but moderate: roughly 2.01 MAE points and 0.045 in R-squared on average. Its country-level R-squared ranged from 0.210 in Italy to 0.443 in Hungary.</p>
<p>The interpretive core of the study used SHAP, or Shapley additive explanations, a technique that decomposes each model prediction into additive contributions from individual features. The authors were careful to frame SHAP summaries as model-based explanations of predictive associations, not causal effects. Because CatBoost performed best, interpretation focused on its SHAP values, computed on held-out test schools and averaged across all ten plausible values to produce PV-robust stability summaries. The headline finding is that books at home was the most stable predictor of all, ranking among the top ten SHAP predictors in all ten countries on average and among the top five in 8.7 countries on average. Highest parental occupational status was nearly as stable, appearing in the top ten in 9.2 countries and the top five in 7.1. Socioeconomic status itself appeared in all ten country-specific models and ranked in the top ten in 8.3 countries on average, though its average top-five count was lower at 4.3.</p>
<p>Beyond the socioeconomic core, the picture became more heterogeneous. Grade placement entered the stable feature sets of seven countries and ranked in the top ten in all seven, but it was not selected in Chinese Taipei, Japan, or Korea. Mathematics learning time appeared in six countries and ranked in the top ten in all six, with an average top-five count of 5.5, and it ranked especially highly in Chinese Taipei, Japan, and the United States. Test effort was retained in only four systems but was consistently influential there, ranking in the top five in all four and even first in Finland, Italy, and Korea in the diagnostic visualization. Directed instruction and disciplinary climate also showed cross-national reach, ranking in the top ten in 7.0 and 5.8 countries on average respectively. The authors argue that this mix of stability and country-specificity is precisely why a single pooled importance ranking would be misleading.</p>
<p>Sensitivity analyses reinforced the robustness of these conclusions. Comparing SHAP summaries based on the first plausible value with summaries averaged across all ten showed only small differences, with maximum absolute gaps of 0.8 countries for top-ten counts, 1.2 countries for top-five counts, and 1.25 rank positions for mean rank, confirming that the main pattern was not an artifact of PV1MATH. A second sensitivity analysis removed grade placement from the stable feature sets where present and refitted the CatBoost models. Performance dropped modestly, with a mean R-squared change of -0.028 and a mean MAE increase of 1.17 score points, the largest decline occurring in Chile, where R-squared fell by 0.091 and MAE rose by 3.92 points. Crucially, the SHAP stability pattern remained broadly similar without grade: books at home, parental occupational status, learning time, socioeconomic status, test effort, directed instruction, and disciplinary climate all stayed among the most stable predictors.</p>
<p>The authors are explicit about the limits of their contribution. They do not claim to have discovered new determinants of mathematics achievement; domains such as socioeconomic resources, home literacy environments, learning time, and classroom climate are already well established. The value lies in showing how these familiar predictors behave under a rigorous, country-specific, survey-weighted, plausible-value-aware explainable machine-learning framework, and in separating three questions that are often conflated: which models predict best, which predictors are stable across systems, and which are context-specific. They also caution that the analysis is predictive rather than causal, that the ten-system sample is not statistically representative of global education, that the 33-predictor pool cannot exhaust all relevant influences such as school policies or teacher characteristics, and that cross-national questionnaire comparisons may be affected by response styles, translation, and measurement comparability.</p>
<p>The implications reach well beyond this dataset. The framework, with its train-only feature selection, school-level holdout evaluation, replicate-weight variance estimation, and all-plausible-value pooling, offers a reproducible template that could be extended to more PISA systems, to other assessments such as TIMSS and PIRLS, to school-level and system-level predictors, and to fairness analyses examining whether prediction errors or SHAP patterns differ across gender, socioeconomic, immigrant, or language groups. Repeated-cycle analyses could test whether cross-national stability patterns persist as PISA evolves. For a field increasingly drawn to black-box prediction, the study makes a compelling case that explainability and methodological rigor are not optional extras but the foundation of responsible machine learning in comparative education.</p>
<p><strong>Subject of Research:</strong> Explainable machine learning applied to predicting and interpreting mathematics achievement in PISA 2018 across ten education systems.</p>
<p><strong>Article Title:</strong> Explainable AI for predicting and interpreting mathematics achievement: a cross-national analysis of PISA 2018</p>
<p><strong>Article References:</strong> Liu, L., &amp; Dai, R. (2026). Explainable AI for predicting and interpreting mathematics achievement: a cross-national analysis of PISA 2018. <em>Large-scale Assessments in Education, 14</em>(1), Article 46. <a href="https://doi.org/10.1186/s40536-026-00320-y" rel="noopener noreferrer">https://doi.org/10.1186/s40536-026-00320-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40536-026-00320-y" rel="noopener noreferrer">10.1186/s40536-026-00320-y</a></p>
<p><strong>Keywords:</strong> PISA 2018, mathematics achievement, explainable artificial intelligence, SHAP, CatBoost, machine learning, survey weights, plausible values, large-scale assessment, socioeconomic status, comparative education, feature selection</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202868</post-id>	</item>
		<item>
		<title>AI Optimism in Principals&#8217; Offices Does Not Translate Into Better Student Digital Skills, Landmark 12-Country Study Finds</title>
		<link>https://scienmag.com/ai-optimism-in-principals-offices-does-not-translate-into-better-student-digital-skills-landmark-12-country-study-finds/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 15:29:51 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[AI in school leadership]]></category>
		<category><![CDATA[ChatGPT expectations]]></category>
		<category><![CDATA[comparative education research on AI adoption]]></category>
		<category><![CDATA[computer and information literacy]]></category>
		<category><![CDATA[cross-country education system analysis]]></category>
		<category><![CDATA[digital inequality]]></category>
		<category><![CDATA[effectiveness of AI integration in classrooms]]></category>
		<category><![CDATA[false discovery rate]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[generative AI influence on education]]></category>
		<category><![CDATA[ICILS 2023]]></category>
		<category><![CDATA[ICT in education]]></category>
		<category><![CDATA[impact of principals' AI expectations]]></category>
		<category><![CDATA[influence of school policies on digital skills]]></category>
		<category><![CDATA[international assessment]]></category>
		<category><![CDATA[International Computer and Information Literacy Study 2023]]></category>
		<category><![CDATA[Multilevel modeling]]></category>
		<category><![CDATA[plausible values]]></category>
		<category><![CDATA[role of school leadership in digital literacy]]></category>
		<category><![CDATA[school digital conditions]]></category>
		<category><![CDATA[school digital conditions and student literacy]]></category>
		<category><![CDATA[socioeconomic status]]></category>
		<category><![CDATA[student computer and information literacy measurement]]></category>
		<category><![CDATA[student digital skills assessment]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=195923</guid>

					<description><![CDATA[A multilevel analysis of ICILS 2023 data from 24,273 students in 12 education systems finds no reliable direct association between principals' ChatGPT expectations, ICT expectations, or reported school digital hindrances and students' assessed computer and information literacy.]]></description>
										<content:encoded><![CDATA[<p>When school leaders embrace artificial intelligence, do their students become better at actually using computers to find, evaluate, and create information? A sweeping new analysis of data from 12 education systems suggests the answer, at least so far, is no. Drawing on the International Computer and Information Literacy Study 2023, or ICILS 2023, researchers examined whether five school-level digital conditions, including principals&#8217; expectations about ChatGPT, were associated with students&#8217; assessed computer and information literacy, commonly abbreviated CIL. The verdict is striking: after rigorous statistical adjustment and strict control for multiple testing, none of the five school conditions showed a reliable direct association with student CIL in any of the participating systems.</p>
<p>The study, led by Sukanya Chaemchoy of Chulalongkorn University together with Thanakrit Supsin, analyzed data from 24,273 students nested in 1,049 schools across Chile, Cyprus, Denmark, Greece, Korea, Norway, Romania, the Slovak Republic, Slovenia, Sweden, Chinese Taipei, and Uruguay. These were the 12 systems that administered ICILS 2023&#8217;s optional questionnaire on generative AI, which asked principals how likely they believed ChatGPT and similar tools were to help or harm students&#8217; learning in their schools. The researchers linked these leadership reports with achievement data, principal questionnaires, and surveys completed by schools&#8217; ICT coordinators, constructing an unusually rich picture of the organizational digital climate surrounding each student.</p>
<p>The five focal conditions were carefully distinguished. Positive ChatGPT expectations measured principals&#8217; anticipated benefits, such as greater student interest in learning, better written work, and improved critical evaluation of information. Negative ChatGPT expectations captured anticipated harms, including shallow conceptual understanding, submission of work that is not the student&#8217;s own, and dependence on AI tools. Instructional ICT-use expectations reflected whether teachers were expected to integrate digital technology into teaching, assessment, and monitoring of progress, while ICT collaboration expectations concerned professional communication and collaboration via technology. Finally, pedagogical ICT hindrances, reported by ICT coordinators, captured constraints such as insufficient teacher skills, limited preparation time, weak pedagogical support, and restrictive policies.</p>
<p>Methodologically, the study is a masterclass in caution. The authors estimated separate weighted two-level models for each education system, treating students as nested within schools and adjusting for student sex, home internet access, computer experience, within-school socioeconomic background, and school socioeconomic composition. Because CIL was measured using five plausible values, each model was run five times and the results formally combined. The family of 60 primary tests, five school conditions across 12 systems, was then subjected to Benjamini-Hochberg false discovery rate control, a correction that dramatically raises the bar for what counts as a credible finding in large-scale educational research.</p>
<p>The outcome was unambiguous at the top line: not a single school condition produced a false-discovery-rate-retained association with CIL. Three coefficients did carry raw p values below 0.05. In Korea, higher ICT collaboration expectations were associated with roughly 9.55 points lower CIL, and pedagogical ICT hindrances were associated with about 4.11 points lower CIL. In the Slovak Republic, the direction reversed dramatically: ICT collaboration expectations were associated with 7.82 points higher CIL. Yet all three signals carried an FDR-adjusted q value of 0.730, meaning they are best read as nominal, system-specific hints for future replication rather than established associations.</p>
<p>The Korea-Slovak Republic contrast is perhaps the most intriguing descriptive finding. The same survey instrument, measuring the same construct, produced estimates with opposite signs and non-overlapping confidence intervals in the two systems. The authors are careful not to overclaim: no formal test of slope heterogeneity was conducted, and the study did not measure the national policies, governance arrangements, curricula, or implementation histories that might explain the divergence. But the pattern echoes earlier ICILS research from 2013, which found that the school-level conditions linked to teachers&#8217; ICT use differed across Australia, the Czech Republic, Germany, and Norway, suggesting that identical school-scale scores may be embedded in fundamentally different institutional realities.</p>
<p>A secondary analysis reinforced this caution about pooled summaries. When the researchers constrained the focal slopes to be equal across systems, using equal total weights for each education system, all five common-slope confidence intervals included zero. The near-zero pooled estimate for ICT collaboration expectations simply cannot represent both Korea&#8217;s negative and the Slovak Republic&#8217;s positive coefficient. As the authors note, adjusting for system mean differences through country indicators does not demonstrate that school-level relationships are homogeneous, a point long emphasized in methodological work on multilevel modeling of country effects.</p>
<p>Six prespecified families of sensitivity analyses, covering 350 focal comparisons, tested whether the conclusions depended on weighting choices, socioeconomic decomposition, coding decisions, complete-case selection, influential schools, or survey-design variance estimation. Every sensitivity confidence interval overlapped its primary counterpart, and no comparison met the prespecified material-sensitivity criterion. The three nominal signals kept their direction in every available comparison, but they never escaped the multiplicity adjustment. In short, the null finding for school digital conditions is not a fragile artifact of one particular model specification.</p>
<p>What did matter, consistently, was student background. Within-school socioeconomic background showed positive adjusted associations with CIL in all 12 systems, with coefficients ranging from 7.78 to 26.52 CIL points per index point. School socioeconomic composition was also positive everywhere, at 30.57 to 64.49 points per index point, though the authors treat it cautiously because the aggregated school mean had reliability below 0.90 in seven systems. Computer experience was positively associated with CIL in every system, and female students outperformed male students in adjusted comparisons across the board. These results align with prior meta-analytic evidence of a positive, if modest, relationship between socioeconomic status and ICT literacy.</p>
<p>The implications reach beyond academia. As governments pour resources into digital infrastructure and school leaders form opinions about generative AI, this study warns against conflating leadership expectations with classroom reality. A principal who expects ChatGPT to boost learning is not necessarily leading a school where students actually develop stronger digital competencies, and none of the measured school conditions reliably distinguished high-CIL schools from low-CIL schools once socioeconomic and experiential factors were accounted for. The indirect pathway from leadership vision through organizational conditions, teacher practice, and student learning opportunities remains largely unmeasured, and this analysis explicitly did not test whether expectations influenced implementation or whether teacher practices mediated any association.</p>
<p>The authors also flag important limitations. The design is cross-sectional, so no causal or temporal claims are possible, and the 12 systems were defined by participation in the optional ChatGPT questionnaire rather than representative sampling of countries. The focal measures were principal and ICT coordinator reports rather than observations of actual teaching or student AI use, and included students had higher weighted mean CIL than excluded students in every system. Several systems, notably Chile, Norway, and Denmark, contributed fewer than 50 schools, widening school-level confidence intervals. These constraints mean the findings generalize only to the participating systems and samples analyzed.</p>
<p>Still, the study&#8217;s core message is a timely corrective to technological optimism. At a moment when generative AI is reshaping debates about homework, assessment, and information literacy, the largest multilevel evidence base yet assembled on principals&#8217; AI expectations and students&#8217; actual digital skills finds no common direct link between the two. What predicts assessed computer and information literacy, in these data, is not what school leaders expect of ChatGPT or their teachers, but the socioeconomic circumstances and accumulated computer experience that students bring with them. Until future research connects leadership expectations to real implementation, teacher enactment, and students&#8217; digital activities, the study suggests, expectations about AI in schools should be read as organizational commentary, not as predictors of learning.</p>
<p><strong>Subject of Research:</strong> Associations between school digital conditions and student computer and information literacy across 12 ICILS 2023 education systems</p>
<p><strong>Article Title:</strong> School digital conditions and student computer and information literacy across 12 education systems: system-specific multilevel evidence from ICILS 2023</p>
<p><strong>Article References:</strong> Chaemchoy, S., &amp; Supsin, T. (2026). School digital conditions and student computer and information literacy across 12 education systems: system-specific multilevel evidence from ICILS 2023. <em>Large-scale Assessments in Education, 14</em>(1), Article 42. <a href="https://doi.org/10.1186/s40536-026-00315-9" rel="noopener noreferrer">https://doi.org/10.1186/s40536-026-00315-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40536-026-00315-9" rel="noopener noreferrer">10.1186/s40536-026-00315-9</a></p>
<p><strong>Keywords:</strong> ICILS 2023, computer and information literacy, ChatGPT expectations, school digital conditions, ICT in education, generative AI, multilevel modeling, socioeconomic status, digital inequality, international assessment, plausible values, false discovery rate</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">195923</post-id>	</item>
	</channel>
</rss>
