<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>publication bias in meta-analyses of educational interventions &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/publication-bias-in-meta-analyses-of-educational-interventions/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 05 Sep 2026 12:10:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>publication bias in meta-analyses of educational interventions &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Rigor Replaces RAMsing in STEM Education Meta-Analysis Methods</title>
		<link>https://scienmag.com/rigor-replaces-ramsing-in-stem-education-meta-analysis-methods/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Sat, 05 Sep 2026 12:10:31 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[challenges in synthesizing STEM education research]]></category>
		<category><![CDATA[consequences of inflated findings in STEM education policy]]></category>
		<category><![CDATA[effects of moderator analyses on evidence validity]]></category>
		<category><![CDATA[effects of publication bias on STEM education]]></category>
		<category><![CDATA[evidence distortion in education policy]]></category>
		<category><![CDATA[impact of RAMSing in educational research]]></category>
		<category><![CDATA[impact of RAMSing on STEM intervention studies]]></category>
		<category><![CDATA[improving accuracy of educational meta-analyses]]></category>
		<category><![CDATA[improving reliability of meta-analytic results in]]></category>
		<category><![CDATA[influence of selective reporting on STEM intervention effectiveness]]></category>
		<category><![CDATA[influence of significant moderators on research findings]]></category>
		<category><![CDATA[meta-analytic methodology in STEM]]></category>
		<category><![CDATA[methodological issues in STEM education meta-analyses]]></category>
		<category><![CDATA[moderator effect size inflation]]></category>
		<category><![CDATA[moderator effect size reporting bias]]></category>
		<category><![CDATA[publication bias in meta-analyses of educational interventions]]></category>
		<category><![CDATA[reporting bias in educational research]]></category>
		<category><![CDATA[selective reporting in meta-analyses]]></category>
		<category><![CDATA[statistical practices in educational psychology]]></category>
		<category><![CDATA[statistical significance and reporting practices in education studies]]></category>
		<category><![CDATA[STEM education meta-analysis]]></category>
		<category><![CDATA[transparency and accuracy in educational research synthesis]]></category>
		<category><![CDATA[transparency in education research synthesis]]></category>
		<guid isPermaLink="false">https://scienmag.com/rigor-replaces-ramsing-in-stem-education-meta-analysis-methods/</guid>

					<description><![CDATA[Meta-analyses of STEM education interventions have been quietly inflating their most celebrated findings, according to a new study that names and quantifies a pervasive reporting practice the authors call RAMSing—reporting after moderators are significant. The research, published in Educational Psychology Review, analyzed nearly 1,800 moderator effect sizes drawn from 89 meta-analyses and found that statistically [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Meta-analyses of STEM education interventions have been quietly inflating their most celebrated findings, according to a new study that names and quantifies a pervasive reporting practice the authors call RAMSing—reporting after moderators are significant. The research, published in Educational Psychology Review, analyzed nearly 1,800 moderator effect sizes drawn from 89 meta-analyses and found that statistically non-significant moderator results may be only a fraction as likely to appear in the published record as their significant counterparts, distorting the evidence that educators, policymakers, and researchers rely on when deciding what works, for whom, and under what conditions.</p>
<p>Moderator analyses are the workhorses of research synthesis in education. When a meta-analysis pools dozens or hundreds of studies testing an intervention—say, a new instructional model or an educational technology tool—the individual study results rarely agree perfectly. Moderators, tested through subgroup analyses or meta-regressions, are meant to explain that between-study heterogeneity by asking whether effects vary with study characteristics such as instructional approach, sample size, student age, disciplinary focus, or assessment format. In theory, they transform a single average effect into a nuanced map of boundary conditions. In practice, the new study argues, they are vulnerable to a form of selective reporting that operates inside the meta-analysis itself.</p>
<p>The team, led by Mehmet Bicakci of Friedrich Alexander University Erlangen-Nürnberg together with Fabian Heller, Heidrun Stoeger of the University of Regensburg, and Albert Ziegler, defines RAMSing as the practice of running multiple moderator tests but publishing only those that reach conventional significance thresholds such as p &lt; .05, or that exceed familiar effect-size benchmarks like Cohen&#8217;s d of 0.20 or 0.50. The practice is distinct from p-hacking and HARKing because it occurs after the analysis, at the reporting stage, and leaves little trace of the broader analytic search that produced the surviving results. During full-text screening of more than 100 STEM education meta-analyses, the researchers repeatedly encountered warning signs: moderators announced in introductions but missing from results sections, inconsistencies in summary tables, and vague statements such as &#8220;no other moderators were significant&#8221; unaccompanied by any statistics.</p>
<p>To measure the scope of the problem, the team conducted an umbrella review, systematically searching ERIC, PsycINFO, Web of Science, and PSYNDEX for the period 2000 through 2023. After screening 5,103 unique records with inter-rater reliabilities between 0.79 and 0.95, they assembled a final dataset of 1,786 moderator effect sizes from 89 meta-analyses of STEM education interventions. The corpus was dominated by mathematics research, which accounted for nearly half of the corpus, followed by multidisciplinary STEM interventions and science-specific studies. Academic outcomes appeared in over 88 percent of cases, and instructional models plus educational technology were the most common intervention categories.</p>
<p>The centerpiece of the analysis was a pair of bias-robust Bayesian frameworks known as RoBMA-PSMA and RoBMA-Regression. Rather than assuming the published literature is complete, these methods model publication selection explicitly, averaging across dozens of candidate models that weight findings according to their p-values and adjust for small-study effects. The results were stark. The models estimated that moderator effect sizes with p-values between 0.05 and 0.50 were only 43 percent as likely to appear in the published record as highly significant results—a relative visibility deficit of roughly 57 percent. For p-values between 0.50 and 1.00, the estimated selection weight fell to 0.11, meaning clearly non-significant moderator findings were about 89 percent less likely to be visible. The Bayes factor for publication-bias-type selection reached an overwhelming 4.72 × 10³⁶, and the overall visibility gap across non-significant results was summarized at approximately ω ≈ 0.40.</p>
<p>Despite the bias, the moderator effects themselves were not an illusion. The bias-adjusted mean moderator effect was g = 0.27, with a 95 percent credible interval of 0.20 to 0.34, supported by a Bayes factor of nearly three million. But the effects were strikingly heterogeneous, with a between-effect standard deviation of τ = 0.48. When the researchers classified moderators into five broad themes—intervention specifications, sample and context variables, disciplinary focus, study design and quality, and publication metadata—the classification substantially improved model fit with a Bayes factor of 350 and modestly reduced residual heterogeneity to τ = 0.44. Disciplinary focus produced the largest expected effects at g = 0.438, followed by sample and context variables at g = 0.364, while design and quality indicators yielded the smallest at g = 0.197. The finding suggests that collapsing science, technology, engineering, and mathematics into a single undifferentiated category risks obscuring meaningful differences across fields.</p>
<p>The bias also contaminated the yardsticks by which effect sizes are judged. The raw published distribution of moderator effects, right-skewed and leptokurtic with a mean of 0.59, yielded empirical quartiles of 0.29, 0.53, and 0.77. But when the team generated 10,000 posterior-predictive draws from the publication-bias-adjusted model, the quartiles dropped to 0.18, 0.38, and 0.64. Notably, visible discontinuities in the raw histogram clustered around Cohen&#8217;s conventional thresholds of 0.20, 0.50, and 0.80, and these jumps disappeared entirely after bias adjustment—raising the possibility that the thresholds function not only as interpretive guides but as implicit reporting incentives. The authors suggest that some moderator findings may be selectively reported not just for statistical significance but for appearing large enough to satisfy magnitude conventions, a secondary form of RAMSing that miscalibrates what a field considers &#8220;small,&#8221; &#8220;typical,&#8221; or &#8220;large.&#8221;</p>
<p>Power analysis added a further sobering dimension. Under a planning scenario using the bias-adjusted median effect of g = 0.38, heterogeneity of τ² = 0.24, and a two-sided test at α = 0.05, detecting a typical moderator effect with 80 percent power requires approximately 14 independent effect sizes per moderator level. Yet about 60 percent of the moderator units in the published corpus fell below that threshold, indicating that underpowered moderator tests are the norm rather than the exception. Low power produces unstable estimates that are especially vulnerable to selective reporting, because apparently interesting findings persist in the literature while null results drift toward the file drawer.</p>
<p>The study&#8217;s remedy is a four-component &#8220;anti-RAMSing toolkit,&#8221; openly available on the Open Science Framework as fully annotated R scripts with a worked example dataset. The workflow begins with structural mapping—a qualitative cataloguing of every intervention-moderator-outcome combination in a literature, visualized as an alluvial diagram—followed by robust Bayesian bias adjustment, bias-aware benchmarking, and prospective power calibration. The authors also propose a 13-item checklist covering moderator justification, preregistration, adequate effect sizes per level, complete reporting of all tested moderators including null results, and sensitivity analyses. Among their practical recommendations: journals could require structured tables listing all prespecified moderators, and reporting guidelines such as PRISMA could add a RAMSing checkpoint demanding disclosure of every moderator test with its accompanying statistics.</p>
<p>The implications reach well beyond STEM education. The authors argue that the vulnerabilities they document—extensive moderator probing, uneven reporting standards, and incentive structures that reward novelty over completeness—are likely wherever meta-analysts probe many study characteristics, including psychology, management research, and the health sciences. Indeed, prior meta-research suggests only about half of all moderator analyses ever reach publication. The deeper problem, the study contends, is the absence of a defensible framework for selecting, testing, and reporting moderators, which leaves the practice partly arbitrary and invites selective disclosure.</p>
<p>The researchers caution that their bias adjustments are model-based corrections rather than direct observations of the complete evidential record, and that they depend on the reporting quality of the published studies themselves—many moderator effects had to be excluded because variances, standard errors, or confidence intervals were missing. Had precision information been consistently reported, the dataset would have been roughly 25 percent larger. Still, sensitivity analyses across narrower and wider prior settings left the substantive conclusions essentially unchanged.</p>
<p>The study&#8217;s central message is that the credibility of moderator evidence depends not only on statistical technique but on what the literature chooses to reveal. Selective visibility, the authors conclude, can inflate apparent moderator effects, encourage overconfident claims about what works for whom, and bias the effect-size estimates used for power planning and research investment decisions. A more credible future for meta-analysis, they argue, requires not more moderator tests, but better ones—preregistered where possible, grounded in theory, powered adequately, and reported completely regardless of outcome.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Selective reporting bias (RAMSing) in moderator analyses within STEM education meta-research</p>
<p><strong>Article Title:</strong> Rigor Replaces RAMsing in STEM Education Meta-Analysis Methods</p>
<p><strong>Article References:</strong> Bicakci, M., Heller, F., Stoeger, H., &amp; Ziegler, A. (2026). From RAMSing to Rigor: Improving Moderator Analysis in STEM Education Meta-Research. <em>Educational Psychology Review, 38</em>(1), Article 75. <a href="https://doi.org/10.1007/s10648-026-10172-1" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s10648-026-10172-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10648-026-10172-1" target="_blank" rel="noopener noreferrer">10.1007/s10648-026-10172-1</a></p>
<p><strong>Keywords:</strong> challenges in synthesizing STEM education research, consequences of inflated findings in STEM education policy, effects of moderator analyses on evidence validity, impact of RAMSing in educational research, improving reliability of meta-analytic results in, influence of selective reporting on STEM intervention effectiveness, methodological issues in STEM education meta-analyses, moderator effect size reporting bias, publication bias in meta-analyses of educational interventions, statistical significance and reporting practices in education studies, STEM education meta-analysis, transparency and accuracy in educational research synthesis</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">187988</post-id>	</item>
	</channel>
</rss>
