<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>psychometric evaluation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/psychometric-evaluation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 11 Sep 2026 05:33:13 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>psychometric evaluation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Studying differential item functioning in Colombia&#8217;s large scale SABER 11 tests</title>
		<link>https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Fri, 11 Sep 2026 05:33:10 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[assessment comparability across cohorts]]></category>
		<category><![CDATA[Colombian high school examinations]]></category>
		<category><![CDATA[Colombian high-school testing]]></category>
		<category><![CDATA[differential item functioning]]></category>
		<category><![CDATA[educational equity in Colombia]]></category>
		<category><![CDATA[educational equity in testing]]></category>
		<category><![CDATA[educational evaluation in Colombia]]></category>
		<category><![CDATA[fairness in standardized testing]]></category>
		<category><![CDATA[fairness testing challenges in large assessments]]></category>
		<category><![CDATA[impact of socioeconomic factors on test performance]]></category>
		<category><![CDATA[large-scale assessments in education]]></category>
		<category><![CDATA[national student assessment validity]]></category>
		<category><![CDATA[private vs public school performance]]></category>
		<category><![CDATA[psychometric analysis of test items]]></category>
		<category><![CDATA[psychometric evaluation]]></category>
		<category><![CDATA[public vs private school performance]]></category>
		<category><![CDATA[SABER 11 examination]]></category>
		<category><![CDATA[SABER 11 test analysis]]></category>
		<category><![CDATA[statistical challenges in DIF analysis]]></category>
		<category><![CDATA[statistical methods in DIF analysis]]></category>
		<category><![CDATA[test fairness and validity]]></category>
		<category><![CDATA[test form comparability]]></category>
		<guid isPermaLink="false">https://scienmag.com/studying-differential-item-functioning-in-colombias-large-scale-saber-11-tests/</guid>

					<description><![CDATA[When Colombian high-school students sit the national SABER 11 examination each year, they are taking two different versions of the test: one in March, aimed mostly at students whose academic year begins in August, and one in September for the far larger cohort whose school year starts in February. The two cohorts come from strikingly [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>When Colombian high-school students sit the national SABER 11 examination each year, they are taking two different versions of the test: one in March, aimed mostly at students whose academic year begins in August, and one in September for the far larger cohort whose school year starts in February. The two cohorts come from strikingly different educational worlds, with the March group dominated by private schools and the September group drawn overwhelmingly from public schools, so the two groups differ substantially in ability. That makes it essential to verify that the two parallel test forms are truly comparable, and a new study in the journal Large-scale Assessments in Education shows how to do it, while revealing that some of the standard statistical wisdom about fairness testing breaks down under the extreme conditions that large-scale assessments routinely create.</p>
<p>The study, conducted by John Alexander Calderón and Nelson Andrés Rodríguez of Colombia&#8217;s National Institute for Educational Evaluation (Icfes) and Víctor H. Cervantes of the University of Illinois at Urbana-Champaign, examined whether items on the Mathematics test of SABER 11 functioned differently for the two testing populations. This question is central to a psychometric property called differential item functioning, or DIF, the phenomenon in which examinees of equal underlying ability have different probabilities of answering an item correctly depending on which group they belong to. If items show DIF between groups, then score comparisons between those groups may reflect bias rather than genuine differences in the measured trait, undermining the fairness of any conclusions drawn from the results.</p>
<p>DIF has been a concern in testing since at least the 1960s, but interest intensified after the 1984 &#8220;Golden Rule&#8221; settlement in the United States, which pushed the testing industry to distinguish statistically between real group differences and bias against particular groups. Under item response theory, the framework most large-scale assessments use for scoring, DIF is defined precisely: for an item scored correct or incorrect, a two-parameter logistic model assigns each item a discrimination parameter, describing how sharply the item separates high-ability from low-ability examinees, and a difficulty parameter, locating the ability level at which an examinee has a fifty percent chance of answering correctly. If either parameter differs across groups, the item&#8217;s characteristic curves diverge. Uniform DIF corresponds to a difference in difficulty alone, producing a constant shift between the curves; non-uniform DIF involves a difference in discrimination, so the curves cross and the group advantage varies with ability; and mixed DIF involves both.</p>
<p>Because SABER 11 already uses item response theory for its scaling and scoring, the researchers chose an IRT-based DIF procedure, the non-compensatory DIF (NCDIF) index from Raju&#8217;s Differential Functioning of Items and Tests framework, complemented by the widely used Mantel–Haenszel procedure. The NCDIF index quantifies the area between the two groups&#8217; item characteristic curves, weighted by the distribution of ability in the focal group, meaning the group of special interest. This weighting has an attractive property: it emphasizes parameter differences where they matter most for the focal group&#8217;s actual scores. The statistical test on the NCDIF index is conducted through parametric bootstrap, known in the DFIT framework as item parameter replication, in which the item parameters are repeatedly re-estimated from simulated data to build the distribution of the statistic under the null hypothesis of no DIF.</p>
<p>Following a framework laid out by Sireci and Rios for tailoring DIF analyses to specific testing contexts, the team had to make a series of practical decisions: which detection method to use, how to define the comparison groups and their sample sizes, how to construct the matching variable on which the two groups are compared, whether to incorporate effect size measures, and at what level to analyze the results. Most of these choices could be settled from the existing literature. But two could not, because no published research had explored the performance of the NCDIF index under conditions that match SABER 11: sample sizes reaching roughly 40,000 examinees in the majority group against about 1,500 in the minority group, a ratio of up to 1:25, combined with a moderate gap in mean ability between the two populations.</p>
<p>To resolve this, the researchers ran a series of simulation studies in which they generated response data for test forms mirroring the real structure of SABER 11, with 40-item forms sharing an anchor of half their items, using item parameters drawn from the actual operational pool of SABER 11 mathematics items rather than from artificially clean, well-distributed parameter sets. They manipulated sample size and ratio, the impact between groups (shifting the focal group&#8217;s mean ability from zero up to plus or minus 0.8 standard deviations), the purification of the matching variable, and the use of effect size classifications, with 200 replications per condition. The purification step matters because the matching variable itself must be free of DIF contamination: the two groups&#8217; ability scales are linked using common items, and if items with DIF are included in the linking, the resulting bias can masquerade as or mask genuine DIF. The researchers used a two-stage purification in which the linking is first performed with all items, suspect items are removed, and the linking is repeated with the purified set.</p>
<p>The simulation results overturned a piece of conventional advice. Previous studies had suggested that DIF analyses work best when the two groups being compared have similar sample sizes, and it had been recommended that the larger group be subsampled to match the smaller one. But in these simulations, the Type I error rates, meaning the rate of incorrectly flagging items as showing DIF when they truly do not, were no worse for the extreme 1,500-versus-40,000 condition than for equal-sized groups of 40,000 each. In fact, the error rates remained less well controlled precisely in the conditions with equal sample sizes at the 40,000 level. The contrast between the 1,500-to-1,500 and 1,500-to-40,000 conditions showed no statistically significant difference in either Type I error or power after purification. The practical conclusion was clear: there was no reason to discard data and subsample the larger group.</p>
<p>The second major finding concerned the interplay of impact and effect sizes. Before purification, Type I error rates ballooned in the largest samples, especially when there was a true difference in mean ability between the groups, an asymmetry that also depended on whether the focal group was favored or disadvantaged. After purification, the NCDIF index&#8217;s error rates approached nominal levels, but the inflation caused by impact remained. Crucially, applying effect size guidelines, which classify flagged items by the practical magnitude of the difference rather than statistical significance alone, drove the false positive rates to nearly zero across almost all conditions, while barely reducing power for detecting moderate or large DIF given the enormous sample sizes involved. For the NCDIF procedure, power was near-perfect for items with genuinely non-negligible DIF. An analysis of variance confirmed that nearly all interactions among the experimental factors were significant, with the largest effects tied to the interplay of sample size ratio, the number of DIF items, and the use of effect sizes.</p>
<p>With these design decisions settled, the team applied the full protocol to actual SABER 11 data: the Mathematics test form administered on the second date of 2018, with 39,377 examinees as the reference group, and the form from the first date of 2019, with 1,508 examinees as the focal group. After purification, the item parameter replication test flagged eight to ten of the 22 common items, and the Mantel–Haenszel procedure flagged three to four. But when the effect size classifications were applied, only a single item, item 18, was classified as showing non-negligible DIF. Its NCDIF value was 0.03096 and its Mantel–Haenszel delta was 1.6386, placing it in the &#8220;large DIF&#8221; category. Once this item was removed from the common set, the Stocking–Lord scale linking between the two forms yielded the transformation constants needed to place both groups&#8217; abilities on a single scale, with the residual mean difference of 0.59 favoring the August-cohort group, consistent with the historical advantage of that subpopulation.</p>
<p>The content review of item 18 proved instructive. The item asks students to judge whether three different algebraic procedures for solving the equation (x + 2)(x + 3) = 5(x + 3) were performed correctly, with each student&#8217;s work shown step by step. The researchers found that although two of the procedures were on track toward the correct solution, none of them actually reached a final answer, and the point at which the student Nelson&#8217;s work stopped could appear incorrect to examinees from the September cohort more frequently than to those from the March cohort. For an examinee of ability 1.0 on the scale where the March group&#8217;s abilities were standardized, the probability of a correct response was roughly 0.62 for the March group but below 0.35 for the September group. The item&#8217;s characteristic curves crossed at an ability value of about −0.824, meaning the difference was small near the average of the focal group but grew rapidly with ability. Icfes&#8217;s mathematics team has since revised the item, making the three procedures more explicit and ensuring each reaches a final answer.</p>
<p>The study&#8217;s implications extend well beyond Colombia. The authors emphasize that detecting DIF is only the first step toward fairness, and that content review, think-aloud protocols with students, teacher interviews, and curricular analyses should follow to understand why an item functions differently. They also caution that future simulation studies of DIF statistics, whether NCDIF, Mantel–Haenszel, or any other index, should draw item parameters from realistic operational pools rather than sanitized, evenly distributed sets, since the distribution of item difficulties within a test interacts with detection behavior in ways that clean simulated pools fail to capture. In their data, the average difficulty of common items was shifted by 0.66 relative to non-common items, an asymmetry that may itself interact with detection rates. For practitioners running large-scale assessments anywhere in the world, the message is that statistical significance alone, with samples in the tens of thousands, can manufacture bias where none exists, and that effect size classification, scale purification, and realistic item parameter pools are the practical safeguards against turning fairness checks into false alarms.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Differential item functioning analysis in large-scale assessments, applied to the Mathematics test of Colombia&#8217;s SABER 11 examination</p>
<p><strong>Article Title:</strong> Differential item functioning analysis in large scale assessments: a case study for DIF in SABER 11</p>
<p><strong>Article References:</strong> Calderón, J. A., Rodríguez, N. A., &amp; Cervantes, V. H. (2026). Differential item functioning analysis in large scale assessments: a case study for DIF in SABER 11. <em>Large-scale Assessments in Education, 14</em>(1), Article 23. <a href="https://doi.org/10.1186/s40536-026-00294-x" target="_blank" rel="noopener noreferrer">https://doi.org/10.1186/s40536-026-00294-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40536-026-00294-x" target="_blank" rel="noopener noreferrer">10.1186/s40536-026-00294-x</a></p>
<p><strong>Keywords:</strong> differential item functioning, DIF, large scale assessments, SABER 11, item response theory, NCDIF index, Mantel–Haenszel, test fairness, effect size, scale purification, validity evidence, psychometrics</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">192428</post-id>	</item>
		<item>
		<title>Cross-Cultural Value Measures: PVQ Studies in Turkey</title>
		<link>https://scienmag.com/cross-cultural-value-measures-pvq-studies-in-turkey/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 13 Dec 2025 03:21:20 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[cross-cultural psychology]]></category>
		<category><![CDATA[cultural interconnectivity and values]]></category>
		<category><![CDATA[Eastern Western cultural dynamics]]></category>
		<category><![CDATA[globalization and values]]></category>
		<category><![CDATA[human values measurement]]></category>
		<category><![CDATA[methodology in cultural research]]></category>
		<category><![CDATA[Portrait Values Questionnaire]]></category>
		<category><![CDATA[psychometric evaluation]]></category>
		<category><![CDATA[psychometrics in non-Western contexts]]></category>
		<category><![CDATA[PVQ studies in Turkey]]></category>
		<category><![CDATA[value frameworks in sociology]]></category>
		<guid isPermaLink="false">https://scienmag.com/cross-cultural-value-measures-pvq-studies-in-turkey/</guid>

					<description><![CDATA[In an era marked by unprecedented globalization and cultural interconnectivity, understanding human values across diverse societies has become more critical than ever. The recent study by Karadag, Ergin-Kocaturk, and Nasir published in BMC Psychology delves into this intricate landscape by evaluating the efficacy and applicability of three well-established value measurement instruments—the Portrait Values Questionnaire Revised [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In an era marked by unprecedented globalization and cultural interconnectivity, understanding human values across diverse societies has become more critical than ever. The recent study by Karadag, Ergin-Kocaturk, and Nasir published in <em>BMC Psychology</em> delves into this intricate landscape by evaluating the efficacy and applicability of three well-established value measurement instruments—the Portrait Values Questionnaire Revised (PVQ-RR), PVQ-40, and PVQ-21—specifically within the sociocultural context of Turkey. This research stands at the intersection of cross-cultural psychology and psychometrics, offering profound insights into how value frameworks translate amidst varying societal norms, beliefs, and traditions.</p>
<p>Values, as fundamental guiding principles in human behavior and cognition, serve as anchors for decision-making, social interactions, and individual life trajectories. However, retrieving reliable, valid measurements of these abstract constructs in non-Western cultures presents a significant methodological challenge. The PVQ series, developed within Schwartz&#8217;s theory of basic human values, has been a cornerstone in this endeavor, yet its application remains under-explored in diverse cultural fabrics like Turkey’s, which blends Eastern tradition with Western modernity. This study embarks on a rigorous psychometric evaluation to fill that critical gap.</p>
<p>At the core of this work is the comparative analysis of three different versions of the Portrait Values Questionnaire. The PVQ-40, a 40-item tool, offers a comprehensive, nuanced mapping of Schwartz&#8217;s ten value types, enabling detailed exploration of subtle nuances in value priorities. The PVQ-21 condenses these items into a more manageable format, potentially enhancing participant engagement and data collection efficiency. Meanwhile, the recently revised PVQ-RR, an augmented and theoretically refined iteration, promises heightened sensitivity and clarity in value distinctions.</p>
<p>The researchers employed a methodologically robust design, drawing upon a demographically diverse Turkish sample. Given Turkey’s unique geopolitical position bridging Europe and Asia, with a rich tapestry of cultural, religious, and social influences, the dataset provides fertile ground for psychometric scrutiny. The study incorporated thorough statistical analyses including confirmatory factor analysis (CFA), internal consistency assessment via Cronbach’s alpha, and measurement invariance testing to discern whether the instruments maintained construct equivalence across subpopulations within Turkey.</p>
<p>One of the landmark findings of this investigation lies in the demonstration that while all three instruments generally replicate the theoretical value structures proposed by Schwartz, significant variations emerge in their psychometric robustness and cultural fit. The PVQ-RR notably surfaces as the instrument with superior model fit indices, underscoring its enhanced sensitivity to cultural semantics intrinsic to Turkish society. This suggests the revision process successfully tuned the tool for detecting and distinguishing value dimensions within non-Western contexts that frequently challenge Western-developed scales.</p>
<p>Delving deeper, the study highlights subtle cultural variations that affect specific value priorities. For example, notions tied to collectivism versus individualism manifest distinctly in response patterns, reflecting Turkey&#8217;s collectivist social fabric influenced by family-centric and community-oriented values. The flexibility and adaptability of the PVQ-RR allowed for better accommodation of such cultural idiosyncrasies without sacrificing theoretical coherence. Conversely, the more concise PVQ-21, while easier to administer, showed limitations in capturing these nuanced contrasts, potentially leading to oversimplification in cross-cultural value assessments.</p>
<p>Moreover, the research underscores the critical importance of measurement invariance – the degree to which a psychological construct is assessed equivalently across groups. The PVQ-RR emerged superior in achieving partial invariance across gender and age groups in Turkey, reassuring researchers of its robustness and utility for comparative studies within the population. These findings carry profound implications for both theoretical psychology and applied social research, suggesting that culturally sensitive tools can sharpen the precision of value measurement and interpretation.</p>
<p>Beyond technical psychometric evaluations, this work also contributes to broader theoretical implications concerning the universality of value structures. While Schwartz’s theory enjoys widespread acceptance, this study demonstrates that cultural calibration is indispensable. Universal frameworks must be flexibly interpreted and adapted to each societal context to maintain analytic integrity and avoid ethnocentric biases. Thus, advances in culturally attuned measurement instruments like PVQ-RR pave the way for more equitable and meaningful cross-cultural comparisons.</p>
<p>This research further informs policy-making, organizational behavior, and intercultural communications by clarifying which value facets resonate most strongly within Turkish society. For multinational corporations, NGOs, and governmental bodies operating in or partnering with Turkey, these insights provide actionable intelligence to tailor messaging strategies, motivate personnel, and craft culturally consonant interventions. Understanding value priorities enables the design of programs that better align with intrinsic cultural motivations, thereby enhancing effectiveness and social acceptance.</p>
<p>The study also calls attention to the evolving nature of human values in an increasingly interconnected world. By providing a replicable evaluation framework applicable to other non-Western populations, it encourages future research to adopt similar methodological rigor. Capturing the fluid dynamics of value change over time—amid globalization, technological diffusion, and sociopolitical transformations—requires instruments that are both theoretically grounded and culturally attuned, a balance well embodied by the PVQ-RR.</p>
<p>Intriguingly, this study opens avenues for further exploration into the psychological underpinnings of culture-specific value expressions and their neurological correlates. Future interdisciplinary collaborations might probe how cultural context modulates neural circuits linked to value-based decision-making, employing tools like functional MRI alongside sophisticated psychometric instruments. Such integrative endeavors promise a more holistic comprehension of human values bridging psychology, neuroscience, and anthropology.</p>
<p>Methodologically, the researchers emphasize the necessity of ongoing refinement and validation cycles. Psychological measurement is far from static; evolving cultural contexts, language nuances, and emergent sociocultural phenomena require instruments to remain adaptable and receptive to feedback from diverse populations. The demonstrated superiority of the PVQ-RR in this study exemplifies the benefits of iterative enhancement rooted in empirical data and theoretical rigor.</p>
<p>Ethically, this research underscores the imperative for cross-cultural psychologists to engage deeply with the populations they study, fostering mutual respect and collaborative dialogue. Developing and validating measurement tools should move beyond surface translation towards meaningful cultural immersion and consultation to safeguard against misinterpretation and cultural bias. These ethical commitments enrich scientific validity while honoring participants’ lived experiences.</p>
<p>In conclusion, the study by Karadag, Ergin-Kocaturk, and Nasir represents a landmark contribution to cross-cultural value psychology. By rigorously assessing three prominent value measurement instruments within Turkey, it offers compelling evidence favoring the enhanced PVQ-RR as a culturally sensitive, psychometrically robust tool capable of capturing the rich, multifaceted nature of human values in a uniquely complex society. This research not only advances methodological sophistication but also enriches our collective understanding of how values shape and are shaped by cultural contexts.</p>
<p>As societies evolve and interconnect at ever-accelerating rates, reliable and culturally aligned measurements of values will remain foundational for research, policy, and practice worldwide. The pioneering work presented here sets critical new standards and inspires continued innovation toward harmonizing global psychological constructs with local cultural realities—heralding a future where science genuinely embraces and interprets the diversity of human value systems.</p>
<hr />
<p><strong>Subject of Research</strong>:<br />
Cross-cultural evaluation of human value measurement instruments in Turkey using PVQ-RR, PVQ-40, and PVQ-21.</p>
<p><strong>Article Title</strong>:<br />
Evaluating value measures across cultures: study on PVQ-RR, PVQ-40, and PVQ-21 in Turkey.</p>
<p><strong>Article References</strong>:<br />
Karadag, E., Ergin-Kocaturk, H. &amp; Nasir, S. Evaluating value measures across cultures: study on PVQ-RR, PVQ-40, and PVQ-21 in Turkey. <em>BMC Psychology</em> (2025). <a href="https://doi.org/10.1186/s40359-025-03725-6">https://doi.org/10.1186/s40359-025-03725-6</a></p>
<p><strong>Image Credits</strong>:<br />
AI Generated</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">116929</post-id>	</item>
	</channel>
</rss>
