<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Schwartz Value Theory &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/schwartz-value-theory/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 16:36:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Schwartz Value Theory &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Best–Worst Scaling Beats Ranking for Reliable Preference Measurement, Study Finds</title>
		<link>https://scienmag.com/best-worst-scaling-beats-ranking-for-reliable-preference-measurement-study-finds/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 16:36:14 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[Behavior Research Methods]]></category>
		<category><![CDATA[behavioral economics]]></category>
		<category><![CDATA[best-worst scaling]]></category>
		<category><![CDATA[decision-making]]></category>
		<category><![CDATA[Kendall's Tau]]></category>
		<category><![CDATA[life values]]></category>
		<category><![CDATA[preference elicitation]]></category>
		<category><![CDATA[psychological measurement]]></category>
		<category><![CDATA[ranking]]></category>
		<category><![CDATA[Schwartz Value Theory]]></category>
		<category><![CDATA[survey methods]]></category>
		<category><![CDATA[test-retest reliability]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217222</guid>

					<description><![CDATA[New experiments show that best–worst scaling produces more consistent and more informative measurements of human preferences than traditional ranking tasks, even over intervals of just minutes.]]></description>
										<content:encoded><![CDATA[<p>Ask a family to rank the best musical decade of all time and you will quickly discover how unstable human preferences can be. The uncle who swears by The Beatles may reverse half his list within the hour, while the middle positions—albums nobody loves or hates—shuffle endlessly. That familiar dinner-table chaos turns out to be a genuine scientific problem. Researchers who study how people choose between health care options, consumer products, or political candidates depend on measuring preferences accurately, and a new study published in Behavior Research Methods suggests that the most popular measurement tool may be the least reliable one.</p>
<p>Garston Liang of the University of Technology Sydney, together with Mackenzie Glover and Guy E. Hawkins of the University of Newcastle, compared two widely used methods of preference elicitation: ranking and best–worst scaling, often abbreviated as BWS. In a ranking task, respondents order a full list of items from most to least preferred, a format familiar from hiring panels, consumer surveys, and even Australia&#8217;s preferential voting system. In best–worst scaling, by contrast, respondents see only small subsets of items at a time and simply pick the one they favor most and the one they favor least. Across three experiments involving hundreds of participants, the team found that both methods largely agree on the ordering of preferences, but that BWS delivers markedly better test–retest reliability—meaning people&#8217;s answers stay more consistent when measured repeatedly, even over intervals as short as a few minutes.</p>
<p>The theoretical stakes of this comparison are considerable. One long-standing view in psychology and behavioral economics assumes that people hold stable internal representations of their preferences, which measurement merely reveals, with any variability attributable to the tool being used. A competing perspective, known as the constructed preferences view, holds that preferences are assembled on the spot, packaged with the conditions of the moment—the available options, the framing of the question, even the mood of the respondent. Under that view, elicited preferences need not be coherent over time at all. Both perspectives agree, however, that the choice of elicitation method matters enormously, which makes the question of whether ranking and BWS actually measure the same thing a critical one for anyone designing surveys.</p>
<p>Each method carries distinct strengths and weaknesses. Ranking is quick, intuitive, and computationally cheap, and it works well when only a handful of options are involved. But its demands grow brutally with set size: ordering a top-ten list requires the equivalent of 45 pairwise comparisons, and marketing researchers routinely confront sets of dozens of competing products. Rankings also force respondents to discriminate between options they may genuinely feel indifferent about, manufacturing distinctions where none exist. BWS sidesteps both problems. Because people are naturally good at identifying what they love and what they loathe, asking for best and worst choices within small subsets keeps cognitive load low regardless of how many items are in play. Over repeated, carefully balanced choice sets, preferences for every item accumulate incrementally, and each choice constrains the possible orderings of items not even shown in that round.</p>
<p>To compare the two methods fairly, the researchers needed a domain where true preferences are unlikely to shift during the experiment. They settled on life values, using a validated set of ten values drawn from Schwartz&#8217;s Value Theory—concepts such as benevolence, hedonism, power, security, and self-direction. Fundamental changes in personal values unfold over years, dwarfing the minutes required for a survey, so any inconsistency observed between two measurements could be attributed to the measurement process itself rather than to genuine change in what people believe. The team also manipulated how the values were labeled: some participants saw broad dimension labels like &#8220;benevolence,&#8221; while others saw example item labels like &#8220;helpful, honest, forgiving.&#8221; Because both label types target the same underlying concept, switching labels between measurements provided an additional stress test of each method&#8217;s stability.</p>
<p>The experimental design was elegant in its simplicity. In Experiment 1a, 101 participants completed the ranking task twice; in Experiment 1b, 100 participants completed the BWS task twice; and in Experiment 2, 148 participants completed both tasks in randomized order. Participants were recruited through the online platform Prolific, restricted to native English speakers with high approval ratings, and paid at a rate commensurate with £9.00 per hour. The BWS task used a balanced pairwise design in which each of the ten life values appeared exactly six times across eleven subsets of five or six options, with any pair of values co-occurring exactly three times. Agreement between measurement occasions was quantified using Kendall&#8217;s Tau, a rank-correlation statistic bounded between −1 and 1, where a value of 1 indicates perfect agreement between two ordered lists.</p>
<p>The headline result was unambiguous. Average individual-level agreement between repeated measurements was substantially higher for BWS than for ranking, with mean Kendall&#8217;s Tau values of 0.69 versus 0.47—a difference supported by Bayes factors exceeding 1000, indicating decisive statistical evidence. The pattern of errors in the ranking task revealed exactly why. The most unstable position in the ranked list was rank six, smack in the middle: 81 percent of rank-six assignments changed between measurements. Consistency improved monotonically toward the extremes, with rank one changing for only 54 percent of responses. BWS, by capitalizing on people&#8217;s reliable judgments at the extremes, effectively avoids forcing respondents to manufacture precision in the muddled middle—precisely where rankings are weakest.</p>
<p>Importantly, the two methods were not measuring different things. The aggregate preference orderings agreed almost perfectly: benevolence emerged as the most important value and power as the least important in both tasks, and the top five values were ordered identically across methods, with only a single pair of middling values trading positions. The BWS scores, however, carried extra information that rankings cannot provide. Because normalized best–worst scores range continuously from −1 to 1, they quantify the strength of preferences, not just their order. In Experiment 2, the score gaps between first and second place and between ninth and tenth place were roughly 0.40 and 0.38, dwarfing the average gap of about 0.11 between adjacent middle ranks—confirming that middle-ranked options were nearly indistinguishable in preference strength. This added resolution has real-world value: in a prior oncology study cited by the authors, BWS revealed that surgeon training was twice as important to patients as surgeon experience and more than three times as important as hospital reputation.</p>
<p>The label manipulation produced a subtle but instructive finding. When the labels changed between measurement occasions—say, from &#8220;benevolence&#8221; to &#8220;helpful, honest, forgiving&#8221;—agreement dropped sharply for both methods, indicating that the two label sets evoke genuinely different interpretations of the same value. Yet the direction and even the relative magnitude of those interpretive shifts were consistent across both elicitation methods: benevolence was preferred as an item label and security as a dimension label, to similar degrees, whether measured by ranking or by BWS. This convergence suggests that the reliability advantage of BWS stems from the method itself rather than from any accidental alignment between the method and the particular stimulus set used.</p>
<p>The practical implications extend well beyond survey methodology. BWS is not free—completion times in the study ran about 8.6 minutes for the BWS task versus 4.6 minutes for ranking, because larger item sets require more choice rounds even though each round stays cognitively light. But the authors argue the added reliability is worth the cost, particularly in applied settings such as health care priority-setting, marketing research, and personality assessment, where item sets of 16 to 57 options make full rankings unwieldy. BWS also naturally captures ties, revealing when respondents genuinely hold two options equal, something a forced ranking can never show. For researchers who have long treated the humble ranked list as the default instrument for measuring what people want, the message of this study is clear: if you care whether your measurements would come out the same way twice, asking people only for their best and their worst may be the most honest question you can ask.</p>
<p><strong>Subject of Research:</strong> Comparing the test–retest reliability of best–worst scaling and ranking methods for eliciting personal preferences</p>
<p><strong>Article Title:</strong> Time after time, the best–worst scaling method is more reliable than ranking</p>
<p><strong>Article References:</strong> Liang, G., Glover, M., &amp; Hawkins, G. E. (2026). Time after time, the best–worst scaling method is more reliable than ranking. <em>Behavior Research Methods, 58</em>(10), Article 284. <a href="https://doi.org/10.3758/s13428-026-03167-x" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03167-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03167-x" rel="noopener noreferrer">10.3758/s13428-026-03167-x</a></p>
<p><strong>Keywords:</strong> best–worst scaling, ranking, preference elicitation, test–retest reliability, survey methods, life values, Schwartz Value Theory, Kendall&#x27;s Tau, decision making, behavioral economics, psychological measurement, Behavior Research Methods</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217222</post-id>	</item>
	</channel>
</rss>
