<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Prolific &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/prolific/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 21:18:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Prolific &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Bayesian Reasoning Problems Could Expose AI Bots Hiding in Online Surveys</title>
		<link>https://scienmag.com/bayesian-reasoning-problems-could-expose-ai-bots-hiding-in-online-surveys/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 21:18:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI detection in crowdsourced research]]></category>
		<category><![CDATA[Bayesian reasoning]]></category>
		<category><![CDATA[Bayesian reasoning in online survey validation]]></category>
		<category><![CDATA[Bayesian reasoning problems revealing AI bots]]></category>
		<category><![CDATA[behavioral research methods]]></category>
		<category><![CDATA[capability-gap test]]></category>
		<category><![CDATA[capability-gap testing for AI identification]]></category>
		<category><![CDATA[chatbot identification in research]]></category>
		<category><![CDATA[cognitive psychology methods for AI detection]]></category>
		<category><![CDATA[crowdsourcing]]></category>
		<category><![CDATA[data contamination]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[human versus AI performance in Bayesian tasks]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in behavioral studies]]></category>
		<category><![CDATA[large language models influencing survey responses]]></category>
		<category><![CDATA[natural frequencies]]></category>
		<category><![CDATA[online behavioral research integrity]]></category>
		<category><![CDATA[online research]]></category>
		<category><![CDATA[online survey data contamination]]></category>
		<category><![CDATA[positive predictive value]]></category>
		<category><![CDATA[predictive value calculation in survey validation]]></category>
		<category><![CDATA[Prolific]]></category>
		<category><![CDATA[signal detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212579</guid>

					<description><![CDATA[A new study proposes using Bayesian reasoning problems, with their well-established human performance ceilings, as calibrated detectors of large language model contamination in online research samples.]]></description>
										<content:encoded><![CDATA[<p>Online behavioral research is facing a quiet crisis. As large language models become woven into everyday life, researchers who recruit participants through crowdsourcing platforms increasingly suspect that some of their respondents are not human at all — or are humans outsourcing their answers to chatbots. A new study published in Behavior Research Methods proposes an elegant solution to this problem, and it comes from an unexpected corner of cognitive psychology: the humble Bayesian reasoning problem, a puzzle that humans have famously struggled with for fifty years.</p>
<p>The study, conducted by independent researcher Vera Wilde, introduces what she calls a capability-gap test. The logic is deceptively simple. Certain problems, such as calculating a positive predictive value from base-rate information, have been studied so extensively in humans that scientists know, with meta-analytic precision, exactly how well people can perform. When participants in an online study dramatically exceed those well-established human ceilings, the most plausible explanation is not superhuman cognition but machine assistance — a canary in the data coalmine, signaling that the sample may be contaminated by large language models.</p>
<p>The empirical basis for the proposal comes from two preregistered pilot studies of a Bayesian reasoning training tool, with a combined sample of 148 participants recruited through the Prolific platform. The tool was designed to teach people how to solve Bayesian inference problems, the kind of task exemplified by medical diagnosis questions: given a disease with a certain prevalence, a test with a certain sensitivity and false-positive rate, what is the probability that a person who tests positive actually has the disease? Decades of research, dating back to classic work by Daniel Kahneman and Amos Tversky and extended by Gerd Gigerenzer and Ulrich Hoffrage, have shown that most people fail such problems, even when the numbers are presented in natural frequency formats that make the underlying logic easier to grasp.</p>
<p>That failure is precisely what makes the task useful as a detector. A meta-analysis by McDowell and Jacobs found that only about 24 percent of people can solve a single Bayesian reasoning problem presented in natural frequency format — and that figure represents a ceiling, a level at which achieving a perfect score on a battery of five such problems is effectively unattainable for genuine human respondents. Yet in Wilde&#8217;s pilots, participants&#8217; accuracy on positive predictive value calculation problems reached roughly three times the established human performance ceiling. In the second pilot, 57 percent of participants achieved perfect 5-for-5 scores, a result that should be extraordinarily rare in an uncontaminated human sample.</p>
<p>The technical heart of the approach lies in distinguishing two outcome measures: accuracy and algorithm use. Accuracy refers simply to whether the participant produced the correct numerical answer. Algorithm use, by contrast, refers to evidence in the participant&#8217;s response that they actually followed the Bayesian reasoning process — for example, constructing a frequency tree, counting cases, or showing the intermediate steps of the calculation. This distinction matters because a training intervention designed to improve Bayesian reasoning should, if it works, change both measures in tandem. A large language model, however, can produce correct answers without any visible reasoning process, or with a reasoning process that does not respond to the training manipulation in the way human learning does.</p>
<p>By tracking both measures simultaneously, researchers can separate two rival explanations for suspiciously high performance. If accuracy spikes but algorithm use does not, contamination is the likelier culprit, because the model supplies answers without the participant acquiring the underlying skill. If both accuracy and algorithm use rise together, the pattern is consistent with authentic learning effects, and the treatment signal can be preserved rather than discarded. In this way, the capability-gap test does not merely flag bad data; it helps researchers decide which parts of their dataset reflect genuine psychological phenomena and which parts reflect machine-generated noise.</p>
<p>Wilde frames the detection problem itself as a signal detection problem, structurally analogous to the mass screenings for low-prevalence conditions — such as disease screening — around which the Bayesian reasoning literature was originally developed. Just as a medical test must balance hits against false alarms, a contamination detector must catch bot-driven responses without wrongly excluding honest participants who happen to be statistically savvy. The known reference distributions from the Bayesian reasoning literature make this calibration possible in a way that ad hoc attention checks cannot. Rather than relying on generic screening questions, researchers can compare observed performance against quantified human benchmarks and estimate the probability that a given response pattern arose from machine assistance.</p>
<p>The approach offers four practical advantages over existing data-quality tools. First, the human performance ceilings are grounded in meta-analyses rather than informal intuition, giving researchers a defensible threshold for suspicion. Second, the human–large language model performance gap on these problems is large, which increases the sensitivity of the test. Third, the known reference distributions allow nuanced assessment rather than crude pass–fail judgments. Fourth, Bayesian reasoning problems are easy to embed in existing surveys, requiring no special software or platform cooperation. Together, these properties make the method deployable at scale across the many fields — psychology, marketing, political science, epidemiology — that increasingly depend on online samples.</p>
<p>The stakes are considerable. Prior research on crowd work has documented substantial and growing use of large language models by online workers, and studies of data contamination in machine learning itself show how memorized content can masquerade as genuine capability. If a meaningful fraction of respondents in an online study are completing tasks with chatbot help, effect sizes may be distorted, replication attempts may fail for reasons that have nothing to do with the underlying science, and the credibility of entire literatures built on crowdsourced data could be undermined. The problem echoes an older statistical concern: John Tukey&#8217;s foundational work on sampling from contaminated distributions warned that even small amounts of contamination can seriously mislead inference drawn from nominally clean data.</p>
<p>Wilde is careful to note the provenance of the idea: neither pilot study was originally designed to validate a contamination detection method, and the proposal emerged from post hoc analysis of unexpectedly strong results. That origin makes the capability-gap test a promising hypothesis rather than a fully validated diagnostic, and the author provides practical recommendations for researchers who wish to use these problems as data-quality diagnostics while the validation literature matures. Both studies were preregistered on the Open Science Framework, and all data, materials, and analysis code are publicly available, allowing other teams to scrutinize and extend the approach. If the method holds up under broader testing, Bayesian reasoning problems — long a symbol of human statistical frailty — may find a second career as guardians of scientific integrity, ensuring that the data feeding behavioral science come from human minds rather than the machines trained on those minds&#8217; collective output.</p>
<p><strong>Subject of Research:</strong> Detecting large language model contamination in online behavioral research samples using Bayesian reasoning problems as capability-gap tests</p>
<p><strong>Article Title:</strong> An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples</p>
<p><strong>Article References:</strong> Wilde, V. (2026). An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples. <em>Behavior Research Methods, 58</em>(11), Article 301. <a href="https://doi.org/10.3758/s13428-026-03184-w" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03184-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03184-w" rel="noopener noreferrer">10.3758/s13428-026-03184-w</a></p>
<p><strong>Keywords:</strong> large language models, Bayesian reasoning, data contamination, online research, data quality, signal detection, natural frequencies, positive predictive value, crowdsourcing, behavioral research methods, capability-gap test, Prolific</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212579</post-id>	</item>
		<item>
		<title>Stress Rewires Money Choices Even Among High Earners, Study Finds</title>
		<link>https://scienmag.com/stress-rewires-money-choices-even-among-high-earners-study-finds/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 13:10:40 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[behavioral economics]]></category>
		<category><![CDATA[decision-making]]></category>
		<category><![CDATA[delay discounting]]></category>
		<category><![CDATA[effects of stress on time and risk perception]]></category>
		<category><![CDATA[gender differences]]></category>
		<category><![CDATA[health behavior]]></category>
		<category><![CDATA[high-income earners]]></category>
		<category><![CDATA[impact of stress on monetary choices]]></category>
		<category><![CDATA[impulsive decision-making]]></category>
		<category><![CDATA[impulsivity]]></category>
		<category><![CDATA[income]]></category>
		<category><![CDATA[Journal of Behavioral Medicine]]></category>
		<category><![CDATA[perceived stress]]></category>
		<category><![CDATA[probability discounting]]></category>
		<category><![CDATA[Prolific]]></category>
		<category><![CDATA[psychological factors in financial behavior]]></category>
		<category><![CDATA[reward magnitude]]></category>
		<category><![CDATA[socioeconomic status and financial behavior]]></category>
		<category><![CDATA[stress and financial decision-making]]></category>
		<category><![CDATA[stress and reward valuation]]></category>
		<category><![CDATA[stress-induced changes in risk assessment]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=194743</guid>

					<description><![CDATA[A study of high-income US adults found that greater perceived stress was associated with steeper delay and probability discounting for small rewards, suggesting the stress-decision link exists independently of socioeconomic disadvantage.]]></description>
										<content:encoded><![CDATA[<p>When people feel stressed, their relationship with time and certainty appears to change in measurable ways, and a new study suggests that this phenomenon is not simply a byproduct of financial hardship. Researchers report that among adults earning at least $100,000 a year, higher levels of perceived stress were linked to steeper devaluation of delayed and uncertain monetary rewards, at least when the stakes were modest. The findings, published in the Journal of Behavioral Medicine, offer some of the clearest evidence yet that the psychological weight of stress shapes impulsive decision-making independently of socioeconomic disadvantage, a confound that has clouded this area of research for decades.</p>
<p>The study focused on two fundamental processes in behavioral economics. The first, delay discounting, describes how the perceived value of a reward shrinks as its arrival is pushed into the future. A person with high delay discounting might take fifty dollars today rather than one hundred dollars in six months, effectively trading a larger outcome for immediate gratification. The second, probability discounting, captures how people devalue rewards that are uncertain. Offered a guaranteed fifty dollars or a hundred dollars with only a sixty percent chance of payout, someone with high probability discounting would gravitate toward the smaller, certain option. Both processes have been tied to a striking range of maladaptive outcomes, including tobacco and other substance use, poor adherence to preventive medical care and treatment regimens, and psychiatric conditions such as major depressive and bipolar disorders.</p>
<p>Perceived stress, the psychological variable at the center of the new work, is distinct from a single bad day or an acute laboratory stressor. It reflects the degree to which people experience their lives as unpredictable, uncontrollable, or overwhelming over an extended period, capturing the cumulative burden of multiple stressors. Earlier observational studies consistently found that adults and adolescents reporting greater perceived stress also showed steeper delay discounting, and similar patterns appeared in people recovering from substance use disorders. Yet experimental studies that deliberately induced acute stress, such as through public speaking or cold exposure, generally failed to reproduce the effect, leaving researchers with a puzzling inconsistency. More importantly, both perceived stress and delay discounting are reliably associated with low income and scarcity, making it difficult to tell whether stress genuinely alters decision-making or whether both simply track poverty.</p>
<p>To break that deadlock, the research team, led by Mary J. King of Virginia Tech alongside Michelle Rockwell, John Epling, and Jeffrey S. Stein, turned to a methodological strategy borrowed from epidemiology: restricting the sample. Instead of statistically adjusting for income after the fact, they recruited only participants whose combined household income was at least $100,000 per year, placing them roughly within the top two income quintiles in the United States. By draining most of the variability out of the income variable, the researchers effectively neutralized socioeconomic disadvantage as an alternative explanation. Participants were recruited through the online crowdsourcing platform Prolific and were further limited to adults aged 28 to 75 who worked outside medicine and related fields, because the study doubled as a comparison group for a separate investigation of primary care clinicians.</p>
<p>A total of 247 participants completed the survey in August 2024, with a median completion time of just under six minutes. After excluding ten participants who failed attention checks and fifteen who reported incomes below the pre-screening threshold, 222 participants remained for analysis. The sample had a median age of 48 years, a slight female majority, and a median Perceived Stress Scale score of 14, corresponding to moderate stress. The Perceived Stress Scale, a widely used ten-item instrument asking how often respondents felt nervous, overwhelmed, or in control over the past month, showed excellent internal consistency in this sample.</p>
<p>Discounting was measured with brief, six-trial adjusting tasks administered for two reward magnitudes, $100 and $10,000, in randomized order. In the delay discounting task, participants repeatedly chose between a smaller immediate amount and a larger delayed amount, with the delay adjusted up or down across trials until the point of indifference, expressed as an ED50, was reached. Its inverse, the rate parameter k, indexes how steeply delayed rewards lose value. The probability discounting task worked analogously, adjusting the chance of receiving a larger immediate reward until indifference, yielding the rate parameter h. Because smaller rewards are typically discounted more steeply over time while larger rewards are discounted more steeply under uncertainty, testing both magnitudes provided an internal validity check: the classic magnitude effects were indeed replicated, confirming the tasks behaved as expected.</p>
<p>The results were strikingly selective. For the $100 rewards, perceived stress significantly predicted steeper delay discounting in both univariate and multiple regression models that controlled for age, education, gender, and household income. The same held true for probability discounting at $100, where higher stress scores were associated with greater devaluation of uncertain rewards. At the $10,000 magnitude, however, perceived stress lost its association with both delay and probability discounting entirely. The researchers suggest this magnitude-specific pattern may reflect the more extreme discounting rates observed for large rewards, which can compress the range of individual differences and reduce sensitivity to modest cross-sectional effects. Notably, the association between stress and probability discounting at $100 did not fully survive a sensitivity analysis that retained outlier participants, a caveat the authors acknowledge.</p>
<p>One demographic factor proved more durable than stress. Across both reward magnitudes, female participants showed higher probability discounting than male participants, preferring smaller, guaranteed rewards over larger, uncertain ones. This aligns with prior findings of gender differences in decision-making under uncertainty, although the literature on gender and delay discounting remains less consistent. The persistence of the gender effect at magnitudes where the stress effect vanished underscores that different psychological and demographic influences on risky choice may operate at different scales of reward.</p>
<p>The study&#8217;s limitations are worth noting. The cross-sectional design precludes any claim about causality or temporal ordering, and the sample was drawn entirely from a research volunteer platform and restricted to non-healthcare workers with high incomes, limiting generalizability. The brief probability discounting task has received less validation than its delay counterpart, though its successful replication of the magnitude effect lends credibility. The authors also caution that their analyses did not adjust for multiple comparisons, and they call for future longitudinal work across diverse populations, income levels, and geographic regions, where cost of living may translate identical salaries into very different financial realities.</p>
<p>Even with those caveats, the implications are substantial. Health-related decisions, from smoking cessation to medication adherence, hinge on how individuals weigh immediate certainty against delayed or probabilistic benefit. If the everyday experience of stress nudges people toward impatience and risk aversion even in financially secure populations, then stress management could represent an underappreciated target for interventions designed to improve long-term health behavior. What this study establishes is that the well-documented link between stress and impulsive choice is not merely poverty wearing a psychological disguise; it persists where economic hardship is largely absent, provided the rewards on the table are small enough to feel like everyday money.</p>
<p><strong>Subject of Research:</strong> The relationship between perceived stress and delay and probability discounting in a high-income adult sample</p>
<p><strong>Article Title:</strong> Examining relationships between delay and probability discounting and perceived stress in a high-income online sample</p>
<p><strong>Article References:</strong> King, M. J., Rockwell, M., Epling, J., &amp; Stein, J. S. (2026). Examining relationships between delay and probability discounting and perceived stress in a high-income online sample. <em>Journal of Behavioral Medicine</em>. <a href="https://doi.org/10.1007/s10865-026-00711-0" rel="noopener noreferrer">https://doi.org/10.1007/s10865-026-00711-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10865-026-00711-0" rel="noopener noreferrer">10.1007/s10865-026-00711-0</a></p>
<p><strong>Keywords:</strong> delay discounting, probability discounting, perceived stress, decision-making, behavioral economics, impulsivity, income, Prolific, health behavior, Journal of Behavioral Medicine, reward magnitude, gender differences</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">194743</post-id>	</item>
	</channel>
</rss>
