<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>probability judgments &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/probability-judgments/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 01:37:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>probability judgments &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Cognitive science trick makes crowdsourced AI training data dramatically more accurate</title>
		<link>https://scienmag.com/cognitive-science-trick-makes-crowdsourced-ai-training-data-dramatically-more-accurate/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 01:37:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[biases in medical image annotation]]></category>
		<category><![CDATA[cognitive bias]]></category>
		<category><![CDATA[Cognitive science techniques for improving AI training data accuracy]]></category>
		<category><![CDATA[crowdsourced data annotation biases in machine learning]]></category>
		<category><![CDATA[crowdsourcing]]></category>
		<category><![CDATA[data annotation]]></category>
		<category><![CDATA[ethical considerations in crowdsourced AI training]]></category>
		<category><![CDATA[human cognitive constraints in data labeling]]></category>
		<category><![CDATA[human judgment]]></category>
		<category><![CDATA[impact of human biases on AI model training]]></category>
		<category><![CDATA[improving data quality in AI through cognitive science]]></category>
		<category><![CDATA[large-scale human annotation in AI development]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[methods to reduce bias in crowdsourced AI data]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[probability judgments]]></category>
		<category><![CDATA[psychology-based interventions for better data labeling]]></category>
		<category><![CDATA[recalibration]]></category>
		<category><![CDATA[satellite image labeling accuracy]]></category>
		<category><![CDATA[scalable micro-task crowdsourcing platforms for AI]]></category>
		<category><![CDATA[training data]]></category>
		<category><![CDATA[wisdom of crowds]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224914</guid>

					<description><![CDATA[Researchers show that recalibrating crowdsourced probability judgments with a cognitive model of human bias makes medical image datasets and the AI models trained on them significantly more accurate.]]></description>
										<content:encoded><![CDATA[<p>Every artificial intelligence system that reads medical images, screens airport baggage, or flags suspicious patterns in satellite photographs owes its abilities to a largely invisible workforce: the human annotators who label the training data. Because machine learning models typically need thousands or even millions of labeled examples, researchers have increasingly turned to online crowdsourcing platforms that break the annotation job into micro-tasks distributed across large pools of workers. The approach is fast and cheap, and the data-labeling industry behind it has become a multi-billion-dollar business, with companies such as Scale AI valued at 13.8 billion dollars in 2024. But there is a catch that has long troubled the field. The people doing the labeling are human, and humans bring cognitive constraints and biases to every judgment they make. Those biases do not simply vanish when the labels are aggregated; they can be inherited by the datasets and quietly embedded in the models trained on them, shaping the behavior of AI systems from the very start of the development pipeline.</p>
<p>A team of researchers led by Gunnar P. Epping and Jennifer S. Trueblood of Indiana University, working with Andrew Caplin of New York University, Erik Duhaime of Centaur Labs, and Daniel Martin of the University of California, Santa Barbara, has now demonstrated a strikingly simple fix. In a study published in Behavior Research Methods, they introduce what they call cognitive-inspired data engineering: the idea that empirical findings and models from cognitive science can be applied directly to crowdsourced judgments to correct systematic biases before the data ever reaches a machine learning algorithm. Rather than de-biasing models after training or modifying training objectives, the team intervenes upstream, at the level of the individual annotation, and shows that this single correction improves both the quality of crowdsourced datasets and the accuracy of the convolutional neural networks trained on them.</p>
<p>The core insight concerns how annotators express uncertainty. The industry-standard approach asks workers to provide a single binary label for each image, a method that discards valuable information whenever an annotator is unsure. An alternative is to elicit subjective probability judgments, asking annotators how likely they think it is that an image belongs to a given category. Probability judgments convey far more information, but they also open the door to more cognitive biases. Beyond simple response biases, where a rater favors one side of the scale, probability judgments are vulnerable to overconfidence and underconfidence, phenomena cognitive scientists have documented for decades. The hard-easy effect, for instance, leads people to be overconfident on difficult classification tasks and underconfident on easy ones. As a result, subjective probabilities rarely match the actual likelihood of being correct, a mismatch known as poor calibration. Crucially, however, these judgments remain correlated with correctness, which means the information they carry can be rescued if the biases are stripped away.</p>
<p>The researchers used a well-established cognitive model called the linear in log odds, or LLO, function to perform that rescue. The function transforms each annotator&#8217;s raw probability judgments into log-odds, fits a logistic regression against ground-truth labels from a small calibration set, and then maps the annotator&#8217;s remaining judgments onto recalibrated probabilities. The two free parameters of the function carry direct psychological interpretations. The slope parameter corrects for overconfidence or underconfidence: when its magnitude exceeds one, the function spreads judgments toward the extremes of the scale, indicating the annotator was underconfident, while a magnitude below one compresses judgments toward the center, correcting overconfidence. The intercept parameter corrects overall response tendency, shifting judgments up or down the probability scale to undo a systematic leaning toward one class. Because biases differ substantially across individuals, the researchers fit a unique recalibration function to each annotator separately, effectively normalizing every worker&#8217;s judgments before combining them into crowd labels.</p>
<p>To test the approach, the team chose a task with real clinical stakes: classifying white blood cells as blast cells or non-blast cells, a critical step in diagnosing malignant blood diseases such as leukemia and lymphoma. The stimuli were 549 digital images of Wright-stained cells from anonymized patient blood smears at Vanderbilt University Medical Center, captured with an automated cell morphology instrument and ground-truthed by three hematopathology faculty members who had to agree unanimously on each classification. The task is an ideal sandbox for studying high-skill perceptual decision-making: novices can reach roughly 65 percent accuracy with minimal training, the task remains challenging even for experts, and machine learning models can be trained on only a few hundred images rather than the thousands required for comparable problems.</p>
<p>In the first experiment, 400 Amazon Mechanical Turk workers recruited through CloudResearch judged the probability that each image contained a blast cell. After training with labeled example images and practice trials with feedback, participants completed testing blocks in which a small subset of images with known labels served as the calibration set for the LLO function. The results were revealing. Recalibration barely changed individual accuracy, which hovered around 65 percent before and after the transformation. But it dramatically improved calibration. Before correction, participants clustered their judgments at the extremes of the scale, and those extremes were badly miscalibrated: images judged to have a zero percent chance of being a blast cell were in fact blast cells roughly 28 percent of the time, while images judged to be certain blasts were truly blasts only 68 percent of the time. After recalibration, the response distribution spread across the scale and the calibration curve moved much closer to the ideal identity line. Most fitted slope parameters fell between negative one and one, confirming that the typical annotator was overconfident, and a few negative slopes even flipped the judgment ordering of participants who performed below chance.</p>
<p>The payoff appeared when individual judgments were aggregated into crowd labels. Using all available judgments, recalibration lifted crowd-label accuracy from 81.6 percent to 85.1 percent, even though no individual annotator had become more accurate. The researchers attribute this to an implicit variance-weighting mechanism: by compressing the judgments of overconfident annotators toward 50 percent, recalibration reduces their outsized influence after binarization, while stretching the judgments of underconfident annotators toward the extremes increases theirs. Recalibrated labels were also more efficient, reaching any given accuracy level with fewer judgments per image, and the benefit of calibration leveled off surprisingly quickly. Just six judgments on calibration images, three per class, captured most of the improvement, raising accuracy by 2.2 percentage points, with only a further 0.4 percentage point gain from extending to ten judgments.</p>
<p>The second experiment moved the study into a real-world setting using DiagnosUs, a crowdsourcing platform specializing in medical data annotation where skilled users, many of them medical students and healthcare professionals, compete in contests for prize money. The platform&#8217;s existing structure, in which gold-standard images with known labels are shuffled into unlabeled sets to provide feedback, supplied a natural calibration set at no extra cost. A total of 175 participants took part, split between a probability-judgment condition and an industry-standard binary-choice condition. Individual accuracy was higher than in the novice experiment, around 69 percent, and the two elicitation methods produced crowd labels of nearly identical accuracy, 88.3 percent each. Recalibration, however, pushed the probability-based crowd labels to 96.7 percent, an 8.4 percentage point gain that exceeded the 3.5 point gain seen with novices. The researchers attribute the larger benefit to the more accurate judgments available for fitting the recalibration function, not to the larger number of judgments, since gains again plateaued at roughly ten calibration judgments per annotator. Notably, the fitted intercept parameters skewed negative in this experiment, suggesting that medically trained annotators erred toward diagnosing cells as cancerous when unsure, consistent with the clinical convention that false positives are less costly than false negatives.</p>
<p>The improvements propagated directly to the machine learning models. The team trained GoogLeNet convolutional neural networks via transfer learning on crowd-labeled datasets, repeating the full training and evaluation procedure across 100 independently resampled datasets with five-fold cross-validation. At every number of judgments per image, models trained on recalibrated labels outperformed those trained on raw probability judgments or binary choices with respect to the true expert labels. Remarkably, models trained on just five recalibrated judgments per image matched the accuracy of models trained on the full set of uncalibrated judgments, a result with immediate economic implications. Because annotation cost scales with the number of judgments collected, recalibrated datasets achieve any target accuracy at a fraction of the annotations, directly attacking what the authors identify as the primary bottleneck in developing medical AI: the shortage of large, accurately labeled datasets in high-skill domains.</p>
<p>The authors are careful to frame the work as a proof of concept rather than a universal solution. The experiments used a single set of stimuli sampled with balanced class prevalence, so future work must test whether recalibration can also correct prevalence-induced errors that arise when target categories are rare, a notorious source of mistakes in visual search. The team also binarized crowd judgments before training, and training directly on continuous probability estimates may preserve even more of the uncertainty information that recalibration cleans up. Still, the broader message is hard to ignore. Cognitive scientists have spent decades cataloguing the systematic distortions in human probability judgment, and this study shows that knowledge can be converted into a simple, label-efficient transformation applicable to virtually any crowdsourcing task involving subjective probabilities. As the authors put it, the approach requires no changes to model architectures or training procedures; it simply fixes the data. In an era when the quality of AI systems is increasingly limited by the quality of human labels, teaching machines to think may begin with correcting how humans report what they think.</p>
<p><strong>Subject of Research:</strong> Using cognitive-science-based recalibration of crowdsourced probability judgments to improve AI training data quality</p>
<p><strong>Article Title:</strong> Improving crowdsourcing for AI through cognitive-inspired data engineering</p>
<p><strong>Article References:</strong> Epping, G. P., Caplin, A., Duhaime, E., Holmes, W. R., Martin, D., &amp; Trueblood, J. S. (2026). Improving crowdsourcing for AI through cognitive-inspired data engineering. <em>Behavior Research Methods, 58</em>(10), Article 283. <a href="https://doi.org/10.3758/s13428-026-03160-4" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03160-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03160-4" rel="noopener noreferrer">10.3758/s13428-026-03160-4</a></p>
<p><strong>Keywords:</strong> crowdsourcing, machine learning, cognitive bias, recalibration, medical imaging, data annotation, probability judgments, wisdom of crowds, neural networks, training data, human judgment, AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224914</post-id>	</item>
	</channel>
</rss>
