<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>wisdom of crowds &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/wisdom-of-crowds/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 03 Oct 2026 01:27:57 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>wisdom of crowds &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Crowds Track Human Moral Judgment by Voting, Not Talking</title>
		<link>https://scienmag.com/ai-crowds-track-human-moral-judgment-by-voting-not-talking/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 03 Oct 2026 01:27:57 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI alignment]]></category>
		<category><![CDATA[AI collective decision-making]]></category>
		<category><![CDATA[AI crowds]]></category>
		<category><![CDATA[AI ensemble decision accuracy]]></category>
		<category><![CDATA[AI voting vs. conversation]]></category>
		<category><![CDATA[benchmark alignment]]></category>
		<category><![CDATA[collective decision-making]]></category>
		<category><![CDATA[collective intelligence in machine ethics]]></category>
		<category><![CDATA[ensemble aggregation]]></category>
		<category><![CDATA[ethical AI research methodologies]]></category>
		<category><![CDATA[false consensus]]></category>
		<category><![CDATA[human-majority moral benchmarks]]></category>
		<category><![CDATA[language models for moral judgment]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large-scale ethical scenario evaluation]]></category>
		<category><![CDATA[machine morality simulation]]></category>
		<category><![CDATA[memory sharing]]></category>
		<category><![CDATA[moral decision-making in AI]]></category>
		<category><![CDATA[moral judgment]]></category>
		<category><![CDATA[multi-agent systems]]></category>
		<category><![CDATA[nuclear crisis moral judgments]]></category>
		<category><![CDATA[nuclear crisis simulation]]></category>
		<category><![CDATA[perspectival diversity]]></category>
		<category><![CDATA[wisdom of crowds]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=230007</guid>

					<description><![CDATA[A massive new study finds that groups of AI agents align with majority human moral judgment through independent voting rather than information sharing, but at the cost of erasing legitimate human disagreement into false consensus.]]></description>
										<content:encoded><![CDATA[<p>When large language models are asked to make moral decisions, does putting many of them in a room make the group wiser? A new study published in AI &amp; Society suggests that the answer is yes, but not for the reasons many researchers assumed. The work, carried out by June Christoph Kang of Korea University and the Empathy Research Institute, borrows its name from the Tachikoma, the small, curious AI tanks of the anime series Ghost in the Shell: Stand Alone Complex. In the show, these agents share identical hardware yet develop distinct moral personalities through divergent experiences, periodically syncing their memories. The study asks whether a similar architecture, many independent agents whose judgments are pooled, brings machine collectives closer to the moral judgment of the human majority than any single model can manage.</p>
<p>To answer that question, Kang ran an unusually large experiment: more than 1.5 million evaluation runs across three comparably sized language models, covering 7,591 scenarios drawn from three validated moral-judgment benchmarks, ETHICS, Scruples, and the Moral Machine dataset, plus an expanded set of 42 nuclear-crisis scenarios. The key metric is benchmark alignment, defined as the agreement between a collective&#8217;s confidence-weighted majority decision and the human-majority reference label on each scenario. The author is careful to frame this as an empirical proxy for aggregated human moral judgment, not as a direct measure of anything as grand as the common good. Still, the metric offers a rigorous way to test whether groups of machines converge on what most people consider right.</p>
<p>Four main findings emerge from the data. First, the Tachikoma effect is real but small: pooling independent agents improves alignment with human majorities by at most four to five percentage points, with an effect size of roughly eta-squared 0.002, and the benefit varies by model. Second, when the analysis was re-estimated using scenario-clustered statistical models that correct for the non-independence of repeated scenarios, the effect looked less like a Condorcet-style gain from group size and more like ordinary variance reduction over a highly correlated ensemble. The mean inter-agent error correlation was about 0.8, meaning the agents tend to make the same mistakes at the same time. When members of a crowd err together, adding more members buys you far less than classical jury theorems would predict.</p>
<p>Third, and perhaps most striking for anyone hoping that diversity is the secret ingredient, a fixed-group-size homogeneity ablation showed that prompted perspectival diversity, instructing agents to adopt different viewpoints, pushes results in the expected direction but only weakly, reaching statistical significance in just one of the three models. Sharing memory between agents did not help at all. Shared memory raised consensus among the agents without improving accuracy, a pattern that echoes classic findings on informational cascades and groupthink in human groups, where communication can homogenize opinion without making it more correct. The practical implication is counterintuitive for the multi-agent AI community: letting agents deliberate and share information appears to add conformity, not wisdom.</p>
<p>The study&#8217;s fourth finding moves from ethics benchmarks to geopolitics. On the expanded set of 42 nuclear-crisis scenarios, multi-agent groups significantly de-escalated simulated crises for two of the three models tested. This result connects to a growing literature on escalation risks from language models in military and diplomatic decision-making, including work from researchers at SIPRI and elsewhere warning that frontier models can exhibit sophisticated but unpredictable reasoning in simulated conflicts. If collective architectures genuinely dampen escalatory tendencies, that would be a safety-relevant dividend of aggregation, though the effect was not uniform across models and the scenarios remain simulations rather than real command-and-control environments.</p>
<p>Underlying all of this is a methodological point that matters for how such results should be read. Much of the apparent benefit of collectives evaporates or shrinks once the non-independence of benchmark scenarios is properly modeled. Because the same scenarios are evaluated many times across agent counts and configurations, treating each evaluation as an independent observation inflates statistical confidence. By re-estimating with scenario-clustered models, the study shows that the Tachikoma effect is best understood as aggregation and variance reduction over a correlated ensemble, not as a magical property of group size. This is a cautionary tale for the broader field of multi-agent LLM research, where dramatic claims about debate and deliberation improving reasoning have sometimes rested on fragile statistical footing.</p>
<p>But the most unsettling finding concerns what alignment costs. Collectives reliably track the human majority, yet they systematically under-represent legitimate human disagreement. On scenarios where humans are nearly evenly split, the AI collectives returned unanimous verdicts 73 to 87 percent of the time. In other words, the very mechanism that makes groups look aligned, confidence-weighted majority voting, collapses genuine moral disagreement into false consensus. Benchmark alignment improves partly by erasing minority positions. A system that reports a confident unanimous verdict on an issue where humans are 50-50 is not capturing collective wisdom; it is manufacturing certainty that does not exist. Supplementary analyses reinforce this: alignment tracks human consensus strength monotonically, and a 20.6 percentage-point alignment gap between controversial and non-controversial Scruples scenarios persists across all agent counts, confirming that collective architecture cannot resolve genuine moral ambiguity.</p>
<p>The study also found that alignment-optimized training amplifies social responsiveness, meaning models tuned with human feedback to be more agreeable and socially attuned respond more strongly to the collective setting. This interacts with known tendencies toward sycophancy in language models, where training can make systems overly eager to mirror perceived human preferences. In a multi-agent context, such social responsiveness may further push agents toward consensus, compounding the false-consensus problem. The author&#8217;s recommendation follows directly: deployed systems should aggregate independent judgments while preserving and reporting disagreement, rather than presenting a single confident group verdict that hides the distribution of views beneath it.</p>
<p>Why does the anime metaphor fit? The Tachikoma of Ghost in the Shell are beloved precisely because they combine independence with periodic synchronization, developing quirky individual perspectives through divergent experience while occasionally pooling what they have learned. The study&#8217;s results suggest the fiction got the architecture partly right and partly wrong. Independence and divergent experience do contribute something, but the synchronization step, the analogue of shared memory and deliberation, adds consensus without adding accuracy. The wisdom, such as it is, comes from the vote of independent minds, not from the conversation between them. For engineers designing multi-agent systems for morally consequential domains, from content moderation to medical triage to crisis diplomacy, the lesson is to keep agents independent, aggregate their confidence-weighted votes, and surface the disagreements rather than smoothing them away.</p>
<p>The broader significance of the work lies in reframing. Collective moral behavior in language models is not primarily a problem of information sharing, deliberation, or emergent group intelligence. It is a problem of robust, disagreement-preserving aggregation, a question social choice theory has grappled with since Condorcet, and one that thinkers from Amartya Sen to Cass Sunstein have shown is fraught when diversity of opinion is treated as noise rather than signal. As large language models are increasingly deployed in morally consequential domains, the temptation will be to tune them until they agree with the majority on everything. This study warns that doing so would trade away something essential: the honest representation of a pluralistic human moral landscape, in which reasonable people, and reasonable machines, sometimes disagree.</p>
<p><strong>Subject of Research:</strong> Collective moral decision-making in multi-agent large language model systems</p>
<p><strong>Article Title:</strong> The Tachikoma effect: aggregation of independent agents, not information sharing, aligns LLM collectives with majority human moral judgment</p>
<p><strong>Article References:</strong> Kang, J. C. (2026). The Tachikoma effect: aggregation of independent agents, not information sharing, aligns LLM collectives with majority human moral judgment. <em>AI &amp;amp; SOCIETY</em>. <a href="https://doi.org/10.1007/s00146-026-03346-6" rel="noopener noreferrer">https://doi.org/10.1007/s00146-026-03346-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00146-026-03346-6" rel="noopener noreferrer">10.1007/s00146-026-03346-6</a></p>
<p><strong>Keywords:</strong> large language models, multi-agent systems, moral judgment, AI alignment, wisdom of crowds, collective decision-making, ensemble aggregation, memory sharing, nuclear crisis simulation, false consensus, benchmark alignment, perspectival diversity</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">230007</post-id>	</item>
		<item>
		<title>Cognitive science trick makes crowdsourced AI training data dramatically more accurate</title>
		<link>https://scienmag.com/cognitive-science-trick-makes-crowdsourced-ai-training-data-dramatically-more-accurate/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 01:37:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[biases in medical image annotation]]></category>
		<category><![CDATA[cognitive bias]]></category>
		<category><![CDATA[Cognitive science techniques for improving AI training data accuracy]]></category>
		<category><![CDATA[crowdsourced data annotation biases in machine learning]]></category>
		<category><![CDATA[crowdsourcing]]></category>
		<category><![CDATA[data annotation]]></category>
		<category><![CDATA[ethical considerations in crowdsourced AI training]]></category>
		<category><![CDATA[human cognitive constraints in data labeling]]></category>
		<category><![CDATA[human judgment]]></category>
		<category><![CDATA[impact of human biases on AI model training]]></category>
		<category><![CDATA[improving data quality in AI through cognitive science]]></category>
		<category><![CDATA[large-scale human annotation in AI development]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Medical Imaging]]></category>
		<category><![CDATA[methods to reduce bias in crowdsourced AI data]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[probability judgments]]></category>
		<category><![CDATA[psychology-based interventions for better data labeling]]></category>
		<category><![CDATA[recalibration]]></category>
		<category><![CDATA[satellite image labeling accuracy]]></category>
		<category><![CDATA[scalable micro-task crowdsourcing platforms for AI]]></category>
		<category><![CDATA[training data]]></category>
		<category><![CDATA[wisdom of crowds]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224914</guid>

					<description><![CDATA[Researchers show that recalibrating crowdsourced probability judgments with a cognitive model of human bias makes medical image datasets and the AI models trained on them significantly more accurate.]]></description>
										<content:encoded><![CDATA[<p>Every artificial intelligence system that reads medical images, screens airport baggage, or flags suspicious patterns in satellite photographs owes its abilities to a largely invisible workforce: the human annotators who label the training data. Because machine learning models typically need thousands or even millions of labeled examples, researchers have increasingly turned to online crowdsourcing platforms that break the annotation job into micro-tasks distributed across large pools of workers. The approach is fast and cheap, and the data-labeling industry behind it has become a multi-billion-dollar business, with companies such as Scale AI valued at 13.8 billion dollars in 2024. But there is a catch that has long troubled the field. The people doing the labeling are human, and humans bring cognitive constraints and biases to every judgment they make. Those biases do not simply vanish when the labels are aggregated; they can be inherited by the datasets and quietly embedded in the models trained on them, shaping the behavior of AI systems from the very start of the development pipeline.</p>
<p>A team of researchers led by Gunnar P. Epping and Jennifer S. Trueblood of Indiana University, working with Andrew Caplin of New York University, Erik Duhaime of Centaur Labs, and Daniel Martin of the University of California, Santa Barbara, has now demonstrated a strikingly simple fix. In a study published in Behavior Research Methods, they introduce what they call cognitive-inspired data engineering: the idea that empirical findings and models from cognitive science can be applied directly to crowdsourced judgments to correct systematic biases before the data ever reaches a machine learning algorithm. Rather than de-biasing models after training or modifying training objectives, the team intervenes upstream, at the level of the individual annotation, and shows that this single correction improves both the quality of crowdsourced datasets and the accuracy of the convolutional neural networks trained on them.</p>
<p>The core insight concerns how annotators express uncertainty. The industry-standard approach asks workers to provide a single binary label for each image, a method that discards valuable information whenever an annotator is unsure. An alternative is to elicit subjective probability judgments, asking annotators how likely they think it is that an image belongs to a given category. Probability judgments convey far more information, but they also open the door to more cognitive biases. Beyond simple response biases, where a rater favors one side of the scale, probability judgments are vulnerable to overconfidence and underconfidence, phenomena cognitive scientists have documented for decades. The hard-easy effect, for instance, leads people to be overconfident on difficult classification tasks and underconfident on easy ones. As a result, subjective probabilities rarely match the actual likelihood of being correct, a mismatch known as poor calibration. Crucially, however, these judgments remain correlated with correctness, which means the information they carry can be rescued if the biases are stripped away.</p>
<p>The researchers used a well-established cognitive model called the linear in log odds, or LLO, function to perform that rescue. The function transforms each annotator&#8217;s raw probability judgments into log-odds, fits a logistic regression against ground-truth labels from a small calibration set, and then maps the annotator&#8217;s remaining judgments onto recalibrated probabilities. The two free parameters of the function carry direct psychological interpretations. The slope parameter corrects for overconfidence or underconfidence: when its magnitude exceeds one, the function spreads judgments toward the extremes of the scale, indicating the annotator was underconfident, while a magnitude below one compresses judgments toward the center, correcting overconfidence. The intercept parameter corrects overall response tendency, shifting judgments up or down the probability scale to undo a systematic leaning toward one class. Because biases differ substantially across individuals, the researchers fit a unique recalibration function to each annotator separately, effectively normalizing every worker&#8217;s judgments before combining them into crowd labels.</p>
<p>To test the approach, the team chose a task with real clinical stakes: classifying white blood cells as blast cells or non-blast cells, a critical step in diagnosing malignant blood diseases such as leukemia and lymphoma. The stimuli were 549 digital images of Wright-stained cells from anonymized patient blood smears at Vanderbilt University Medical Center, captured with an automated cell morphology instrument and ground-truthed by three hematopathology faculty members who had to agree unanimously on each classification. The task is an ideal sandbox for studying high-skill perceptual decision-making: novices can reach roughly 65 percent accuracy with minimal training, the task remains challenging even for experts, and machine learning models can be trained on only a few hundred images rather than the thousands required for comparable problems.</p>
<p>In the first experiment, 400 Amazon Mechanical Turk workers recruited through CloudResearch judged the probability that each image contained a blast cell. After training with labeled example images and practice trials with feedback, participants completed testing blocks in which a small subset of images with known labels served as the calibration set for the LLO function. The results were revealing. Recalibration barely changed individual accuracy, which hovered around 65 percent before and after the transformation. But it dramatically improved calibration. Before correction, participants clustered their judgments at the extremes of the scale, and those extremes were badly miscalibrated: images judged to have a zero percent chance of being a blast cell were in fact blast cells roughly 28 percent of the time, while images judged to be certain blasts were truly blasts only 68 percent of the time. After recalibration, the response distribution spread across the scale and the calibration curve moved much closer to the ideal identity line. Most fitted slope parameters fell between negative one and one, confirming that the typical annotator was overconfident, and a few negative slopes even flipped the judgment ordering of participants who performed below chance.</p>
<p>The payoff appeared when individual judgments were aggregated into crowd labels. Using all available judgments, recalibration lifted crowd-label accuracy from 81.6 percent to 85.1 percent, even though no individual annotator had become more accurate. The researchers attribute this to an implicit variance-weighting mechanism: by compressing the judgments of overconfident annotators toward 50 percent, recalibration reduces their outsized influence after binarization, while stretching the judgments of underconfident annotators toward the extremes increases theirs. Recalibrated labels were also more efficient, reaching any given accuracy level with fewer judgments per image, and the benefit of calibration leveled off surprisingly quickly. Just six judgments on calibration images, three per class, captured most of the improvement, raising accuracy by 2.2 percentage points, with only a further 0.4 percentage point gain from extending to ten judgments.</p>
<p>The second experiment moved the study into a real-world setting using DiagnosUs, a crowdsourcing platform specializing in medical data annotation where skilled users, many of them medical students and healthcare professionals, compete in contests for prize money. The platform&#8217;s existing structure, in which gold-standard images with known labels are shuffled into unlabeled sets to provide feedback, supplied a natural calibration set at no extra cost. A total of 175 participants took part, split between a probability-judgment condition and an industry-standard binary-choice condition. Individual accuracy was higher than in the novice experiment, around 69 percent, and the two elicitation methods produced crowd labels of nearly identical accuracy, 88.3 percent each. Recalibration, however, pushed the probability-based crowd labels to 96.7 percent, an 8.4 percentage point gain that exceeded the 3.5 point gain seen with novices. The researchers attribute the larger benefit to the more accurate judgments available for fitting the recalibration function, not to the larger number of judgments, since gains again plateaued at roughly ten calibration judgments per annotator. Notably, the fitted intercept parameters skewed negative in this experiment, suggesting that medically trained annotators erred toward diagnosing cells as cancerous when unsure, consistent with the clinical convention that false positives are less costly than false negatives.</p>
<p>The improvements propagated directly to the machine learning models. The team trained GoogLeNet convolutional neural networks via transfer learning on crowd-labeled datasets, repeating the full training and evaluation procedure across 100 independently resampled datasets with five-fold cross-validation. At every number of judgments per image, models trained on recalibrated labels outperformed those trained on raw probability judgments or binary choices with respect to the true expert labels. Remarkably, models trained on just five recalibrated judgments per image matched the accuracy of models trained on the full set of uncalibrated judgments, a result with immediate economic implications. Because annotation cost scales with the number of judgments collected, recalibrated datasets achieve any target accuracy at a fraction of the annotations, directly attacking what the authors identify as the primary bottleneck in developing medical AI: the shortage of large, accurately labeled datasets in high-skill domains.</p>
<p>The authors are careful to frame the work as a proof of concept rather than a universal solution. The experiments used a single set of stimuli sampled with balanced class prevalence, so future work must test whether recalibration can also correct prevalence-induced errors that arise when target categories are rare, a notorious source of mistakes in visual search. The team also binarized crowd judgments before training, and training directly on continuous probability estimates may preserve even more of the uncertainty information that recalibration cleans up. Still, the broader message is hard to ignore. Cognitive scientists have spent decades cataloguing the systematic distortions in human probability judgment, and this study shows that knowledge can be converted into a simple, label-efficient transformation applicable to virtually any crowdsourcing task involving subjective probabilities. As the authors put it, the approach requires no changes to model architectures or training procedures; it simply fixes the data. In an era when the quality of AI systems is increasingly limited by the quality of human labels, teaching machines to think may begin with correcting how humans report what they think.</p>
<p><strong>Subject of Research:</strong> Using cognitive-science-based recalibration of crowdsourced probability judgments to improve AI training data quality</p>
<p><strong>Article Title:</strong> Improving crowdsourcing for AI through cognitive-inspired data engineering</p>
<p><strong>Article References:</strong> Epping, G. P., Caplin, A., Duhaime, E., Holmes, W. R., Martin, D., &amp; Trueblood, J. S. (2026). Improving crowdsourcing for AI through cognitive-inspired data engineering. <em>Behavior Research Methods, 58</em>(10), Article 283. <a href="https://doi.org/10.3758/s13428-026-03160-4" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03160-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03160-4" rel="noopener noreferrer">10.3758/s13428-026-03160-4</a></p>
<p><strong>Keywords:</strong> crowdsourcing, machine learning, cognitive bias, recalibration, medical imaging, data annotation, probability judgments, wisdom of crowds, neural networks, training data, human judgment, AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224914</post-id>	</item>
		<item>
		<title>Algorithmic Monoculture May Not Be So Bad, MIT Study Finds</title>
		<link>https://scienmag.com/algorithmic-monoculture-may-not-be-so-bad-mit-study-finds/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:39:19 +0000</pubDate>
				<category><![CDATA[Mathematics]]></category>
		<category><![CDATA[AI decision-making in recruitment]]></category>
		<category><![CDATA[AI hiring algorithms]]></category>
		<category><![CDATA[algorithmic bias]]></category>
		<category><![CDATA[algorithmic monoculture]]></category>
		<category><![CDATA[algorithmic monoculture in industry]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[automated decision-making]]></category>
		<category><![CDATA[automation in hiring processes]]></category>
		<category><![CDATA[consequences of single-algorithm reliance]]></category>
		<category><![CDATA[diversity of AI tools in industry]]></category>
		<category><![CDATA[echo chambers]]></category>
		<category><![CDATA[effects of algorithm deployment in employment]]></category>
		<category><![CDATA[ensemble algorithms]]></category>
		<category><![CDATA[ethical implications of algorithmic monopolies]]></category>
		<category><![CDATA[hiring algorithms]]></category>
		<category><![CDATA[impact of automated resume screening]]></category>
		<category><![CDATA[labor markets]]></category>
		<category><![CDATA[Mit]]></category>
		<category><![CDATA[MIT research on AI decision-making]]></category>
		<category><![CDATA[Philosophical Perspectives]]></category>
		<category><![CDATA[philosophical perspectives on AI algorithms]]></category>
		<category><![CDATA[resume screening]]></category>
		<category><![CDATA[risks of uniform AI systems]]></category>
		<category><![CDATA[wisdom of crowds]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224326</guid>

					<description><![CDATA[MIT researchers mathematically show that the harms of algorithmic monoculture in hiring depend on design details, with ensemble algorithms potentially outperforming diverse systems.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence systems are quietly taking over decisions that were once made by human beings, and few areas illustrate this shift more clearly than hiring. Resume screening algorithms now sit at the front door of many large employers, filtering applications with a speed and consistency that no human recruiter could match. As these tools proliferate, a growing chorus of scholars has warned about a scenario they call algorithmic monoculture: a future in which a single algorithm, or near-identical copies of it, makes essentially all the decisions in a given industry. The fear is intuitive. If one company&#8217;s algorithm rejects your resume, and every other company relies on the same system, you could find yourself shut out of an entire job market by a single automated judgment. But new research from the Massachusetts Institute of Technology suggests that this widely shared anxiety may be overstated, and that the real effects of algorithmic monoculture depend heavily on the details of how such systems are deployed.</p>
<p>The study, published in the journal Philosophical Perspectives, was conducted by Brian Hedden, a professor in MIT&#8217;s Department of Linguistics and Philosophy who holds a shared position with the Schwarzman College of Computing and the Department of Electrical Engineering and Computer Science, and Manish Raghavan, the Drew Houston Career Development Professor at the MIT Sloan School of Management and in EECS. Both are principal investigators in the Laboratory for Information and Decision Systems. Rather than treating algorithmic monoculture as a monolithic threat, the pair systematically evaluated the major objections that have been raised against it, testing each one against formal models that capture multiple hiring scenarios. Their conclusion is striking: many of the standard arguments against monoculture either fail outright or are not decisive against all of its forms, while the most serious drawback turns out to be something quite different from what critics have emphasized.</p>
<p>The researchers begin by confronting the objection that has dominated the conversation: systematic exclusion. The worry is that when firms share the same screening algorithm, a candidate rejected by one employer will be rejected by all of them, compounding a single algorithmic mistake into a career-defining barrier. Yet when Hedden and Raghavan modeled this situation across a series of scenarios, they found the argument less compelling than it first appears. The total number of people hired, they show, is not affected by whether firms use the same algorithm or different ones. All the jobs still get filled, and the same number of people end up employed. What changes is the distribution of outcomes and, intriguingly, the balance of power in the labor market. Because firms using an identical algorithm end up competing over the same pool of approved candidates, job seekers may actually gain bargaining leverage, which can drive wages up. In this sense, monoculture could paradoxically strengthen the position of workers rather than weaken it.</p>
<p>The authors then turned to objections rooted in individual agency. Consider a scenario in which a job candidate submits an application and their resume is automatically forwarded to every firm using the same hiring algorithm. In that case, the candidate never gets a chance to learn from early rejections and revise their materials before the next round. This loss of feedback and iteration seems like a genuine harm, and Hedden acknowledges it as a strong objection to certain bad forms of monoculture. But the harm, he argues, is not intrinsic to monoculture itself. If the shared system allows candidates to revise and resubmit their applications, the objection loses its force. The design of the platform, not the mere fact of algorithmic uniformity, determines whether applicants retain meaningful control over how they present themselves to the market.</p>
<p>A related concern involves gaming. When everyone knows that a particular resume format produces better results under a particular algorithm, job seekers have an obvious incentive to reformat their documents to exploit that knowledge. Critics worry that a monocultural system would invite widespread manipulation. Hedden counters that this is not obviously true. If many different firms use many different algorithms, a strategic applicant might simply target a couple of those systems and game them, gaining a small advantage with a few employers while leaving the rest untouched. The incentive to game, in other words, does not disappear under polyculture; it merely fragments. Whether uniformity amplifies manipulation or concentrates it into a single, more transparent target is an empirical question, not a settled point in favor of diversity.</p>
<p>The most substantive objection the researchers identified concerns information. Drawing on the wisdom of crowds, a well-established idea from social psychology holding that a diverse group of independent decision makers can outperform any single individual, Hedden explains that firms equipped with different hiring algorithms can collectively assemble a higher-quality pool of new hires. Monoculture threatens this advantage by creating what the researchers describe as informational echo chambers. When every firm consults the same algorithm, candidates with the same characteristics and credentials get hired every time by every firm, and potentially superior alternatives go undiscovered. The researchers prove this tendency mathematically: monoculture tends to suppress the exploration that diversity of judgment would otherwise provide. In hiring, this could mean the best candidates are less likely to find jobs, not because they are excluded outright, but because no one is looking in the places where their talents might reveal themselves.</p>
<p>Raghavan notes that this discovery problem may matter more in some domains than others. It is not clear whether reduced discovery is always a bad thing, he says, but it is definitely a worry when designing AI for applications like science, art, or writing, where the value of an output often lies in its novelty. The researchers are careful to scope their claims accordingly. Their analysis focuses on hiring, with lending as a natural parallel, since credit decisions were once made by independent bankers but now rest on standardized FICO credit scores derived from a single algorithm. A handful of resume screening tools are similarly common across Fortune 500 companies. But they caution that other domains, such as generative AI content creation or AI-guided scientific research, may behave differently, and that monoculture in those settings could prove considerably more problematic.</p>
<p>Crucially, the researchers also show that the echo chamber problem has a potential technical fix. One approach is to build randomness into a monocultural platform, inducing a higher level of exploration and preventing the system from converging on the same candidates every time. Another, more surprising remedy is to lean further into monoculture rather than away from it. By packaging multiple firms&#8217; hiring algorithms into a single ensemble algorithm, which assigns each candidate a score based on an average across the constituent models, the drawbacks of any one system can be smoothed out. Through a series of simulations of different hiring situations, the researchers confirmed that such an ensemble can sometimes outperform the use of multiple independent algorithms, and can even allow a monoculture to perform as well as, or better than, a polyculture in which every firm runs its own distinct system. The performance of monoculture, they emphasize, also depends on the accuracy of the underlying algorithm: if one algorithm is substantially more accurate than the many alternatives, uniformity may simply be the better bet.</p>
<p>Several open questions remain before these findings can guide policy or practice. Hedden notes that it is not yet clear how feasible algorithmic ensembling would be in the real world, where competing firms may be reluctant to share or combine their proprietary screening tools. Raghavan adds that many of the answers about the promises and pitfalls of algorithmic monoculture are going to be contextual, and that substantial empirical work is still needed to understand how these concerns play out in actual markets. The theoretical models capture the structure of the problem, but a job market is a complex system with frictions, asymmetries, and human behaviors that no simulation fully reproduces. The researchers hope their work will inspire further study of the long-term consequences of algorithmic uniformity, as well as investigations into the real-world complexities that determine whether shared algorithms help or harm the people they evaluate.</p>
<p>The broader lesson of the study is a call for nuance in a debate that has often been framed in absolutes. A trend toward algorithmic monoculture is a realistic scenario and an important consequence of AI adoption, Hedden says, but it is hard to say in the abstract whether monoculture would be a bad thing; it depends on the details, including the domain in question and the accuracy of the algorithm itself. For job seekers, the findings offer a measure of reassurance that the nightmare of universal automated rejection is not an inevitable feature of shared algorithms, and even a hint that common systems could raise wages by intensifying competition among employers. For the architects of AI systems, the message is more constructive: the harms of monoculture are design problems, not destiny, and mechanisms like ensembles and built-in randomness may preserve the efficiency of shared tools while recovering the exploratory benefits of diverse judgment.</p>
<p><strong>Subject of Research:</strong> The effects of algorithmic monoculture in automated hiring decisions</p>
<p><strong>Article Title:</strong> The effects of an “algorithmic monoculture” depend on the details</p>
<p><strong>Article References:</strong> The effects of an “algorithmic monoculture” depend on the details. (n.d.). <a href="https://www.eurekalert.org/news-releases/1146252" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> algorithmic monoculture, hiring algorithms, artificial intelligence, MIT, ensemble algorithms, resume screening, wisdom of crowds, labor markets, algorithmic bias, automated decision-making, Philosophical Perspectives, echo chambers</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224326</post-id>	</item>
		<item>
		<title>Breaking Problems Apart Helps Deliberative Crowds Make Wiser Decisions</title>
		<link>https://scienmag.com/breaking-problems-apart-helps-deliberative-crowds-make-wiser-decisions/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 05 Aug 2026 19:45:25 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[cognitive load reduction]]></category>
		<category><![CDATA[collaborative reasoning]]></category>
		<category><![CDATA[collective intelligence strategies]]></category>
		<category><![CDATA[collective problem-solving]]></category>
		<category><![CDATA[complex problem analysis]]></category>
		<category><![CDATA[decision accuracy enhancement]]></category>
		<category><![CDATA[deliberative group decisions]]></category>
		<category><![CDATA[group decision-making]]></category>
		<category><![CDATA[group discussion dynamics]]></category>
		<category><![CDATA[problem decomposition in crowds]]></category>
		<category><![CDATA[subgroup analysis in decision making]]></category>
		<category><![CDATA[wisdom of crowds]]></category>
		<guid isPermaLink="false">https://scienmag.com/breaking-problems-apart-helps-deliberative-crowds-make-wiser-decisions/</guid>

					<description><![CDATA[A new study in Nature Communications is drawing attention to a deceptively simple idea with potentially far-reaching consequences: groups may make better decisions when they first divide a complex problem into smaller, more manageable parts. Researchers Francisco Barrera-Lemarchand, Valentina Lescano-Charreau, Juan Ruiz and colleagues report that collective problem decomposition can improve the “wisdom” of deliberative [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>A new study in <em>Nature Communications</em> is drawing attention to a deceptively simple idea with potentially far-reaching consequences: groups may make better decisions when they first divide a complex problem into smaller, more manageable parts. Researchers Francisco Barrera-Lemarchand, Valentina Lescano-Charreau, Juan Ruiz and colleagues report that collective problem decomposition can improve the “wisdom” of deliberative crowds, offering a possible way to make group reasoning more accurate without requiring every participant to become an expert.</p>
<p>The finding addresses a long-standing question in collective intelligence. Crowds can sometimes outperform individuals because different people bring different information, intuitions and analytical strategies to the same problem. Yet group discussion can also produce the opposite result. Participants may anchor on an early suggestion, follow confident voices, repeat shared assumptions or converge on an attractive but incorrect answer. The challenge is therefore not simply to gather more opinions, but to organize interaction so that useful diversity survives the discussion.</p>
<p>Problem decomposition provides one possible solution. Instead of asking a group to solve a complicated question in a single step, the method separates it into distinct subproblems. Participants can then analyze specific components before their insights are recombined into a collective answer. In technical terms, decomposition reduces the cognitive dimensionality of the task: a large decision space is transformed into a set of smaller spaces that may be easier to evaluate, compare and aggregate.</p>
<p>The approach is especially relevant to deliberative crowds, in which people do more than submit independent estimates. They communicate, exchange arguments and revise their views. Deliberation can improve performance when discussion reveals new evidence or corrects mistakes, but it can also create correlated errors. Once participants influence one another, their judgments may become less independent, weakening one of the mechanisms behind the classic wisdom-of-crowds effect. Structuring the problem before discussion may help preserve complementary perspectives while still allowing information to circulate.</p>
<p>The researchers’ central contribution is to connect the architecture of a task with the quality of collective reasoning. A crowd is not an unchanging source of intelligence; its performance depends on how questions are framed, how information is shared and how individual contributions are combined. By assigning attention to separate components of a problem, decomposition may prevent participants from competing over a single vague conclusion and instead encourage them to contribute specialized pieces of analysis.</p>
<p>This distinction matters because many real-world questions are not single questions at all. Assessing a public-health intervention, forecasting an economic outcome, evaluating a scientific hypothesis or deciding how to respond to an environmental threat typically requires several judgments at once. Evidence may be incomplete, uncertainty may differ across components and the final decision may depend on how those components interact. Decomposition can make those hidden structures explicit, allowing a group to identify where it agrees, where it disagrees and which uncertainties matter most.</p>
<p>The result is not merely a matter of dividing labor. In a well-designed collective process, subproblem answers must eventually be integrated. That integration can involve averaging estimates, weighting evidence, comparing competing explanations or using a formal decision rule. The quality of the final outcome therefore depends on both stages: the accuracy of the partial judgments and the method used to recombine them. The study’s emphasis on collective problem decomposition suggests that improving the first stage can substantially strengthen the second.</p>
<p>The work also offers a possible response to a familiar problem in online discussion. Digital platforms can assemble thousands of opinions, but volume alone does not guarantee reliable knowledge. Large groups may amplify misinformation, reward rhetorical confidence or become polarized around simplified narratives. A decomposition-based design could ask users to address clearly defined aspects of an issue, expose the reasoning behind each contribution and organize the results before a final collective judgment is formed. Such systems could be useful in citizen science, policy consultation, forecasting platforms and collaborative research.</p>
<p>For artificial intelligence developers, the implications may be equally significant. Many AI systems already use ensembles, multi-agent reasoning or chains of intermediate tasks to tackle difficult problems. The study’s message is closely aligned with that direction: complex reasoning may become more robust when it is distributed across specialized steps rather than attempted as one undifferentiated act. Human groups and AI agents could potentially work through decomposed tasks separately and then compare or synthesize their outputs, creating hybrid systems designed to reduce shared blind spots.</p>
<p>The researchers’ finding does not mean that every problem should be split into smaller pieces, or that group discussion automatically produces correct answers. Decomposition can fail if the subproblems are defined poorly, if important dependencies are ignored or if the final synthesis gives excessive weight to unreliable components. Its value lies in making collective reasoning more deliberate and transparent. In a world increasingly reliant on decisions made by committees, online communities and human-machine teams, the study suggests that the path to smarter crowds may begin not with more people, but with better questions.</p>
<p><strong>Subject of Research</strong>: Collective problem decomposition and the wisdom of deliberative crowds</p>
<p><strong>Article Title</strong>: Collective problem decomposition improves the wisdom of deliberative crowds</p>
<p><strong>Article References</strong>: Barrera-Lemarchand, F., Lescano-Charreau, V., Ruiz, J. <i>et al.</i> Collective problem decomposition improves the wisdom of deliberative crowds. <i>Nature Communications</i> (2026). <a href="https://doi.org/10.1038/s41467-026-76365-y">https://doi.org/10.1038/s41467-026-76365-y</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: 10.1038/s41467-026-76365-y</p>
<p><strong>Keywords</strong>: collective intelligence, wisdom of crowds, deliberation, problem decomposition, group decision-making, social learning, collective reasoning, forecasting, human collaboration</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">177106</post-id>	</item>
	</channel>
</rss>
