<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>multi-agent AI systems &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/multi-agent-ai-systems/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 26 Jul 2026 14:20:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>multi-agent AI systems &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI language models may surpass collaboration benefits as they scale</title>
		<link>https://scienmag.com/ai-language-models-may-surpass-collaboration-benefits-as-they-scale/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 26 Jul 2026 14:20:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI collaboration limitations]]></category>
		<category><![CDATA[AI model diversity vs redundancy]]></category>
		<category><![CDATA[AI model ensemble effects]]></category>
		<category><![CDATA[AI system feedback loops]]></category>
		<category><![CDATA[AI system misalignment and misunderstandings]]></category>
		<category><![CDATA[collaborative AI task outcomes]]></category>
		<category><![CDATA[generative sampling errors]]></category>
		<category><![CDATA[impact of model collaboration on reasoning]]></category>
		<category><![CDATA[large language models performance]]></category>
		<category><![CDATA[LLM cooperation challenges]]></category>
		<category><![CDATA[multi-agent AI systems]]></category>
		<category><![CDATA[reinforcement of errors in AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/ai-language-models-may-surpass-collaboration-benefits-as-they-scale/</guid>

					<description><![CDATA[A fresh study in July 2026 challenges a tempting assumption about today’s “capable” AI systems: that combining tools or models always improves results. In experiments reported by Kim, Gu, Park and colleagues, large language models (LLMs) were evaluated under cooperative settings designed to leverage multiple agents, prompts, or intermediate steps. Instead of consistently outperforming single-model [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>A fresh study in July 2026 challenges a tempting assumption about today’s “capable” AI systems: that combining tools or models always improves results. In experiments reported by Kim, Gu, Park and colleagues, large language models (LLMs) were evaluated under cooperative settings designed to leverage multiple agents, prompts, or intermediate steps. Instead of consistently outperforming single-model approaches, the collaborative strategies sometimes failed to deliver—and could even reduce overall quality.</p>
<p>The researchers frame the problem as a mismatch between collaboration incentives and actual task structure. LLMs generate text by predicting likely continuations, so “help” from another model can be interpreted as plausible-but-misaligned language rather than corrective information. When agents exchange outputs, the system may amplify superficial patterns, reinforce early errors, or converge on a shared misunderstanding.</p>
<p>Crucially, the paper distinguishes between improvements driven by genuine diversity and degradation caused by redundant signals. If partner models produce largely overlapping reasoning trajectories, the ensemble-like interaction behaves less like an informed committee and more like a feedback loop. Under such conditions, collaboration can increase confidence in incorrect directions—an effect related to compounding errors and confirmation bias within generative sampling.</p>
<p>To probe when collaboration helps versus hurts, the team analyzes performance across tasks requiring reasoning, consistency, and structured problem solving. They find that some cooperative methods do enhance accuracy, but the benefits are not universal. For certain prompts and difficulty levels, the “outgrowth” phenomenon emerges: as models become more capable, they may no longer need external guidance to reach strong answers, yet the added coordination overhead still introduces distortion.</p>
<p>The study also highlights how evaluation metrics can mask failure modes. Even when final answers appear coherent, internal traces may indicate that agents are negotiating language rather than improving the underlying solution. The authors argue that designers should measure not only end outputs but also consistency checks, error propagation, and how reasoning signals change after interaction.</p>
<p>From a technical standpoint, the results have implications for multi-agent prompting, tool-using assistants, and systems that route tasks among specialized models. The findings suggest that cooperative architectures should be adaptive, selecting collaboration only when it is likely to add complementary information. Otherwise, the extra agents act like “noise injectors” that reshape probability distributions without improving decision quality.</p>
<p>As AI capabilities accelerate, this work serves as a reminder that bigger models do not automatically benefit from more coordination. Sometimes, the simplest path—well-calibrated single-model reasoning with careful prompting—can outperform elaborate collaboration schemes.</p>
<p>The research underscores a broader lesson for viral science and AI watchers alike: smarter collaboration is not the same as more collaboration. For developers and researchers, the next step is designing collaboration protocols that explicitly manage diversity, detect misalignment early, and prevent feedback loops from taking over.</p>
<p><strong>Subject of Research</strong>: Capable language models and collaboration effects in multi-agent settings</p>
<p><strong>Article Title</strong>: Capable language models can outgrow the benefits of collaboration.</p>
<p><strong>Article References</strong>: Kim, Y., Gu, K., Park, C. <em>et al.</em> Capable language models can outgrow the benefits of collaboration. <em>Nat Mach Intell</em> <strong>8</strong>, 1157–1172 (2026). <a href="https://doi.org/10.1038/s42256-026-01268-y">https://doi.org/10.1038/s42256-026-01268-y</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: 10.1038/s42256-026-01268-y</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">173893</post-id>	</item>
		<item>
		<title>Collaborative AI Successfully Clears U.S. Medical Licensing Exams</title>
		<link>https://scienmag.com/collaborative-ai-successfully-clears-u-s-medical-licensing-exams/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Thu, 09 Oct 2025 18:19:07 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[addressing hallucinations in AI responses]]></category>
		<category><![CDATA[advancements in medical licensing tests]]></category>
		<category><![CDATA[AI models accuracy improvement]]></category>
		<category><![CDATA[Collaborative AI in healthcare]]></category>
		<category><![CDATA[collective intelligence in medical assessments]]></category>
		<category><![CDATA[enhancing trustworthiness in clinical AI]]></category>
		<category><![CDATA[GPT-4 applications in medicine]]></category>
		<category><![CDATA[multi-agent AI systems]]></category>
		<category><![CDATA[open-access research in digital health]]></category>
		<category><![CDATA[overcoming AI inconsistencies in exams]]></category>
		<category><![CDATA[structured deliberation among AI agents]]></category>
		<category><![CDATA[US Medical Licensing Examination success]]></category>
		<guid isPermaLink="false">https://scienmag.com/collaborative-ai-successfully-clears-u-s-medical-licensing-exams/</guid>

					<description><![CDATA[In a groundbreaking advancement for artificial intelligence in healthcare, a team of researchers has demonstrated that a collaborative approach involving a council of AI models can dramatically enhance the accuracy of medical knowledge assessments. This innovative study reveals that a group of AI agents, based on OpenAI’s GPT-4, working together through structured deliberation, outperforms individual [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking advancement for artificial intelligence in healthcare, a team of researchers has demonstrated that a collaborative approach involving a council of AI models can dramatically enhance the accuracy of medical knowledge assessments. This innovative study reveals that a group of AI agents, based on OpenAI’s GPT-4, working together through structured deliberation, outperforms individual AI models on the notoriously challenging United States Medical Licensing Examination (USMLE). The research, published in the open-access journal PLOS Digital Health, marks a significant step forward in harnessing collective intelligence to address the complexities and variable responses typically encountered in AI-generated medical decisions.</p>
<p>The USMLE is a rigorous, three-stage examination that assesses a physician&#8217;s ability to apply knowledge, concepts, and principles fundamental to the practice of medicine. Traditionally, the exam poses a substantial challenge to AI, partly due to the inconsistency of answers when a single large language model (LLM) is queried multiple times on the same material. These responses often vary in quality, sometimes containing inaccuracies or hallucinations—fabricated information presented as fact—which compromise trustworthiness in clinical settings.</p>
<p>To overcome these challenges, researchers designed a council of AI agents. This ensemble consists of multiple instances of GPT-4, each providing initial answers independently. What sets this approach apart is a novel facilitator algorithm tasked with mediating disagreements among the AI instances. When responses diverge, the facilitator orchestrates a deliberative dialogue that encourages the council to synthesize the varying answers. This iterative exchange results in the generation of a consensus response, which, as the study shows, tends to be notably more accurate than any single model’s reply.</p>
<p>Testing this collaborative AI framework on a set of 325 publicly available USMLE questions spanning Step 1 (foundational biomedical sciences), Step 2 Clinical Knowledge (clinical diagnosis and management), and Step 3 (advanced clinical scenarios), the council achieved remarkable accuracy rates of 97%, 93%, and 94%, respectively. These figures represent a significant improvement over solitary GPT-4 models, demonstrating the potential of collective AI reasoning in complex domains. Even more strikingly, when the council initially failed to reach unanimous agreement, subsequent deliberations resulted in correct consensus answers 83% of the time, showcasing the power of structured dialogue for self-correction.</p>
<p>The implications of these findings extend far beyond standardized exams. In healthcare, where nuanced understanding and precision are paramount, AI tools must be not only powerful but reliable and trustworthy. By enabling multiple AI agents to &#8220;converse&#8221; and refine their outputs, this collaborative approach introduces a quality control mechanism inherently absent in single-model systems. The authors of the study argue that such collective intelligence may redefine how we evaluate AI’s effectiveness and reliability, especially in high-stakes environments like medical diagnosis and treatment planning.</p>
<p>Yahya Shaikh, the lead author from Baltimore, emphasizes the transformative potential of this methodology. “Our research establishes that when AI systems engage in a structured dialogue, they can surpass the accuracy of any one system, achieving unprecedented performance on complex medical licensing exams without specialized training on medical data,” Shaikh says. This insight underscores the viability of leveraging dialogue and diversity among AI models as means to minimize errors and harness the strengths of their varied reasoning pathways.</p>
<p>Importantly, the study also addresses fundamental misconceptions about AI reliability. Conventional wisdom favors consistency in AI outputs as a hallmark of quality, yet the work by Shaikh and colleagues reveals that variability among responses, when channeled appropriately, becomes an asset rather than a liability. Instead of expecting uniformity, the council model embraces diverse perspectives, allowing AI agents to weigh different interpretations and evidence before converging on a final answer. This dynamic mirrors human expert panel discussions and decision-making processes, lending credibility to the concept of AI teamwork.</p>
<p>Another intriguing aspect of the research involves the concept of semantic entropy, a measure used to quantify the diversity and uncertainty in AI-generated answers. Zainab Asiyah, a co-author of the study, notes that semantic entropy provides a narrative of the AI’s internal struggle and eventual resolution, akin to a human cognitive journey. This metric reveals how individual AI agents can influence one another’s viewpoints through conversation, sometimes even challenging incorrect answers and steering the group toward consensus.</p>
<p>While the results are promising, the authors caution that the council approach has yet to be validated in real-world clinical practice. Practical applications will require addressing challenges such as integrating AI councils into existing healthcare workflows, ensuring transparency of deliberation processes, and managing the ethical implications of AI-influenced decisions. Nonetheless, the demonstrated ability of AI systems to self-correct and improve through interaction opens a new horizon for AI deployment in medicine, education, and possibly other knowledge-intensive fields.</p>
<p>Zishan Siddiqui, another co-author, highlights the pragmatic nature of this work by emphasizing that the focus is not merely on showcasing AI’s test-taking abilities but on developing a method that leverages AI’s natural variation to enhance accuracy. “This system’s capability to ‘take a few tries, compare notes, and self-correct’ should be incorporated into future tools designed for education and healthcare, where correctness is non-negotiable,” Siddiqui notes. By fostering redundancy and collaborative reasoning, the council model could serve as a foundation for safer, more effective AI-based decision support systems.</p>
<p>This study stands as a compelling example of the evolving landscape of artificial intelligence research, where the focus shifts from individual model performance to synergistic interactions among multiple agents. Collaborative intelligence presents a promising avenue to overcome the current limitations of LLMs, offering robustness against errors and amplifying strengths through collective insight. As AI continues to integrate into critical domains, approaches that emulate human-like collaboration among AI systems are likely to shape the future trajectory of machine learning applications.</p>
<p>In summary, the research conducted by Yahya Shaikh and colleagues presents a transformative approach by demonstrating that a council of GPT-4-based AI models, working through iterative discussion and consensus-building, can significantly enhance accuracy on medical licensing exams. This work challenges preconceived notions of AI consistency and reliability, introducing a paradigm where variability and constructive debate are harnessed to achieve superior outcomes. As AI tools advance, embracing collaborative intelligence may hold the key to unlocking new levels of accuracy, trust, and utility in both medical and broader scientific contexts.</p>
<hr />
<p><strong>Subject of Research</strong>: Not applicable</p>
<p><strong>Article Title</strong>: Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE</p>
<p><strong>News Publication Date</strong>: October 9, 2024</p>
<p><strong>Web References</strong>: <a href="http://dx.doi.org/10.1371/journal.pdig.0000787">http://dx.doi.org/10.1371/journal.pdig.0000787</a></p>
<p><strong>References</strong>: Shaikh Y, Jeelani-Shaikh ZA, Jeelani MM, Javaid A, Mahmud T, Gaglani S, et al. (2025) Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE. PLOS Digit Health 4(10): e0000787.</p>
<p><strong>Image Credits</strong>: Nguyen Dang Hoang Nhu, Unsplash (CC0)</p>
<p><strong>Keywords</strong>: Artificial Intelligence, Medical Licensing Exam, USMLE, GPT-4, Collaborative Intelligence, Large Language Models, AI Deliberation, Medical AI Accuracy, Structured Dialogue, AI Self-Correction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">88370</post-id>	</item>
	</channel>
</rss>
