<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>large language models performance &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/large-language-models-performance/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 26 Jul 2026 14:20:14 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>large language models performance &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI language models may surpass collaboration benefits as they scale</title>
		<link>https://scienmag.com/ai-language-models-may-surpass-collaboration-benefits-as-they-scale/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 26 Jul 2026 14:20:14 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI collaboration limitations]]></category>
		<category><![CDATA[AI model diversity vs redundancy]]></category>
		<category><![CDATA[AI model ensemble effects]]></category>
		<category><![CDATA[AI system feedback loops]]></category>
		<category><![CDATA[AI system misalignment and misunderstandings]]></category>
		<category><![CDATA[collaborative AI task outcomes]]></category>
		<category><![CDATA[generative sampling errors]]></category>
		<category><![CDATA[impact of model collaboration on reasoning]]></category>
		<category><![CDATA[large language models performance]]></category>
		<category><![CDATA[LLM cooperation challenges]]></category>
		<category><![CDATA[multi-agent AI systems]]></category>
		<category><![CDATA[reinforcement of errors in AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/ai-language-models-may-surpass-collaboration-benefits-as-they-scale/</guid>

					<description><![CDATA[A fresh study in July 2026 challenges a tempting assumption about today’s “capable” AI systems: that combining tools or models always improves results. In experiments reported by Kim, Gu, Park and colleagues, large language models (LLMs) were evaluated under cooperative settings designed to leverage multiple agents, prompts, or intermediate steps. Instead of consistently outperforming single-model [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>A fresh study in July 2026 challenges a tempting assumption about today’s “capable” AI systems: that combining tools or models always improves results. In experiments reported by Kim, Gu, Park and colleagues, large language models (LLMs) were evaluated under cooperative settings designed to leverage multiple agents, prompts, or intermediate steps. Instead of consistently outperforming single-model approaches, the collaborative strategies sometimes failed to deliver—and could even reduce overall quality.</p>
<p>The researchers frame the problem as a mismatch between collaboration incentives and actual task structure. LLMs generate text by predicting likely continuations, so “help” from another model can be interpreted as plausible-but-misaligned language rather than corrective information. When agents exchange outputs, the system may amplify superficial patterns, reinforce early errors, or converge on a shared misunderstanding.</p>
<p>Crucially, the paper distinguishes between improvements driven by genuine diversity and degradation caused by redundant signals. If partner models produce largely overlapping reasoning trajectories, the ensemble-like interaction behaves less like an informed committee and more like a feedback loop. Under such conditions, collaboration can increase confidence in incorrect directions—an effect related to compounding errors and confirmation bias within generative sampling.</p>
<p>To probe when collaboration helps versus hurts, the team analyzes performance across tasks requiring reasoning, consistency, and structured problem solving. They find that some cooperative methods do enhance accuracy, but the benefits are not universal. For certain prompts and difficulty levels, the “outgrowth” phenomenon emerges: as models become more capable, they may no longer need external guidance to reach strong answers, yet the added coordination overhead still introduces distortion.</p>
<p>The study also highlights how evaluation metrics can mask failure modes. Even when final answers appear coherent, internal traces may indicate that agents are negotiating language rather than improving the underlying solution. The authors argue that designers should measure not only end outputs but also consistency checks, error propagation, and how reasoning signals change after interaction.</p>
<p>From a technical standpoint, the results have implications for multi-agent prompting, tool-using assistants, and systems that route tasks among specialized models. The findings suggest that cooperative architectures should be adaptive, selecting collaboration only when it is likely to add complementary information. Otherwise, the extra agents act like “noise injectors” that reshape probability distributions without improving decision quality.</p>
<p>As AI capabilities accelerate, this work serves as a reminder that bigger models do not automatically benefit from more coordination. Sometimes, the simplest path—well-calibrated single-model reasoning with careful prompting—can outperform elaborate collaboration schemes.</p>
<p>The research underscores a broader lesson for viral science and AI watchers alike: smarter collaboration is not the same as more collaboration. For developers and researchers, the next step is designing collaboration protocols that explicitly manage diversity, detect misalignment early, and prevent feedback loops from taking over.</p>
<p><strong>Subject of Research</strong>: Capable language models and collaboration effects in multi-agent settings</p>
<p><strong>Article Title</strong>: Capable language models can outgrow the benefits of collaboration.</p>
<p><strong>Article References</strong>: Kim, Y., Gu, K., Park, C. <em>et al.</em> Capable language models can outgrow the benefits of collaboration. <em>Nat Mach Intell</em> <strong>8</strong>, 1157–1172 (2026). <a href="https://doi.org/10.1038/s42256-026-01268-y">https://doi.org/10.1038/s42256-026-01268-y</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
<p><strong>DOI</strong>: 10.1038/s42256-026-01268-y</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">173893</post-id>	</item>
		<item>
		<title>Can AI Grasp Emotions More Deeply Than Humans?</title>
		<link>https://scienmag.com/can-ai-grasp-emotions-more-deeply-than-humans/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Thu, 22 May 2025 16:45:56 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[advancements in artificial intelligence]]></category>
		<category><![CDATA[AI and human emotional behavior]]></category>
		<category><![CDATA[AI applications in human judgment]]></category>
		<category><![CDATA[AI capabilities in emotional scenarios]]></category>
		<category><![CDATA[AI emotional intelligence]]></category>
		<category><![CDATA[emotional intelligence in technology]]></category>
		<category><![CDATA[emotional intelligence tests in AI]]></category>
		<category><![CDATA[groundbreaking AI studies]]></category>
		<category><![CDATA[human versus AI emotional understanding]]></category>
		<category><![CDATA[implications of AI in emotional contexts]]></category>
		<category><![CDATA[large language models performance]]></category>
		<category><![CDATA[University of Geneva AI research]]></category>
		<guid isPermaLink="false">https://scienmag.com/can-ai-grasp-emotions-more-deeply-than-humans/</guid>

					<description><![CDATA[In a groundbreaking study that challenges conventional wisdom about the limits of artificial intelligence (AI), researchers from the University of Geneva (UNIGE) and the University of Bern (UniBE) have demonstrated that large language models (LLMs) exhibit impressive competence not only in understanding but also in generating emotionally intelligent behaviour. Published recently in Communications Psychology, their [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking study that challenges conventional wisdom about the limits of artificial intelligence (AI), researchers from the University of Geneva (UNIGE) and the University of Bern (UniBE) have demonstrated that large language models (LLMs) exhibit impressive competence not only in understanding but also in generating emotionally intelligent behaviour. Published recently in <em>Communications Psychology</em>, their findings reveal that these AI systems, including the widely known ChatGPT, outperform average human scores on emotional intelligence (EI) tests and can create entirely new evaluative scenarios in mere moments—an achievement that could reshape the future of AI applications in fields traditionally dominated by human judgment.</p>
<p>The study centers on large language models, sophisticated AI frameworks designed to process, interpret, and generate human language using vast datasets and complex algorithms. These models, capable of answering intricate questions and navigating nuanced textual tasks, have primarily been seen as tools for informational retrieval, text synthesis, and problem-solving. However, the question posed by the UNIGE and UniBE team was whether these AI systems could extend their capabilities to the realm of emotional intelligence — a deeply human attribute involving the perception, understanding, and management of emotions both in oneself and others.</p>
<p>To tackle this question, the research team employed a set of five widely recognized emotional intelligence assessments commonly used in both psychological research and corporate environments. These tests are structured around emotionally charged scenarios that require decision-making reflective of emotional understanding and regulation. For example, one scenario describes a situation where a character named Michael is confronted with the fact that a colleague has stolen his idea and is being praised for it—a context calling for a sophisticated emotional and social response. The most emotionally intelligent option, verified by prior human consensus, was to “talk to his superior about the situation” rather than resorting to conflict, silence, or retaliatory theft.</p>
<p>The AI models tested included ChatGPT-4, ChatGPT-3.5 (referred to as ChatGPT-o1), Gemini 1.5 Flash, Copilot 365, Claude 3.5 Haiku, and DeepSeek V3—an array representing the cutting edge of generative AI technology. Upon administering the EI tests to these LLMs, the results were striking: the AIs achieved an average score of 82% correct answers, considerably higher than the human participants’ average of 56%. This disparity strongly suggests that large language models not only comprehend emotional contexts but can also emulate appropriate emotional responses with considerable accuracy.</p>
<p>Beyond merely solving tests, the study went further to explore whether these AI systems could create emotionally nuanced assessments themselves. In a subsequent phase, researchers tasked ChatGPT-4 with generating new emotional intelligence test scenarios from scratch. Remarkably, the scenarios produced by ChatGPT-4 matched the clarity, reliability, and realism of the original human-developed tests, which had taken years of iterative refinement to perfect. Over 400 human participants later took these AI-generated tests, affirming their validity and practical utility. Such rapid generation of new, credible emotional intelligence materials by AI underscores the potential for these systems to aid extensively in educational, psychological, and organizational environments.</p>
<p>This innovative experiment sheds light on several important facets of emotional intelligence as operationalized by AI. The ability of LLMs to grasp subtle emotional cues and recommend behaviours aligned with emotional competence suggests that these models possess a form of “emotional reasoning.” This goes beyond simple pattern recognition into the realm of understanding social norms, context, and the consequences of various emotional responses. The findings challenge the long-held notion that emotional intelligence is exclusively a human domain, mediated by empathy and lived experience.</p>
<p>While the implications are profound, the researchers caution against unregulated reliance on AI for emotionally sensitive roles. The study highlights the necessity of expert oversight and contextual awareness when deploying AI tools in coaching, education, or conflict resolution settings. AI can enhance human capacities, but nuanced judgment, ethical considerations, and cultural sensitivities still require human intervention to ensure appropriate and ethical use.</p>
<p>Technically, the models’ success can be attributed to the extensive training on diverse textual data, enabling pattern extraction from myriad social and emotional contexts embedded in language. By internalizing these patterns, LLMs form probabilistic representations of emotionally intelligent behaviour, allowing them to generalize effectively to novel scenarios, as demonstrated. The capacity to generate new evaluative instruments quickly stems from their generative design, which facilitates creative text production grounded in coherent emotional logic.</p>
<p>The research carried out by Katja Schlegel of UniBE and Marcello Mortillaro of UNIGE exemplifies interdisciplinary collaboration across psychology, affective science, and AI technology. Their methodology, combining rigorous psychological testing frameworks with cutting-edge AI benchmarking, provides a blueprint for future studies on the integration of emotional intelligence and artificial intelligence. This approach could accelerate AI development tailored not just to linguistic proficiency but also socio-emotional competence.</p>
<p>In conclusion, this study marks a pivotal moment in AI research, illustrating that large language models can be both proficient solvers and creators in the emotionally charged dimensions of human behaviour. As AI becomes increasingly embedded in personal and professional spheres, this expanded emotional toolkit within LLMs offers promising avenues for enhanced communication, empathy-driven interactions, and conflict management mediated by technology. The scientific community and industry alike are now tasked with responsibly harnessing these capabilities to amplify human potential without compromising ethical standards.</p>
<hr />
<p><strong>Subject of Research</strong>: Not applicable</p>
<p><strong>Article Title</strong>: Large language models are proficient in solving and creating emotional intelligence tests</p>
<p><strong>News Publication Date</strong>: 22-May-2025</p>
<p><strong>Web References</strong>:<br />
<a href="http://dx.doi.org/10.1038/s44271-025-00258-x"><a href="https://doi.org/10.1038/s44271-025-00258-x">https://doi.org/10.1038/s44271-025-00258-x</a></a></p>
<p><strong>Keywords</strong>: Artificial intelligence, Emotional intelligence, Large language models, ChatGPT, Emotional reasoning, Emotional intelligence tests, Generative AI, Human-AI collaboration, Affective computing, AI in education, Conflict management, Psychological assessment</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">47401</post-id>	</item>
	</channel>
</rss>
