<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI agents &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-agents/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 22:33:51 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI agents &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Agents Are Being Graded Wrong: Landmark Audit Finds No Benchmark Controls All Key Threats</title>
		<link>https://scienmag.com/ai-agents-are-being-graded-wrong-landmark-audit-finds-no-benchmark-controls-all-key-threats/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 22:33:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agent evaluation]]></category>
		<category><![CDATA[AI agent evaluation]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[Artificial Intelligence Review]]></category>
		<category><![CDATA[assessment of AI threat detection and control]]></category>
		<category><![CDATA[autonomous AI decision-making]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[benchmarking limitations in artificial intelligence]]></category>
		<category><![CDATA[challenges in interpreting AI benchmark scores]]></category>
		<category><![CDATA[critique of current AI performance metrics]]></category>
		<category><![CDATA[data contamination]]></category>
		<category><![CDATA[evaluation infrastructure]]></category>
		<category><![CDATA[execution cost]]></category>
		<category><![CDATA[external tool integration in AI systems]]></category>
		<category><![CDATA[impact of benchmark design on AI scoring]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[long-chain reasoning in AI agents]]></category>
		<category><![CDATA[meta-taxonomy]]></category>
		<category><![CDATA[multi-step task planning in language models]]></category>
		<category><![CDATA[non-determinism]]></category>
		<category><![CDATA[PRISMA-ScR]]></category>
		<category><![CDATA[PRISMA-ScR systematic review methodology]]></category>
		<category><![CDATA[systematic review]]></category>
		<category><![CDATA[systematic review of AI benchmarks]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212879</guid>

					<description><![CDATA[A systematic survey of 259 studies finds that none of seventeen prominent AI agent benchmarks jointly controls data contamination, non-determinism, and execution cost, prompting a roadmap for reliability-first evaluation.]]></description>
										<content:encoded><![CDATA[<p>Large language models no longer simply answer questions. They plan multi-step tasks, call external tools, browse environments, and act autonomously across long chains of decisions. Yet according to a sweeping new systematic survey published in Artificial Intelligence Review, the benchmarks used to grade these AI agents produce scores that are far less meaningful than the field assumes. The study, led by Vinoth Nageshwaran of the University of the Cumberlands with colleagues at Indiana University of Pennsylvania, Ton Duc Thang University, and Elizabeth City State University, argues that a benchmark number is not self-interpreting: what it means depends entirely on which capability it measures, how that capability is scored, and where the agent is tested when it is measured.</p>
<p>The research team built their analysis on a reproducible systematic mapping review conducted under the PRISMA-ScR standard, the accepted protocol for systematic reviews in health and computer science research. After a rigorous screening pipeline, 259 primary studies formed the core analytical corpus, with 294 studies in the released living-review corpus. The authors are unusually candid about the limits of their own method: primary-arm records were screened by a single audited automated pass, while human double-screening with an inter-rater agreement of kappa equal to 0.65 was applied only to a supplementary arm. They explicitly describe broad-map attribute shares as provisional heuristic estimates and stress that the corpus is a carefully constructed mapping sample, not a census of the entire field. That kind of methodological honesty is rare in a literature often criticized for overclaiming.</p>
<p>The survey&#8217;s first major contribution is a crisp conceptual boundary called the dependent-step test. The test demarcates genuine autonomous agentic evaluation from ordinary static natural-language-processing evaluation and from prompt engineering. In essence, an evaluation qualifies as agentic only when later steps in a task depend on the outcomes of earlier steps, so that errors compound and the agent must recover, replan, or abandon a strategy mid-trajectory. A model that answers a thousand independent trivia questions is being tested very differently from one that must book a flight, notice the payment failed, diagnose why, and retry with a corrected form. The dependent-step test gives researchers a principled way to decide whether a benchmark is actually measuring agency or merely repackaging static question answering.</p>
<p>The second contribution is a meta-taxonomy that places every agent benchmark in a three-pillar coordinate system: capability, meaning what the benchmark measures; scoring paradigm, meaning how performance is judged; and environment topology, meaning where the agent operates. This framework matters because two benchmarks with identical headline scores may probe entirely different faculties. One may reward factual recall in a sandboxed text environment, while another measures tool orchestration in a live operating system where a single mistyped command can cascade into failure. By locating each benchmark along these three axes, the taxonomy lets researchers compare evaluations like coordinates on a map rather than as isolated leaderboard entries, exposing blind spots where entire capability regions remain untested.</p>
<p>Third, the authors paired an evidence map over the full corpus with a purposive, hand-verified, venue-verified landscape matrix of seventeen prominent benchmarks, achieving an inter-rater reliability of kappa equal to 0.77, a level conventionally regarded as substantial agreement. This matrix functions as a deep-dive companion to the broad map: where the corpus-wide analysis offers breadth with provisional estimates, the seventeen-benchmark matrix offers verified depth on the evaluations that most actively shape the field&#8217;s self-image. The contrast between the two instruments is itself instructive, showing how much confidence a reader should place in each kind of claim.</p>
<p>The survey&#8217;s most striking finding emerges from its critical analysis of that verified matrix. The authors identify a structural trilemma facing agent evaluation: data contamination, non-determinism, and execution cost. Data contamination occurs when benchmark tasks leak into training data, inflating scores without genuine capability. Non-determinism arises because agentic trajectories involve stochastic model behavior and environment interactions, so the same agent can score differently across runs. Execution cost reflects the computational expense of running agents through long, tool-using trajectories, which discourages the repeated trials needed to tame non-determinism. Among the seventeen prominent, verified benchmarks, the study found that none reports evidence that all three threats are jointly controlled, and none reports a complete standardized run-cost record, a 0 out of 17 result that should unsettle anyone quoting leaderboard numbers.</p>
<p>The implications of that 0 out of 17 finding ripple outward. When contamination is uncontrolled, a high score may reflect memorization rather than reasoning. When non-determinism is unquantified, a single reported run may be a lucky draw rather than a stable estimate of ability. When cost is unrecorded, results cannot be reproduced at reasonable expense, and the community cannot even assess whether an evaluation is practical to repeat. Together these gaps mean that many celebrated comparisons between competing agents may be statistically fragile, and that the field&#8217;s rapid narrative of progress rests on measurements whose error bars are largely unknown. The trilemma is structural because addressing any one threat tends to worsen another: more repetitions to handle non-determinism raise cost, and larger task sets that resist contamination raise it further.</p>
<p>The transparency of the review process itself deserves attention as a model for the field. The authors released a companion repository containing their pipeline, including a build script that reproduces every reported count against a committed HTTP response cache, yielding byte-identical, MD5-verified outputs with zero live API calls. They document per-source query status, flagging databases where only counts could be verified and one major index that remained unresolved for lack of credentials. They even ran a capture-recapture diagnostic between OpenAlex and Crossref, then declined to treat its output as a recall estimate because the two indices are heterogeneous targeted searches rather than independent random samples, a violation of the estimator&#8217;s core assumptions. A probe-set validation instead confirmed that canonical benchmarks were captured completely, with Scopus corroborating the result. This is bibliometrics done with unusual care.</p>
<p>Looking forward, the survey proposes a five-direction roadmap that reframes progress not as the invention of yet more individual benchmarks but as a shift toward standardized, reliability-first, cost-aware evaluation infrastructure. In practice that means community-agreed protocols for reporting contamination checks, variance across repeated runs, and the compute cost of every evaluation, so that a benchmark score arrives with the same kind of measurement metadata that a physics experiment or clinical trial would demand. The roadmap effectively asks the field to grow up: to treat agent evaluation as a measurement science with known instruments, calibrated error, and reproducible procedures, rather than as a proliferation of leaderboards each with its own unstated assumptions.</p>
<p>For an industry pouring billions into agentic AI products, the stakes could hardly be higher. Autonomous agents are beginning to write code, manage workflows, and interact with real systems on behalf of users, and deployment decisions increasingly cite benchmark performance as evidence of safety and competence. If, as this survey shows, no prominent benchmark currently demonstrates joint control over contamination, non-determinism, and cost, then the numbers guiding those decisions are weaker than they appear. The study does not claim agents are failing; it claims we cannot yet reliably tell. Turning evaluation from a competitive spectacle into trustworthy infrastructure, the authors argue, is the prerequisite for knowing what today&#8217;s agents can actually do, and the open repository accompanying the paper offers the community a concrete place to start.</p>
<p><strong>Subject of Research:</strong> Evaluation and benchmarking of large language model agents</p>
<p><strong>Article Title:</strong> Large language model agent evaluation and benchmarking: a systematic survey, meta-taxonomy, and critical research roadmap</p>
<p><strong>Article References:</strong> Nageshwaran, V., Ezekiel, S., Tran, T. T., &amp; Lakshmi Narasimhan, V. (2026). Large language model agent evaluation and benchmarking: a systematic survey, meta-taxonomy, and critical research roadmap. <em>Artificial Intelligence Review</em>. <a href="https://doi.org/10.1007/s10462-026-11678-4" rel="noopener noreferrer">https://doi.org/10.1007/s10462-026-11678-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10462-026-11678-4" rel="noopener noreferrer">10.1007/s10462-026-11678-4</a></p>
<p><strong>Keywords:</strong> large language models, AI agents, benchmarking, agent evaluation, meta-taxonomy, systematic review, data contamination, non-determinism, execution cost, PRISMA-ScR, evaluation infrastructure, Artificial Intelligence Review</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212879</post-id>	</item>
		<item>
		<title>How Language Models Learn to Use Tools Like Humans Do</title>
		<link>https://scienmag.com/how-language-models-learn-to-use-tools-like-humans-do/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 22:56:28 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[advancements in artificial intelligence research]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI and human tool use comparison]]></category>
		<category><![CDATA[AI integration with search engines and APIs]]></category>
		<category><![CDATA[API integration]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[Benchmarks]]></category>
		<category><![CDATA[cognitive parallels between humans and AI]]></category>
		<category><![CDATA[development of AI tool learning]]></category>
		<category><![CDATA[evolution from passive text generators to active tool orchestrators]]></category>
		<category><![CDATA[foundation models]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[impact of AI on problem-solving capabilities]]></category>
		<category><![CDATA[Language model tool use]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models as active tool users]]></category>
		<category><![CDATA[natural language processing]]></category>
		<category><![CDATA[neuroscience of tool use in humans]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[role of specialized brain regions in tool use]]></category>
		<category><![CDATA[supervised fine-tuning]]></category>
		<category><![CDATA[systematic survey on AI tool learning]]></category>
		<category><![CDATA[task planning]]></category>
		<category><![CDATA[tool learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=208559</guid>

					<description><![CDATA[A comprehensive new survey maps how large language models learn to plan, select, execute, and integrate external tools, charting the methods and benchmarks driving tool-augmented AI.]]></description>
										<content:encoded><![CDATA[<p>Tool use has long been considered a defining hallmark of human intelligence. From the earliest stone implements to modern software, our species has extended its physical and cognitive reach by designing and manipulating instruments that solve problems beyond our native abilities. Neuroscientists have even identified specialized regions of the human brain, such as the anterior supramarginal gyrus, that activate uniquely during tool use, a capacity absent in our closest primate relatives. Now, a comprehensive new survey published in the open-access journal Vicinagearth argues that artificial intelligence is crossing a strikingly similar threshold. A team of researchers from Fudan University, East China Normal University, the University of Science and Technology of China, and the Institute of Artificial Intelligence at China Telecom has mapped out how large language models are learning to become genuine tool users, transforming them from passive text generators into active orchestrators of calculators, search engines, application programming interfaces, and far more.</p>
<p>The survey, led by Jinyang Chen, Haolun Wu, Jianhong Pang, and Yihua Wang with senior supervision from Dell Zhang and Changzhi Sun, provides the most systematic account to date of what the field calls tool learning. The core idea is deceptively simple: instead of expecting a language model to answer every question from its internal parameters, the model should learn to delegate parts of a problem to external tools that are faster, more accurate, or more current. When a user asks about tomorrow&#8217;s weather in Tokyo, the model should not hallucinate a forecast but call a live weather API. When a calculation involves more than trivial arithmetic, the model should hand the numbers to a code interpreter or symbolic solver. This reframing, the authors argue, turns tool use from physical manipulation into symbolic orchestration, where the language model acts as an intelligent controller that understands both what the user wants and what each tool can do.</p>
<p>At the heart of the survey lies a unified four-stage framework that the researchers use to organize the entire landscape: task planning, tool selection, task execution, and response generation. In the planning stage, the model decomposes a complex, high-level instruction into a series of smaller, solvable subtasks, working out the dependencies and the order in which they must run. Systems like ART build libraries of example tasks that guide this decomposition through few-shot prompts, while HuggingGPT combines specification-based instructions with demonstration parsing to schedule subtasks, and RestGPT refines its plans iteratively in a coarse-to-fine scheme as execution proceeds. Planning, the authors stress, is the cognitive foundation on which everything else depends, because a model that cannot break a problem down correctly will select the wrong tools or call them in the wrong order.</p>
<p>Tool selection is the second stage, and it presents a genuinely difficult engineering problem when the number of candidate tools reaches into the thousands. The survey distinguishes two dominant strategies. Retriever-based approaches first filter the tool library using semantic relevance, drawing on classical term-matching techniques such as TF-IDF and BM25 or on neural models like Sentence-BERT trained specifically for tool retrieval. Newer systems push this further: Craft asks the language model to generate hypothetical tool descriptions from a query and retrieve against them, while COLT applies graph neural networks and explicitly targets completeness of retrieval, ensuring no relevant tool is missed. When the candidate pool is small, LLM-based selection takes over, letting the model reason directly over tool descriptions. Methods like ToolBench combine fine-tuning with example retrieval and system prompts, while ToolVerifier generates comparative questions that force the model to distinguish between superficially similar tools, reducing misselection in ambiguous contexts.</p>
<p>Task execution, the third stage, is where natural language must be converted into structured, machine-readable calls. The survey illustrates this with a simple but revealing example: to convert 100 dollars into euros, the model must recognize that a currency API is required, extract the amount, source currency, and target currency, and emit a well-formed call such as convert with the appropriate named parameters. Getting this right demands schema-guided prompting, as in RestGPT and EasyTool, or more elaborate reasoning strategies like ReverseChain, which works backward from the desired final tool to infer the intermediate steps and parameters. Multi-agent architectures such as ToolNet and ConAgents divide the labor among specialized sub-agents that parse, plan, and execute, negotiating with each other over a shared backbone model. Meanwhile, instruction-tuned systems like Gorilla and ToolkenGPT, trained on massive tool-use corpora, generalize to unseen APIs with remarkable fluency, and verification modules such as ToolVerifier and Themis check proposed calls for validity before anything is actually executed, catching unsafe or ill-formed invocations early.</p>
<p>The final stage, response generation, determines how tool outputs are woven back into the model&#8217;s answer. The survey identifies two broad paradigms. Direct insertion methods, exemplified by TALM and Toolformer, embed the raw tool output at placeholder positions in the prompt and let the language model continue from there; Toolformer went further by fine-tuning on synthetic data in which API calls are interleaved with text, teaching the model when and where to invoke tools during generation. Information integration methods go deeper: ToolLLM routes tool outputs through a dedicated integration module before the main model composes its answer, ReCOMP compresses retrieved and computed information into latent representations for tighter factual grounding, and ConAgents lets multiple agents reinterpret tool results collaboratively before drafting the final response. The trade-off is clear: direct insertion is fast and simple, while deeper integration produces more coherent, context-sensitive answers in multi-tool scenarios.</p>
<p>Perhaps the most consequential part of the survey is its method-centric synthesis of how these capabilities are actually learned. Tuning-free approaches rely purely on prompt engineering and in-context demonstration, making them ideal for proprietary models that cannot be retrained and for rapid prototyping against unseen APIs; ReAct, which interleaves chain-of-thought reasoning with tool calls and incorporates environmental feedback, remains the canonical example. Supervised fine-tuning trains models on curated traces of tool interaction, from Toolformer&#8217;s self-generated annotations to ToolBench&#8217;s multi-stage datasets, with newer methods such as SFT-GO optimizing semantically important tokens separately and rehearsal-based strategies preventing catastrophic forgetting of general abilities. Reinforcement learning, however, is emerging as the frontier. ToolRL showed that fine-grained, step-level rewards outperform coarse final-answer signals and achieved up to a 17 percent improvement over supervised fine-tuning using Group Relative Policy Optimization. OTC penalizes unnecessary tool calls and cut tool usage by up to 73 percent while boosting productivity by 229 percent, and ReTool, which combines code execution with textual reasoning, reached 72.5 percent accuracy on the AIME mathematics competition, surpassing even OpenAI&#8217;s o1-preview baseline while exhibiting emergent self-correction.</p>
<p>Evaluation has matured alongside the methods. The survey reviews a rich ecosystem of benchmarks, from broad, coverage-oriented suites like ToolBench, API-Bank, and APIBench that test generalization across hundreds or thousands of APIs, to diagnostic instruments such as T-Eval, which scores six distinct dimensions of tool competence, and the Berkeley Function Calling Leaderboard, which measures structured invocation and agentic orchestration. Scenario-specific benchmarks probe the edges: ToolQA and ToolTalk test tool-grounded question answering in dialogue, ToolEmu simulates tools safely when real execution is too costly or risky, InjecAgent probes robustness against adversarial injection attacks, and SCITOOLBENCH and RoTBench examine scientific and symbolic tool chaining. The reported results paint a consistent picture. GPT-4-Turbo leads T-Eval with an overall score of 86.4, and the GPT-4 family dominates the GTA benchmark, yet the open-source xLAM series tops single-turn function calling on BFCL with accuracy up to 89.27, and the fine-tuned Lynx-7B model approaches GPT-3.5 performance on API-Bank, demonstrating that high-quality, ability-diverse training data can narrow the gap considerably.</p>
<p>The stakes extend well beyond leaderboards. The authors argue that tool learning makes language models more trustworthy and interpretable in concrete ways: intermediate tool invocations expose the reasoning path, standardized tool interfaces reduce sensitivity to prompt phrasing, and deterministic, externally verifiable tool outputs help suppress the hallucinations that have plagued large language models since their inception. In high-stakes domains such as finance, law, and healthcare, this traceability is not a luxury but a requirement. Tool integration also lets models engage with databases, scientific solvers, and medical systems, producing results that are domain-specific and checkable rather than merely plausible. This transparency and grounding, the survey contends, is what elevates tool learning from a convenient trick to a foundational capability that redefines what machine intelligence can credibly deliver.</p>
<p>Significant open challenges remain, and the survey is candid about them. Safety is paramount: hallucinated API calls or erroneous parameter choices in open-ended environments can produce dangerous or misleading outcomes, demanding runtime checks, input sanitization, and robust fallback strategies. Latency bottlenecks from multi-step, multi-tool pipelines strain user experience and scalability. Seamless multimodal integration across vision, speech, and structured data requires better interface design, and personalization, incorporating user preferences and history into tool selection, is still largely unsolved. The authors point toward multi-agent collaboration, LLM-driven tool creation in which models synthesize their own new functions, and unified abstraction frameworks that standardize model-tool interaction while enhancing generalization and safety. Richer benchmarks that simulate authentic, multi-stage, multimodal task scenarios are needed to capture the complexities of real-world use. If those challenges are met, the researchers conclude, tool learning will not merely augment language models but will stand as a foundational component of autonomous, transparent, and reliable artificial intelligence, marking the moment machines truly began to extend their own capabilities the way humans have extended theirs for millions of years.</p>
<p><strong>Subject of Research:</strong> Tool learning with large language models, covering methods, pipelines, tuning strategies, and benchmarks for tool-augmented AI systems.</p>
<p><strong>Article Title:</strong> Tool learning with language models: a comprehensive survey of methods, pipelines, and benchmarks</p>
<p><strong>Article References:</strong> Tool learning with language models: a comprehensive survey of methods, pipelines, and benchmarks. (n.d.). <a href="https://doi.org/10.1007/s44336-025-00024-x" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00024-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00024-x" rel="noopener noreferrer">10.1007/s44336-025-00024-x</a></p>
<p><strong>Keywords:</strong> tool learning, large language models, artificial intelligence, reinforcement learning, supervised fine-tuning, API integration, benchmarks, task planning, hallucination, AI agents, natural language processing, foundation models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">208559</post-id>	</item>
		<item>
		<title>AI Voices in the Crowd: Separating Social Influence from Built-In Bias in Language Model Opinion Dynamics</title>
		<link>https://scienmag.com/ai-voices-in-the-crowd-separating-social-influence-from-built-in-bias-in-language-model-opinion-dynamics/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 15:39:41 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agent-based simulation]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI conversational agents in online discussions]]></category>
		<category><![CDATA[AI governance]]></category>
		<category><![CDATA[AI language models social influence]]></category>
		<category><![CDATA[biases in language model training data]]></category>
		<category><![CDATA[built-in bias in large language models]]></category>
		<category><![CDATA[collective opinion change in AI]]></category>
		<category><![CDATA[computational social science]]></category>
		<category><![CDATA[consensus formation]]></category>
		<category><![CDATA[distinguishing social influence from model bias]]></category>
		<category><![CDATA[impact of training data on AI opinions]]></category>
		<category><![CDATA[implications of bias in AI-driven social simulations]]></category>
		<category><![CDATA[influence of structured exchange vs inherent bias]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Nature Communications.]]></category>
		<category><![CDATA[network topology and opinion convergence]]></category>
		<category><![CDATA[opinion dynamics]]></category>
		<category><![CDATA[opinion dynamics in artificial agents]]></category>
		<category><![CDATA[polarization]]></category>
		<category><![CDATA[simulated focus groups in AI research]]></category>
		<category><![CDATA[social influence]]></category>
		<category><![CDATA[synthetic respondents]]></category>
		<category><![CDATA[training bias]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=206499</guid>

					<description><![CDATA[New research in Nature Communications provides a controlled framework for separating genuine opinion dynamics between AI agents from the shared biases embedded in the models themselves.]]></description>
										<content:encoded><![CDATA[<p>Large language models are no longer passive tools that answer questions in isolation. Increasingly, they are deployed as conversational agents, simulated focus groups, synthetic survey respondents, and automated participants in online discussions, which means their opinions do not simply sit inside them — they circulate. A study published in Nature Communications tackles a deceptively simple question that follows from this shift: when a population of large language models changes its collective opinion over time, how much of that change comes from genuine interaction between the agents, and how much is merely a reflection of the biases already baked into the models themselves?</p>
<p>The distinction matters because the two effects demand completely different responses. If opinions converge because agents influence one another through structured exchange, that is a dynamical phenomenon — something that depends on network topology, repeated contact, and the rules of communication. If opinions converge because every model shares the same training data and the same fine-tuning choices, that is a bias phenomenon, and no amount of network rewiring will fix it. Conflating the two leads researchers to draw false conclusions about social dynamics whenever they use language models to simulate human populations, a practice that has grown rapidly across computational social science.</p>
<p>Classical opinion dynamics models, from the DeGroot framework to bounded-confidence models such as Deffuant and Hegselmann–Krause, have long separated individual predispositions from interpersonal influence. A human agent starts with a prior position and then updates it as a weighted function of the neighbors&#8217; positions. When all agents start from identical priors, any observed change must come from the interaction term. When agents never interact, any observed alignment must come from shared priors. The new study imports this clean separation into the world of language models, where it is far harder to achieve, because an LLM&#8217;s &#8216;prior&#8217; is opaque and entangled with billions of parameters shaped by training corpora, alignment procedures, and decoding settings.</p>
<p>The methodological core of the work lies in designing controlled experimental conditions that isolate the two channels. In one condition, model agents are allowed to exchange opinions across multiple rounds, each agent seeing and responding to the outputs of others, so that interaction effects can accumulate. In a matched condition, agents are queried in complete isolation, with no exposure to each other&#8217;s outputs, so that any consistency in their answers reflects only the models&#8217; intrinsic tendencies. By comparing the trajectories of these two conditions across topics, model families, and network structures, the researchers can estimate how much of the observed opinion dynamics is attributable to social influence and how much to shared bias.</p>
<p>The results carry an important warning for anyone using LLMs as synthetic participants. Because large commercial models are trained on broadly overlapping corpora and aligned toward similar helpfulness and safety objectives, they exhibit substantial baseline agreement: asked independently, they often cluster around the same positions on political, ethical, and policy questions. When such models are then placed in an interaction network, this shared prior acts like a strong external field, pulling the whole population toward the same attractor regardless of the communication structure. Apparent consensus in an LLM society can therefore be an artifact of homogeneity in the underlying models rather than an emergent product of deliberation — the opposite of what genuine social consensus formation looks like in human groups.</p>
<p>The study also quantifies when interaction does matter. Interaction effects become visible when agents begin from heterogeneous positions, when prompts are constructed to suppress the models&#8217; default leanings, or when network structures channel information asymmetrically so that some agents act as hubs and others as peripheral listeners. Under these conditions, the trajectory of collective opinion diverges measurably from the isolated-query baseline, and classical dynamical concepts — anchoring, threshold effects, polarization into clusters — reappear in recognizable form. This suggests that the tools of decades of opinion dynamics research remain applicable to artificial agents, provided the bias floor is first accounted for.</p>
<p>For the broader scientific community, the findings arrive at a moment of intense debate about the validity of LLM-based social simulation. Several recent papers have shown that synthetic samples generated by language models can reproduce survey response patterns with striking fidelity, raising hopes for cheap, scalable, and ethically uncomplicated substitutes for human participants. Other work has cautioned that such fidelity is skin-deep: models reproduce the central tendencies of human populations while flattening minority viewpoints, amplifying majority biases, and failing to capture the contextual sensitivity of real respondents. The new disentangling framework gives this debate a sharper analytical instrument, allowing researchers to state precisely which portion of a simulated social outcome they are willing to trust.</p>
<p>The practical implications extend to platform governance as well. As AI agents increasingly populate recommendation feeds, comment sections, and automated moderation pipelines, the opinions they express are not neutral background noise; they actively shape the informational environment that human users experience. If those agents converge on uniform positions due to shared training biases rather than deliberative processes, they could function as an invisible consensus machine, nudging public discourse in directions no deliberative body ever chose. Understanding the bias-versus-interaction decomposition is therefore not only a matter of methodological hygiene for simulator designers but a governance question for the emerging mixed human–AI public sphere.</p>
<p>The study points toward concrete best practices. Researchers using LLM agents to model opinion dynamics should always run the isolation control: query each agent independently and establish the bias baseline before interpreting any collective behavior. They should diversify model families, prompt formulations, and initial conditions to prevent a single homogeneous prior from dominating the result. And they should report the decomposition explicitly, distinguishing influence-driven convergence from bias-driven agreement, so that downstream readers know whether an emergent consensus in silico says something about social process or merely about the models themselves. As artificial agents become permanent residents of our information ecosystems, tools like this one offer a way to keep the distinction between what agents say to each other and what they were built to say firmly in view.</p>
<p><strong>Subject of Research:</strong> Disentangling interaction effects from training biases in the collective opinion dynamics of large language model agents</p>
<p><strong>Article Title:</strong> Disentangling interaction and bias effects in opinion dynamics of large language models</p>
<p><strong>Article References:</strong> Brockers, V. C., Ehrlich, D. A., &amp; Priesemann, V. (2026). Disentangling interaction and bias effects in opinion dynamics of large language models. <em>Nature Communications, 17</em>(1), Article 10077. <a href="https://doi.org/10.1038/s41467-026-77340-3" rel="noopener noreferrer">https://doi.org/10.1038/s41467-026-77340-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41467-026-77340-3" rel="noopener noreferrer">10.1038/s41467-026-77340-3</a></p>
<p><strong>Keywords:</strong> large language models, opinion dynamics, AI agents, training bias, computational social science, agent-based simulation, social influence, polarization, consensus formation, Nature Communications, synthetic respondents, AI governance</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">206499</post-id>	</item>
		<item>
		<title>AI Agents Can Now Build and Test Building Energy Models From Plain Language</title>
		<link>https://scienmag.com/ai-agents-can-now-build-and-test-building-energy-models-from-plain-language/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:13:59 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agent benchmark]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI agents in construction planning]]></category>
		<category><![CDATA[AI-driven building design optimization]]></category>
		<category><![CDATA[AI-powered building energy modeling]]></category>
		<category><![CDATA[ASHRAE 90.1]]></category>
		<category><![CDATA[automated energy model creation and modification]]></category>
		<category><![CDATA[building energy modeling]]></category>
		<category><![CDATA[building envelope and equipment simulation]]></category>
		<category><![CDATA[Department of Energy building research]]></category>
		<category><![CDATA[EnergyPlus]]></category>
		<category><![CDATA[HVAC synthesis]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[Measure authoring]]></category>
		<category><![CDATA[Model Context Protocol]]></category>
		<category><![CDATA[natural language interface for energy simulation]]></category>
		<category><![CDATA[open-source EnergyPlus engine integration]]></category>
		<category><![CDATA[OpenStudio SDK]]></category>
		<category><![CDATA[OpenStudio-MCP]]></category>
		<category><![CDATA[OpenStudio-MCP software for energy modeling]]></category>
		<category><![CDATA[physics-based building performance diagnostics]]></category>
		<category><![CDATA[plain language building performance analysis]]></category>
		<category><![CDATA[reducing human coding in energy modeling]]></category>
		<category><![CDATA[sandboxing]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202051</guid>

					<description><![CDATA[Researchers have unveiled OpenStudio-MCP, an open-source server that lets AI agents create, simulate, and diagnose building energy models from plain-language requests.]]></description>
										<content:encoded><![CDATA[<p>A building&#8217;s lifetime energy bill is largely decided before anyone pours a foundation. Choices about form, envelope, equipment, controls, and operation lock in performance years in advance, and building energy modeling is the discipline that prices those choices before construction. Now researchers at the National Laboratory of the Rockies, working for the U.S. Department of Energy, have unveiled a system that hands that entire modeling workflow to artificial intelligence agents. The software, called OpenStudio-MCP, is described in the journal SoftwareX and lets large language models create, modify, simulate, and diagnose physics-based building energy models from nothing more than a plain-language request, with no human writing code at any point.</p>
<p>The open-source EnergyPlus engine, developed by the Department of Energy, performs the underlying calculations that predict how a design will perform. Most tools operate on its text input files, but practitioners typically work one layer up, in the OpenStudio software development kit, which provides a typed model structure and a library of reusable transformation scripts known as Measures. The new server deliberately builds on that SDK layer rather than raw input files, because the typed object structure is more composable and less error-prone to manipulate, and it connects AI agents to the broader ecosystems of OpenStudio-standards, ComStock, and the Building Component Library. Earlier protocol-based servers, such as EnergyPlus-MCP, work only at the input-file layer and defer geometry creation and HVAC loop construction to future work; OpenStudio-MCP tackles the full lifecycle from creation through simulation and evaluation.</p>
<p>The technical heart of the system is the Model Context Protocol, introduced by Anthropic in late 2024, which allows a language model to drive external software by calling validated, schema-typed tools instead of emitting free-form code. OpenStudio-MCP exposes 197 such tools, organized into self-contained skill modules that are discovered automatically at startup. They span model creation, geometry and thermal zoning, construction and schedule assignment, HVAC synthesis, simulation control, results extraction, quality assurance, and even large-scale parametric analysis on OpenStudio-server. A single call to add_baseline_system, for example, wires a complete air- and plant-loop topology corresponding to one of the ten baseline system types defined in ASHRAE Standard 90.1 Appendix G, a task that has historically required careful manual scripting.</p>
<p>Because language models can call operations out of order, select unsuitable systems, or simply invent SDK methods that do not exist, the server wraps the SDK at two levels. Low-level tools expose explicit OpenStudio operations as typed calls, while higher-level tools encode common workflows such as whole-building creation, baseline HVAC assignment, and result extraction. A separate knowledge layer serves curated workflow guides covering object dependencies, ASHRAE system-selection rules, and Measure authoring, which agents retrieve on demand. The server also guards against the finite context window of any language model: instead of dumping a 5,000-space model or a full 8,760-hour annual result into the conversation, it returns compact structured summaries, previews models without loading them, and distills entire simulations into a handful of summary numbers while the model files themselves remain server-side.</p>
<p>Perhaps the most striking capability is automated Measure authoring. Measures are programs, and extending analysis beyond the off-the-shelf library has traditionally required engineers to double as software developers, putting custom analysis out of reach for many architects and engineers who best understand the building. With OpenStudio-MCP, an agent scaffolds a new Measure from a plain-language request, writes its logic, runs its tests, and applies it inside a sandboxed environment that runs child processes as unprivileged users under a fail-closed Landlock filesystem policy, a seccomp filter denying outbound network access, and strict resource limits. The authors stress these are implemented controls rather than security guarantees, since the software has not undergone a formal external penetration test, but targeted adversarial probes using canary listeners and decoy secrets exercised cross-tenant file reads, environment-secret capture, filesystem escapes, and forged transfer requests.</p>
<p>To demonstrate the system end to end, the researchers gave an AI agent a natural-language request specifying a medium office building, its location, a baseline HVAC system, a comfort criterion, and a four-pipe active-chilled-beam retrofit, naming no tools. Working from an empty session with Claude Opus 4.8, the agent created and simulated a 27-zone, three-story, 53,600-square-foot office using the ASHRAE 90.1-2019 template with a variable-air-volume reheat system. The initial model narrowly failed the comfort criterion with 301.7 occupied unmet hours. Digging into sizing reports, the agent found that heating capacity was adequate but a 50 percent reheat-mode airflow cap was constraining morning warm-up after nighttime setback, raised the cap on all 27 terminals, and brought the model to 78.0 unmet hours at a site energy use intensity of 42.4 kBtu per square foot, squarely within the middle half of the observed U.S. office stock from the 2018 Commercial Buildings Energy Consumption Survey.</p>
<p>The agent then authored a Measure that replaced each terminal with a four-pipe beam connected to the existing chilled- and hot-water plants, verified SDK class names, passed its own tests, and validated 27 beam terminals with no errors, all without a human writing or reviewing code. The retrofit maintained comfort but increased site energy by 8.7 percent, driven by a 179 percent surge in fan energy, because the constant-volume beams forfeited the variable-volume air handler&#8217;s part-load fan savings and economizer hours. Crucially, the agent reported this adverse result, explained both mechanisms, and proposed remedies including a right-sized dedicated outdoor air system with energy recovery. The session used 98 calls to 36 tools, three annual simulations, and roughly 21 minutes, with no human intervention after the initial request.</p>
<p>Feasibility is not reliability, so the team built a reproducible benchmark of 16 graded tasks across six families, tested with Claude Opus 4.8, Opus 4.6, Sonnet 4.6, and Haiku 4.5 alongside GPT-5.4 and GPT-5.4-mini. Each trial was graded by two deterministic checks without any AI judging: whether the agent called an acceptable tool, and whether the saved model passed physical checks such as assembly R-values, HVAC loop membership, and pinned EnergyPlus outputs. The distinction proved essential, because in 18 of 23 outcome failures an agent replaced a roof assembly with one up to 1.86 square meters kelvin per watt worse while reporting success, an error visible only in the saved artifact, not in the agent&#8217;s confident report. Under a common configuration with all tool schemas loaded, GPT-5.4 and Opus 4.8 achieved 100 percent outcome rates, with Opus 4.6 at 95.8 percent, GPT-5.4-mini at 93.8, Sonnet at 91.7, and Haiku at 85.4.</p>
<p>The ablation results carry a practical lesson for anyone deploying agentic AI. Loading every tool schema up front raised the weakest model&#8217;s success rate by 12.5 points but inflated costs by 36 to 68 percent for stronger Claude tiers while changing outcomes by at most two tasks, meaning deferred schema discovery saves money for capable models at no accuracy loss. The curated knowledge layer, surprisingly, changed no model&#8217;s completion rate by more than 6.3 points. The completion budget also mattered: at a 120-second limit, 17 of Opus 4.8&#8217;s 18 failures were timeouts, yet it passed every trial in four of five configurations when given 600 seconds, showing that slow but productive work was being misclassified as failure. An unscaffolded baseline without the server showed agents can handle basic OpenStudio operations through direct scripting, but both tested models failed a task requiring exact counting of warnings in an EnergyPlus error file.</p>
<p>The researchers frame the work as broadening access rather than replacing rigor. Engineers and architects could request, test, and run bespoke retrofit analyses without writing Ruby, while organizations could offer shared modeling capacity to design firms, classrooms, or utility programs by issuing authentication tokens instead of provisioning workstations, turning energy modeling from a per-seat desktop activity into shared infrastructure. Validation and quality-assurance tools, audit records of every tool call, and artifact-based grading make the checking explicit, but the authors are careful to note that engineering judgment is not automated: generated artifacts and conclusions require review by a qualified practitioner before use in real engineering decisions. The code, benchmark harness, and archived trial records are openly available under a BSD-3-Clause-style license, inviting the building science community to put AI-driven modeling to the test.</p>
<p><strong>Subject of Research:</strong> An open-source Model Context Protocol server enabling AI agents to perform full-lifecycle building energy modeling with the OpenStudio SDK.</p>
<p><strong>Article Title:</strong> OpenStudio-MCP: a model context protocol (MCP) server for AI agent-driven building energy modeling with the OpenStudio SDK</p>
<p><strong>Article References:</strong> Ball, B. L., Long, N., Fleming, K., &amp; Goldwasser, D. (2026). OpenStudio-MCP: a model context protocol (MCP) server for AI agent-driven building energy modeling with the OpenStudio SDK. <em>SoftwareX, 36</em>, Article 103020. <a href="https://doi.org/10.1016/j.softx.2026.103020" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103020</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103020" rel="noopener noreferrer">10.1016/j.softx.2026.103020</a></p>
<p><strong>Keywords:</strong> OpenStudio-MCP, building energy modeling, Model Context Protocol, large language models, AI agents, EnergyPlus, OpenStudio SDK, HVAC synthesis, Measure authoring, ASHRAE 90.1, agent benchmark, sandboxing</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202051</post-id>	</item>
		<item>
		<title>When Knowledge Graphs Meet Large Language Models: A New Roadmap for Trustworthy AI Reasoning</title>
		<link>https://scienmag.com/when-knowledge-graphs-meet-large-language-models-a-new-roadmap-for-trustworthy-ai-reasoning/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:38:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI fact verification]]></category>
		<category><![CDATA[AI representation gap]]></category>
		<category><![CDATA[cognitive science in AI development]]></category>
		<category><![CDATA[cognitive synergy]]></category>
		<category><![CDATA[explainability in AI]]></category>
		<category><![CDATA[explainable AI]]></category>
		<category><![CDATA[factual hallucination]]></category>
		<category><![CDATA[hierarchical theoretical frameworks for AI]]></category>
		<category><![CDATA[knowledge agents]]></category>
		<category><![CDATA[knowledge graph and language model fusion]]></category>
		<category><![CDATA[Knowledge graph integration with large language models]]></category>
		<category><![CDATA[knowledge graphs]]></category>
		<category><![CDATA[knowledge reasoning]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[model fine-tuning]]></category>
		<category><![CDATA[multi-step reasoning challenges]]></category>
		<category><![CDATA[neuro-symbolic integration]]></category>
		<category><![CDATA[paradigm conflicts in AI]]></category>
		<category><![CDATA[prompt engineering]]></category>
		<category><![CDATA[retrieval-augmented generation]]></category>
		<category><![CDATA[synergy bottleneck in AI systems]]></category>
		<category><![CDATA[trustworthy AI reasoning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201848</guid>

					<description><![CDATA[A new survey maps how knowledge graphs can fix the hallucinations, weak reasoning, and opacity of large language models through five technical pathways.]]></description>
										<content:encoded><![CDATA[<p>Large language models have dazzled the world with their fluency, yet they remain haunted by a familiar trio of flaws: they invent facts, they stumble through multi-step reasoning, and they cannot explain why they said what they said. A new survey published in Knowledge and Information Systems argues that the antidote may already exist in one of artificial intelligence&#8217;s oldest and most reliable inventions, the knowledge graph, and it offers the most systematic map yet of how these two very different kinds of machine intelligence can be fused into something greater than either alone.</p>
<p>The review, authored by Jiale Wu, Yijiang Zhao, Zhuhua Liao, and Min Liu of Hunan University of Science and Technology, goes beyond the usual catalog of techniques that dominates the survey literature. Instead, the team builds a hierarchical theoretical framework, described as a root-phenomenon-consequence structure, grounded in the cognitive science of neuro-symbolic integration. The central claim is that the difficulties of combining knowledge graphs with large language models are not a scattered collection of engineering annoyances but the surface expressions of three deep, interlocking challenges: a representation gap, a synergy bottleneck, and a paradigm conflict.</p>
<p>The representation gap refers to the fundamental mismatch between how knowledge graphs and language models encode meaning. A knowledge graph stores the world as discrete triples, subject, relation, object, arranged in a symbolic network that a machine can traverse with perfect fidelity. A large language model, by contrast, compresses statistical regularities of language into billions of continuous neural parameters. One system reasons with symbols it can inspect; the other reasons with patterns it cannot. Bridging these two encodings without losing the strengths of either is the first root problem the survey identifies.</p>
<p>The synergy bottleneck concerns what happens when the two systems are actually coupled. Simply retrieving a subgraph and pasting it into a prompt does not guarantee that the model will use the evidence correctly, and fine-tuning a model on graph data can degrade the very linguistic competence that made it useful. The paradigm conflict, meanwhile, is more philosophical: symbolic systems are built for exact, verifiable inference, while neural systems are built for tolerant, probabilistic generalization. The survey argues that progress depends on recognizing these tensions explicitly rather than papering over them with ever-larger models.</p>
<p>To organize the technical landscape, the authors propose a five-dimensional taxonomy of integration approaches. The first dimension is prompt engineering, where knowledge graph content is translated into text and placed in the model&#8217;s context window. Methods in this family include chain-of-thought prompting over graphs, frameworks such as KG-GPT and KG-CoT, and systems like Think-on-Graph that guide a language model step by step along relevant knowledge paths. Prompting is attractive because it requires no retraining, but it consumes context space and depends heavily on the quality of the retrieved evidence.</p>
<p>The second dimension, retrieval augmented generation, has become one of the fastest-moving areas in the field. Rather than relying on a model&#8217;s frozen internal memory, these systems fetch structured evidence at question time. The survey traces an evolution from classic dense retrieval to graph-aware pipelines such as GNN-RAG, HyKGE for medical question answering, and the Think-on-Graph series, whose later versions employ multi-agent, dual-evolving context retrieval over heterogeneous graphs. Neurobiologically inspired architectures such as HippoRAG, which model long-term memory as a graph, illustrate how far the retrieval paradigm has drifted from simple keyword search toward something resembling structured recall.</p>
<p>The third dimension is model fine-tuning, in which knowledge graph information is baked into the model&#8217;s parameters. The survey highlights parameter-efficient techniques such as KG-Adapter, infuser-guided knowledge integration, and knowledge-graph-enhanced model editing, along with distillation approaches that transfer reasoning ability from large teachers to smaller students. Fine-tuning promises deeper integration than prompting, but it raises cost, rigidity, and knowledge-timeliness problems: a model trained on yesterday&#8217;s graph cannot easily learn today&#8217;s facts.</p>
<p>The fourth and fifth dimensions mark what the authors see as the field&#8217;s evolutionary frontier. Large reasoning model collaboration pairs knowledge graphs with the new generation of reasoning-heavy models, exemplified by reinforcement-learning-trained systems such as DeepSeek-R1, and by frameworks like KG-o1 and Search-o1 that let a reasoning model consult a graph during extended chains of deliberation. Knowledge agents, the final dimension, go further still: autonomous systems such as KG-Agent, AriGraph with its episodic memory, and Generate-on-Graph treat the language model as an agent that can plan, query, and even extend an incomplete knowledge graph on its own. The survey characterizes the overall trajectory of these five dimensions as a progression from external guidance toward autonomous cognition, a shift with profound implications for how much trust such systems can eventually earn.</p>
<p>The practical stakes are already visible in vertical domains. In health care, graph-augmented frameworks such as Medical Graph RAG, KoSEL, and MedReason ground clinical question answering in curated medical knowledge, reducing the risk of confidently wrong answers in settings where errors can harm patients. In law, systems like ChatLaw combine knowledge graphs with mixture-of-experts architectures to anchor legal reasoning in statutes and precedent, while researchers have explored how graph-based prompting can clarify the legal implications of model outputs. In scientific research, knowledge graphs are being coupled with language models for tasks ranging from biomedical literature mining in Alzheimer&#8217;s studies to automated retrosynthesis planning of macromolecules, where a model proposes reaction routes that a structured chemical knowledge base validates.</p>
<p>None of this amounts to a solved problem, and the survey is candid about the open challenges. Neuro-symbolic alignment remains immature: there is still no principled theory for when symbolic structure should override neural intuition or vice versa. Knowledge timeliness is equally pressing, since real-world facts change far faster than models or graphs can be updated, and stale evidence can be worse than none. The authors point toward lightweight reasoning engines that could make graph-guided inference affordable at scale, and toward high-trust agent systems in which every reasoning step is auditable against an explicit knowledge structure. If those directions mature, the hybrid of neural fluency and symbolic rigor that this survey maps out could become the architecture on which genuinely reliable AI reasoning is built, transforming language models from persuasive improvisers into accountable thinkers that show their work.</p>
<p><strong>Subject of Research:</strong> The integration of knowledge graphs with large language models to improve factual accuracy, reasoning, and explainability in AI systems.</p>
<p><strong>Article Title:</strong> A survey on knowledge graph-augmented large language model reasoning: theoretical challenges, technical pathways, and evolutionary logic</p>
<p><strong>Article References:</strong> Wu, J., Zhao, Y., Liao, Z., &amp; Liu, M. (2026). A survey on knowledge graph-augmented large language model reasoning: theoretical challenges, technical pathways, and evolutionary logic. <em>Knowledge and Information Systems, 68</em>(1), Article 261. <a href="https://doi.org/10.1007/s10115-026-02880-5" rel="noopener noreferrer">https://doi.org/10.1007/s10115-026-02880-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10115-026-02880-5" rel="noopener noreferrer">10.1007/s10115-026-02880-5</a></p>
<p><strong>Keywords:</strong> large language models, knowledge graphs, neuro-symbolic integration, retrieval augmented generation, prompt engineering, model fine-tuning, knowledge agents, factual hallucination, knowledge reasoning, cognitive synergy, AI agents, explainable AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201848</post-id>	</item>
		<item>
		<title>AI Agents and Smart Contracts Could Secure and Speed Last-Mile Medicine Delivery</title>
		<link>https://scienmag.com/ai-agents-and-smart-contracts-could-secure-and-speed-last-mile-medicine-delivery/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 02:03:41 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agent architecture for healthcare supply chain management]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI and blockchain integration in pharmaceutical logistics]]></category>
		<category><![CDATA[AI-powered route planning for pharmaceuticals]]></category>
		<category><![CDATA[blockchain]]></category>
		<category><![CDATA[blockchain smart contracts for medicine logistics]]></category>
		<category><![CDATA[cold chain]]></category>
		<category><![CDATA[crowdsourced healthcare supply chain management]]></category>
		<category><![CDATA[crowdsourcing]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[last-mile delivery]]></category>
		<category><![CDATA[last-mile medicine delivery challenges and solutions]]></category>
		<category><![CDATA[North African pharmaceutical distribution logistics]]></category>
		<category><![CDATA[pharmaceutical last-mile delivery optimization]]></category>
		<category><![CDATA[pharmaceutical logistics]]></category>
		<category><![CDATA[real-time pharmacy dispatching automation]]></category>
		<category><![CDATA[regulatory compliance in pharmaceutical supply chains]]></category>
		<category><![CDATA[smart contracts]]></category>
		<category><![CDATA[supply chain]]></category>
		<category><![CDATA[tamper-evident blockchain records for medicine custody]]></category>
		<category><![CDATA[temperature-sensitive medication transportation]]></category>
		<category><![CDATA[traceability]]></category>
		<category><![CDATA[vehicle routing]]></category>
		<category><![CDATA[verifiable credentials]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193442</guid>

					<description><![CDATA[Researchers have built a blockchain-and-AI-agent system for crowdsourced pharmaceutical delivery that cut simulated route time by 25 percent while keeping the ledger as the source of truth.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers at Hassan II University of Casablanca has combined blockchain smart contracts with large language model agents to orchestrate crowdsourced pharmaceutical last-mile delivery, reporting that in matched simulations the AI-assisted workflow cut route time by 25 percent and planning time by nearly 79 percent compared with a human-operated baseline. The work addresses a stubborn gap in pharmaceutical logistics: while blockchain systems excel at preserving a tamper-evident record of custody, they do little to help dispatchers forecast warehouse readiness, build routes, or choose suitable carriers in real time.</p>
<p>Pharmaceutical delivery differs from ordinary parcel logistics in critical ways. A late medicine can interrupt treatment, a temperature excursion can ruin the product, and a substituted or falsely confirmed delivery can harm a patient. Between a distributor and the final pharmacy or patient, traffic, varied vehicles, returns, and fluctuating demand make these obligations harder to satisfy. The Moroccan and wider North African context adds further complications, since small pharmacies and independent carriers may have uneven digital capacity, and any production system must comply with pharmaceutical, personal-data, and electronic-trust regulations.</p>
<p>The researchers&#8217; answer is a ledger-grounded agent architecture in which the blockchain, not the language model, remains the authoritative source of operational truth. Smart contracts govern identity status, order assignments, capacity bookings, delivery evidence, and returns, while a real-time index and a Neo4j graph database organize contract events into retrievable relationships. An AI agent can only act on an order, carrier, booking, or delivery status if that record exists in verified platform data, substantially reducing the risk of decisions built on fabricated or unsupported information.</p>
<p>Four specialized agents divide the work. A Main Agent monitors the operation, retrieves similar past scenarios from graph memory, and coordinates specialist workflows. A Forecasting Agent estimates order arrivals, warehouse readiness, dispatch needs, and the risk of missing wave cutoffs. A vehicle-routing agent constructs feasible route plans and evaluates urgent insertions, while a Dispatch Agent carries out carrier checks, assignments, bids, bookings, and approved state changes in sequence. Numerical services compute forecasts and routes using portfolios of statistical and deep-learning forecasting methods and optimization solvers including ant colony optimization, particle swarm optimization, genetic algorithms, simulated annealing, proximal policy optimization, and deep Q-learning.</p>
<p>The routing logic applies lexicographic priorities rather than blending metrics through arbitrary weights: unassigned urgent orders come first, then unassigned standard orders, then minutes of lateness, and finally distance as a tiebreaker. Feasibility conditions enforce capacity limits, vehicle and temperature compatibility, valid carrier identity and eligibility, booking consistency, and locked route prefixes. Motorbikes carry either cold-chain or ambient products within a route but never both, while suitably equipped cars can mix both classes when packaging and capacity constraints permit. Urgent orders are prepared within 30 to 45 minutes and dispatched immediately after staging, while standard orders consolidate into waves at 10:00, 15:00, and 17:00.</p>
<p>Crowdsourced carriers introduce a trust problem that the framework tackles through verifiable credentials. Enrollment requires an authorized issuer to check official identity documents, driving licences, photographs with liveness checks, vehicle registration, insurance, and any cold-chain qualification. The approved carrier signs a one-time challenge to prove control of a wallet key, and the issuer issues a short-lived W3C verifiable credential binding the carrier, wallet, and capabilities. The blockchain stores only the credential fingerprint, issuer, validity period, and revocation status; personal documents and photographs remain encrypted off-chain. Pickup, delivery, and return events then link product scans, carrier signatures, time and location evidence, and pharmacy confirmations to the same digital identity.</p>
<p>Evaluation used five fixed-seed simulated operating days totaling 1600 orders, 40 carriers per day, 288 urgent orders, and 448 cold-chain orders, with each day replayed once through the human-operated baseline and once through the agent-optimized flow under identical conditions. The agent flow reduced route time by 25.0 percent and planning time by 78.6 percent, with consistent improvements in distance, urgent-order handling, assignment readiness, and proof completeness. A composite efficiency index reached 126.4 against the baseline&#8217;s 100, and the favorable direction remained stable across sensitivity analyses of the weighting scheme. Smart-contract functions were benchmarked separately on a local Ethereum virtual machine and the Sepolia testnet, quantifying gas, fees, and confirmation times for registration, order creation, booking, bidding, pickup, proof, and completion.</p>
<p>The authors caution that the evidence comes from simulation, not production. Five day-level replications with one distribution center and synthetic pharmacy profiles cannot capture real traffic, carrier misconduct, connectivity loss, sensor behavior, or user interaction, and the study did not conduct a dedicated hallucination test. Cryptographic credentials cannot guarantee document authenticity, prevent credential lending, stop a stolen key before revocation, or certify that the approved person physically performed every delivery. The next phase will involve the collaborating Moroccan distributor, first replaying live ERP and warehouse events in shadow mode, then running a monitored single-center pilot with vetted carriers and pharmacies, and finally scaling load across centers, carriers, and transactions.</p>
<p>For deployment, the researchers propose a phased production architecture that keeps personal data off-chain, uses a permissioned ledger among authorized partners, isolates external model providers behind an API privacy gateway, and routes agent recommendations through a human approval boundary before any smart-contract commitment. Compliance reviews must cover Moroccan pharmaceutical law, personal-data protection under Law 09-08, and electronic trust services under Law 43-20, while WHO good distribution practices make calibrated sensors, excursion alerts, and documented corrective actions operational requirements rather than optional features.</p>
<p>Beyond the specific Moroccan case, the study offers a template for adding operational intelligence to blockchain-based traceability without sacrificing accountability: the ledger supplies durable facts, deterministic tools handle computation, the language model coordinates context and sequencing, and every consequential action passes through explicit rules, auditable evidence, and human oversight. As agentic AI systems move into safety-sensitive logistics, the authors argue that reliability must be measured, not assumed, and future work will include dedicated hallucination tests, model comparisons, and quantitative security evaluation of contracts, carriers, identity, and agent behavior.</p>
<p>The distinction between traceability and orchestration helps explain why earlier blockchain pharmaceutical systems, however robust their records, left dispatchers largely on their own. Prior platforms built on Ethereum and Hyperledger demonstrated that role-based contracts, distributed file storage, and IoT sensing could make missing or inconsistent custody events easier to detect, and recent systems have combined ledger provenance with counterfeit-detection classifiers and certificate validation. Yet none of these contributions centered on real-time dispatch under fluctuating crowdsourced capacity, which is precisely where the Casablanca team positions its work.</p>
<p>The choice of graph-based retrieval rather than conventional text retrieval reflects a deliberate design judgment. In a retrieval-augmented setup built on a property graph, an order can be linked to its products, required temperature class, booking, carrier, route, and delivery proofs as explicit relationships, so the agent retrieves structured context rather than isolated passages of text. This matters in logistics, where the meaning of a record depends on its position in a web of custody events, and it reduces the chance that an agent reasons from fragments disconnected from the operational state.</p>
<p>The architecture also acknowledges that tool interfaces themselves create an attack surface. Malicious or misleading tool schemas, compromised servers, and excessive permissions can all corrupt an agent&#8217;s behavior, which is why the framework relies on allow-lists, authentication, input validation, and trace logging for every external call. This defensive posture extends to the model&#8217;s own tendencies: recent evaluations of language-model factual recall penalize confident wrong answers and reward appropriate abstention, a principle the authors carry into their design by ensuring the agent can abstain or escalate rather than guess when verified data is absent.</p>
<p>The evaluation methodology deserves attention in light of the broader research landscape. Reviews of generative AI in logistics consistently find that empirical field evidence remains scarce and that simulation dominates the literature, a pattern this study follows knowingly. By replaying five fixed-seed days through both workflows under identical demand, carrier pools, and order mixes, the researchers created a controlled comparison in which differences in route time, planning time, and proof completeness can be attributed to the coordination method rather than to environmental noise. The separate benchmarking of contract functions on a local virtual machine and a public testnet similarly isolates the cost of on-chain accountability from the benefits of agent coordination.</p>
<p>The lexicographic routing priorities also connect to a long-standing debate in optimization practice. Weighted objective functions force planners to express trade-offs in commensurable units, and small weight changes can silently reorder priorities in ways operators cannot anticipate or audit. By ranking urgent assignments, standard assignments, lateness, and distance in strict order, the framework makes every routing decision explainable as a sequence of dominance checks, which aligns naturally with the accountability requirements of pharmaceutical custody and with the need for authorized staff to understand and contest agent recommendations.</p>
<p>The identity model likewise reflects emerging standards beyond the pharmaceutical domain. Binding a short-lived verifiable credential to a proven wallet key, storing only fingerprints and revocation status on-chain, and keeping documents encrypted off-chain follows the privacy principle that blockchain should attest to facts without becoming a repository of personal data. Expiry and revocation mechanisms acknowledge that carrier eligibility is a continuing state, not a one-time gate, addressing risks such as credential lending or stolen keys only partially, as the authors themselves note.</p>
<p>Ultimately, the study&#8217;s contribution lies less in any single technique than in the division of labor it enforces: durable facts from the ledger, deterministic computation from numerical services, contextual coordination from the language model, and consequential state changes gated by contracts and human approval. Whether that division survives contact with live traffic, real carriers, and regulatory inspection is the question the planned shadow-mode replay and monitored pilot are designed to answer.</p>
<p><strong>Subject of Research:</strong> Integration of smart contracts and AI agents for secure, traceable crowdsourced last-mile pharmaceutical delivery</p>
<p><strong>Article Title:</strong> Smart contracts and AI agents for secure last-mile pharmaceutical delivery through crowdsourcing</p>
<p><strong>Article References:</strong> Nadime, K. L., Haidar, D., Benabbou, R., &amp; Benhra, J. (2026). Smart contracts and AI agents for secure last-mile pharmaceutical delivery through crowdsourcing. <em>Discover Artificial Intelligence, 6</em>(1), Article 1128. <a href="https://doi.org/10.1007/s44163-026-02206-y" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02206-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02206-y" rel="noopener noreferrer">10.1007/s44163-026-02206-y</a></p>
<p><strong>Keywords:</strong> blockchain, smart contracts, AI agents, last-mile delivery, pharmaceutical logistics, crowdsourcing, large language models, vehicle routing, cold chain, traceability, verifiable credentials, supply chain</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193442</post-id>	</item>
	</channel>
</rss>
