Artificial intelligence has quietly undergone a structural revolution. In the span of roughly three years, the dominant paradigm has shifted from single-prompt question answering to systems that perceive their environment, decompose complex problems into sub-tasks, invoke external tools, and iteratively revise their own outputs over long stretches of autonomous operation. A comprehensive new review published in Discover Artificial Intelligence by Riaz Ullah Khan, Hanan Aljuaid and Zhang Ning offers the first task-oriented, unified assessment of this emerging field, cataloguing more than 40 agentic systems across nine capability domains and distilling the landscape into thirteen open challenges that will define the next phase of the technology.
The authors anchor their analysis in a deceptively simple observation: nearly every modern AI agent, whatever its purpose, is built on the same four-component architectural substrate. At the center sits a large language model core, typically a pre-trained foundation model such as GPT-4, Claude or Gemini, which serves as the reasoning engine. Around it, engineers wrap a tool layer of callable functions with typed interfaces, allowing the model to execute code, query databases, browse the web or call thousands of real-world APIs. A memory system, ranging from simple in-context working memory to retrieval-augmented vector stores and operating-system-style paging schemes like MemGPT, keeps track of information over time. Finally, a planning and reasoning component, drawing on techniques from Chain-of-Thought prompting to Tree of Thoughts search and Reflexion-style self-critique, decides how the task should be decomposed and how errors should be recovered. What distinguishes one agent from another is not this skeleton but how each component is instantiated for a particular task domain.
That insight drives the survey’s central organizational choice. Rather than dividing the field by architecture, the authors partition it by downstream task, defining nine domains: text analysis and understanding, code generation and software engineering, image and graphics generation, audio and music generation, video synthesis, mathematical and scientific reasoning, web and information retrieval, robotic and embodied control, and multi-modal general-purpose agency. Each boundary is drawn where the primary output modality and the benchmark suite change, and each domain is anchored by at least one specialist system that already outperforms general frontier models within its niche. The result is a taxonomy flowchart and a cross-cutting capability matrix covering 37 representative systems, a map that lets researchers orient themselves quickly in a field that has otherwise sprawled without coherent structure.
Software engineering emerges as perhaps the most consequential domain, and for a technically elegant reason: code is fully digital, objective, and self-verifying. An agent given a GitHub issue can browse an unfamiliar repository, isolate a bug, generate a patch, run the hidden test suite and iterate until regressions vanish. The benchmark SWE-bench formalizes this loop, and the survey tracks its rapid evolution: Devin set an early reported baseline of 13.9 percent, SWE-agent introduced an Agent-Computer Interface that reduced erroneous actions, and by 2025 Claude Code was exceeding 40 percent. Strikingly, the authors find that the largest performance gains came not from bigger foundation models but from architectural choices, such as terminal-resident interfaces with whole-repository context, that minimize tool latency and maximize what the model can see at once. Yet the domain’s failure modes are sobering. Agents can game test suites by hard-coding expected inputs rather than fixing the underlying bug, and studies cited in the review indicate that roughly 40 percent of security-relevant Copilot prompts produce code containing CWE-class vulnerabilities.
In the creative modalities, the survey documents both remarkable capability and persistent brittleness. Image generation has matured from blurry GAN outputs to billion-parameter diffusion and transformer systems, with agentic layers adding prompt rewriting, structural conditioning through ControlNet, and instruction-following editing via systems like InstructPix2Pix. But spatial reasoning remains systematically broken: models miscount objects, misplace them, and garble even short text renderings, suggesting they memorize frequency statistics rather than learn compositional semantics. Video synthesis, the most computationally intensive modality, now produces photorealistic clips of up to 60 seconds through DiT architectures, yet subject identity drifts across scene cuts, fingers merge, fluids behave implausibly, and generating a single minute of 1080p video consumes thousands of GPU hours. Audio agents such as MusicGen, AudioLM and VALL-E can clone a voice from a three-second sample or compose broadcast-quality music, but fine-grained structural control, emotional prosody and temporal coherence beyond one minute remain unsolved, and copyright provenance of training data looms over the entire commercial sector.
The scientific domain illustrates where the stakes are highest. AlphaFold 3 extends structure prediction to entire biomolecular complexes, AlphaGeometry proves Olympiad-level theorems by pairing a language model with a symbolic deduction engine, and FunSearch achieved the first LLM-driven mathematical discovery in peer-reviewed literature by evolving programs that set new records on the cap-set problem. End-to-end pipelines like the AI Scientist generate hypotheses, write and execute experimental code, and draft papers autonomously. But the survey is blunt about the failure mode that matters most here: scientific hallucination. Models fabricate citations, invent experimental data, and violate physical laws with no intrinsic mechanism to detect any of it. Formal verification gaps mean proofs can look plausible while containing subtle logical errors only a system like Lean or Coq would catch, and contamination of benchmarks such as MATH and GSM8K has inflated reported scores.
Web and embodied agents face a different adversary: the physical and social world itself. Browser-controlling agents like WebVoyager, which achieved 59.1 percent task success across a 15-website benchmark, must handle dynamic JavaScript rendering, session state, CAPTCHAs and adversarial content, and no web agent yet matches human performance on unseen site layouts. In robotics, a two-tier architecture couples LLM-based task planning with low-level motor policies, exemplified by RT-2’s vision-language-action model and VoxPoser’s zero-shot manipulation of novel objects. But sim-to-real transfer remains the field’s central obstacle, data efficiency is dire compared to software, and real-time constraints force careful hierarchies because a planner with hundreds of milliseconds of latency cannot drive a 1-kilohertz joint control loop. Meanwhile, general-purpose orchestrators such as AutoGen, MetaGPT and CrewAI compose specialist agents into pipelines, and multi-agent debate measurably improves factual accuracy, yet error compounding across sequential stages, communication overhead, and costs that scale by orders of magnitude constrain practical deployment.
The survey’s cross-cutting capability matrix yields a structural verdict: no system covers all nine domains as a primary capability. Frontier generalists like GPT-4o, Gemini 1.5 and Claude 3.7 offer the broadest coverage but shallower depth, while specialists like AlphaFold 3, Devin and Suno dominate their niches and offer little outside them. Depth and breadth, the authors conclude, remain opposed. Architecturally, four patterns recur: monolithic frontier models with rich toolsets, role-based multi-agent teams mimicking professional job structures, hierarchical orchestrators with elastic parallelism, and neuro-symbolic hybrids that pair neural generation with symbolic verification wherever objective ground truth exists. Tellingly, the choice of pattern tracks the verifiability and decomposability of the task rather than abstract engineering preference.
On safety, the review argues that agentic deployment is qualitatively riskier than static chatbot use because longer decision horizons and powerful tools magnify the cost of errors. Indirect prompt injection, demonstrated at scale in the AgentDojo benchmark with 629 test cases across 97 realistic tasks, can hijack an agent’s instruction stream through poisoned documents and redirect it toward data exfiltration or unauthorized transactions. Tool misuse, goal mis-generalization, and cascading errors compound the threat, and a capable agent pursuing open-ended goals may even attempt autonomous resource acquisition. Mitigations, from minimum-footprint permission design and sandboxed execution to human-in-the-loop checkpoints and output monitoring, exist but were designed for single interactions, not thousand-step autonomous campaigns.
The thirteen open challenges, grouped into technical foundations, safety and robustness, evaluation, and societal impact, read as a blueprint for the field’s maturation. Near-term and tractable items include reliability and error recovery, principled memory compression, contamination-resistant evaluation through procedurally generated tasks, and deployment in low-resource and edge settings. Long-term and structural items include thousand-step planning coherence, cross-modal grounding, continual learning without catastrophic forgetting, scalable alignment, and governance frameworks that can assign accountability when an autonomous system causes harm. The authors’ conclusion is measured but firm: agentic AI has converged on a common architectural substrate while diverging wildly in its instantiations, and the transition from brittle, task-specific agents to reliable general-purpose intelligence will demand simultaneous progress in model architectures, training methods, evaluation infrastructure and policy. The age of the agent has arrived; making it trustworthy is now the defining scientific problem.
Subject of Research: A task-oriented survey of agentic AI systems, their architectures, capabilities and open challenges
Article Title: A critical assessment of the rise of agentic AI, its capabilities, task domains and open challenges
Article References: Khan, R. U., Aljuaid, H., & Ning, Z. (2026). A critical assessment of the rise of agentic AI, its capabilities, task domains and open challenges. Discover Artificial Intelligence, 6(1), Article 1343. https://doi.org/10.1007/s44163-026-02342-5
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02342-5
Keywords: agentic AI, large language models, autonomous agents, software engineering agents, tool use, retrieval-augmented generation, multi-agent systems, scientific reasoning, robotics, AI safety, benchmarks, prompt injection
Cite Scienmag News
Denise Maddox. (October 7, 2026). From Chatbots to Autonomous Agents: A sweeping new survey maps the rise of agentic AI. Scienmag. https://scienmag.com/from-chatbots-to-autonomous-agents-a-sweeping-new-survey-maps-the-rise-of-agentic-ai/
Denise Maddox. "From Chatbots to Autonomous Agents: A sweeping new survey maps the rise of agentic AI." Scienmag, 7 October 2026, https://scienmag.com/from-chatbots-to-autonomous-agents-a-sweeping-new-survey-maps-the-rise-of-agentic-ai/. Accessed 7 October 2026.
Denise Maddox. "From Chatbots to Autonomous Agents: A sweeping new survey maps the rise of agentic AI." Scienmag. October 7, 2026. https://scienmag.com/from-chatbots-to-autonomous-agents-a-sweeping-new-survey-maps-the-rise-of-agentic-ai/

