Artificial intelligence is moving rapidly from systems that answer questions to agents that can pursue objectives, make decisions and act across digital environments. That shift is creating a governance problem that cannot be solved by measuring model performance alone, according to a new paper in Nature. Researchers Atoosa Kasirzadeh and Iason Gabriel propose a framework for understanding AI agents through four connected dimensions: autonomy, efficacy, goal complexity and generality. Their central argument is that the word “agent” covers systems with radically different capabilities, risks and oversight needs—and that treating them as one category could leave major gaps in regulation and public safety.
The first dimension, autonomy, describes how independently an AI system can operate. A low-autonomy assistant may generate a draft, recommend an action or wait for a human command before doing anything further. A more autonomous agent may break a broad instruction into subtasks, select tools, execute actions and adapt its behaviour based on changing conditions. At the highest levels, an agent could pursue a goal over extended periods with limited human intervention. Technically, this may involve planning loops, memory systems, tool access, environmental feedback and the ability to revise a sequence of actions. Each additional layer of independence changes the governance challenge, because oversight must move from checking individual outputs to monitoring an ongoing process.
Efficacy is the second dimension and refers to an agent’s ability to achieve intended outcomes in the world. A system can be highly autonomous yet ineffective, repeatedly taking actions without producing useful results. Conversely, an agent with modest independence may be extremely effective within a narrow domain, such as scheduling resources, analyzing specialized data or controlling a well-defined industrial process. Efficacy depends on more than language fluency. It can be shaped by access to reliable information, the quality of an agent’s planning and reasoning, its ability to use external software, the accuracy of its internal representations and the consequences of errors. As efficacy rises, failures may become less frequent—but successful actions can also create greater impact when the system is operating without close supervision.
The third dimension, goal complexity, concerns the structure of the objectives an AI agent is asked to pursue. Simple goals may involve a single, clearly defined task with an obvious completion condition. Complex goals can require an agent to balance competing priorities, interpret ambiguous instructions, coordinate multiple steps and respond to unexpected events. A request such as “find the best available option” may conceal questions about cost, reliability, fairness, timing, privacy and risk. For an AI system, translating such a request into operational decisions is not merely a matter of generating text. It requires selecting subgoals, assigning priorities and determining when a result is good enough. Governance must therefore address how goals are specified, who has authority to define them and how conflicts between objectives are resolved.
Generality forms the fourth axis. Narrow agents are designed for a limited task, environment or professional setting, where developers can test performance against a relatively stable set of conditions. General-purpose agents are intended to function across many domains, potentially moving between research, administration, programming, education and other activities. Greater generality can make systems more useful, but it also expands the range of situations in which their behaviour may be difficult to predict. A narrow system may be evaluated using domain-specific benchmarks and safety checks. A general system requires broader testing, because risks can emerge from combinations of capabilities that were not present in any single training example or evaluation scenario.
Together, the four dimensions produce what the authors call “agentic profiles.” Rather than asking whether a system is simply an AI agent, the framework asks what kind of agent it is and how its properties interact. A task-specific assistant might have low autonomy, moderate efficacy, simple goals and limited generality. An enterprise system that can access databases, coordinate workflows and act on behalf of employees could occupy a different profile, with higher efficacy and goal complexity. A highly autonomous, general-purpose agent would sit at the most demanding end of several axes, raising questions about supervision, accountability, security and control. The profiles are not intended as rigid labels; they are maps for comparing systems that may look similar on the surface but have very different real-world implications.
This approach also highlights why governance cannot rely on one universal safety measure. An autonomous agent with access to financial systems presents different concerns from a highly effective but closely supervised scientific assistant. A system with complex goals may require mechanisms for clarifying instructions, exposing intermediate plans and detecting conflicts between objectives. A general-purpose agent may need continuous monitoring across contexts rather than a one-time certification. Technical safeguards could include permission boundaries, sandboxed tool use, audit logs, human approval checkpoints, uncertainty reporting and emergency shutdown procedures. But the paper’s framework also points beyond engineering: organizations must determine who is responsible when an agent causes harm, how affected people can challenge decisions and what forms of deployment should require public oversight.
The researchers’ proposal arrives as AI developers increasingly describe their products as capable of acting rather than merely responding. Modern agents can combine large language models with retrieval systems, code execution, web browsing, APIs and persistent memory. In a typical control loop, the model receives an objective, observes information from its environment, generates a plan, invokes tools, evaluates results and repeats the process. This architecture can turn a probabilistic text generator into a system with practical influence over external events. The same loop that allows an agent to complete useful multi-step work can also amplify a mistaken assumption, an ambiguous instruction or a malicious input. Understanding an agent’s profile therefore becomes essential before deciding how much access, authority and autonomy it should receive.
The paper presents agentic profiles as a common language for developers, policymakers and the public at a moment when the boundaries between software and decision-maker are becoming increasingly blurred. The framework does not claim that every risk can be predicted from four dimensions, nor does it replace detailed technical evaluation. Its value lies in making differences visible: two systems may both be marketed as assistants, while one only suggests actions and the other executes them; two may be equally capable, while one operates in a narrow environment and the other generalizes across domains. By mapping autonomy, efficacy, goal complexity and generality, the authors argue, society can design governance mechanisms that match the actual power of an AI agent—before usefulness, scale and independence outpace accountability.
Subject of Research: AI agent capabilities and governance
Article Title: Agentic profiles for effective AI governance
Article References: Kasirzadeh, A., Gabriel, I. Agentic profiles for effective AI governance. Nature 656, 320–328 (2026). https://doi.org/10.1038/s41586-026-10805-z
Image Credits: AI Generated
DOI: 10.1038/s41586-026-10805-z
Keywords: artificial intelligence, AI agents, AI governance, autonomy, efficacy, goal complexity, generality, agentic profiles, AI safety, regulation

