For more than a decade, the dominant metaphor for working with artificial intelligence has been the human-in-the-loop: a person standing guard over a machine, checking its outputs, correcting its mistakes, and ultimately deciding what happens next. A new open-access study from researchers at the University of Bologna argues that this picture is now dangerously incomplete. As large language models become capable partners in complex cognitive work, the researchers contend, the weakest link in the human-AI team is often not the algorithm but the person using it. Their answer is a conceptual inversion they call AI-in-the-human-loop, in which the AI agent observes, models, and gently assesses the human’s own reasoning, knowledge, and behaviour, helping the team mature together rather than merely serving as an obedient tool.
The study, published in the Journal of Ambient Intelligence and Humanized Computing by Roberto Casadei, Giovanni Delnevo, Barry Bassi, Chiara Ceccarini, and Silvia Mirri, is a position paper with an empirical backbone. Its central claim is that effective collaboration, like any good teamwork, requires all members to understand the problem, the environment, and each other. Yet while a vast literature examines the strengths and limitations of AI systems, far less attention has been paid to the risks created by a lack of maturity on the human side: users who over-trust plausible-sounding outputs, delegate too much reasoning, or hold mental models of the AI that bear little resemblance to what the system can actually do. The authors propose a formally structured model in which both actors continuously and proactively assess one another, closing the gap through what they describe as bidirectional scaffolding.
Technically, the framework models a human-AI team as a dynamic system state containing the user’s knowledge, behaviour, and configuration, the AI agent’s corresponding state, the shared goals, the task context, and the full process history, including logs, outputs, and plans. Crucially, each actor holds only its own view of this state, and those views may diverge from the ground truth and from each other. Collaboration falters when these internal representations drift apart. The proposed process framework therefore runs a cognitive pipeline: a grounding and calibration phase aligns user intent with AI capabilities; a cognition-tracking layer captures actions, decisions, and reasoning traces through what the authors call a Human Cognitive Interface; and a cognitive state representation feeds a simulation engine that generates candidate plans and hypothetical reasoning trajectories. Adaptive feedback, nudging, and targeted user assessments are then calibrated to the maturity of both partners.
Two novel concepts anchor the vision. The first, AI-in-the-human-loop, positions the AI not as a subordinate awaiting validation but as a reflective partner embedded within the human’s cognitive loop, observing and mirroring mental processes to foster metacognitive awareness and epistemic growth. The second, human-under-test, borrows an analogy from software engineering. Just as a subject-under-test is probed by test cases to verify and validate a system, the human partner can be gently assessed by the AI agent, with tests planned to cover key risks, detect regressions in understanding, and check compliance with shared requirements and factual knowledge. The authors are careful to note that the analogy is meant to highlight goals and procedures, not to dehumanise: the aim is to identify risks and promote quality, never to assign blame.
To test whether any of this is observable in practice, the team analysed a real-world collaboration: a programmer with no prior expertise in statistics who used ChatGPT 5.4 to analyse a questionnaire administered to secondary school students about craft-oriented education and the local footwear sector. The corpus comprised eleven documented chats, thirty human turns and thirty corresponding LLM turns, moving from broad exploratory questions about how to begin, through data auditing and item-type-specific analyses of open-ended, binary, single-choice, multiple-choice, ordinal, and reversed-Likert questions, to synthesis work producing charts, subgroup comparisons, and reusable scripts. The researchers then examined this workflow retrospectively using an LLM-as-a-judge setup, with two models from different providers, GPT 5.4 and Gemini 2.0 Flash, acting as independent evaluators across four dimensions: quality of the initial context, interaction issues, collaboration maturity, and missed opportunities for proactive AI intervention.
The findings are striking precisely because the project was, by conventional measures, a success. Both judges agreed that the initial context was generally good, with goals, data, and deliverables mostly clear. Yet the corpus still revealed non-trivial frictions. The clearest was an incorrect count produced by the operational LLM when processing a JSON file containing duplicate keys, an error caught only because the human explicitly checked it. Judges also flagged mismatches between the granularity of the model’s coding of open responses and the coding later adopted by the human, difficulties with artefact availability and format, and the need to explicitly request reusable outputs. Human researchers added further methodological concerns, including the formal treatment of grid-style questionnaire items and the validation of some statistical outputs, suggesting that automated judges may still need complementing by expert review.
Both evaluators also detected a maturation trajectory: early chats were exploratory and framing-oriented, while later interactions became structured by question type and oriented toward concrete deliverables. Interestingly, the two judges disagreed on where the collaboration ultimately landed. GPT 5.4 offered a cautious assessment, placing the human-AI dyad mainly between the Repeatable and Defined levels of a Humphrey-style maturity scale, while Gemini was more optimistic, situating it between Defined and Managed. The authors preserved rather than averaged this disagreement, noting that it may reflect differences between the judge models themselves. The most robust result, they argue, is the convergence in identifying recurring patterns and concrete points where a more proactive model, clarifying file semantics, agreeing on coding granularity, confirming which CSV columns to analyse, or setting minimal rules for subgroup comparisons, would have helped.
The study is candid about its limits. It examines a single analyst, a single operational model, a single task, and eleven chats; the analysis is retrospective rather than embedded in real-time interaction; the LLM-as-a-judge methodology may carry model-specific biases; and the framework itself has not yet been implemented as a complete system. Generalisability across users, domains, and tasks therefore remains an open question. Still, the exploratory evidence supports the paper’s core motivation: even a successful project can harbour maturity issues, collaborative frictions, and methodological misalignments that no one notices until an external assessor looks for them. In a world where millions of knowledge workers now lean on conversational AI daily, that observation alone is unsettling.
The broader implications reach into some of the most pressing debates in contemporary AI research. The paper situates its proposal against evidence that users tend to over-rely on AI advice even when it conflicts with contextual information, that LLM sycophancy can inflate trust without improving reliability, and that hallucinations persist even in retrieval-augmented systems, with models sometimes internally encoding the correct answer while generating a wrong one. It also draws on work showing that accurate user mental models improve decision-making, and that effective systems should promote calibrated trust rather than maximal trust. The authors’ framework treats these not as isolated problems but as symptoms of a missing feedback loop on the human side of the partnership, one that agentic AI, with its growing autonomy, makes increasingly urgent to close.
Many open questions remain before AI-in-the-human-loop becomes a practical architecture. What kinds of tests suit different aspects of a human-under-test, and how should they be generated contextually? How can expectations be defined, compliance quantified, and coverage estimated when human behaviour is fundamentally non-deterministic? How can testing be planned to avoid bothering the user, and designed to respect privacy and ethics? And, pointedly, who evaluates the evaluators? The Bologna team frames these as research challenges rather than settled answers, and calls for prototypes and quantitative studies to validate the vision. But the conceptual shift they propose is already provocative: AI systems designed not merely to perform tasks with users, but to help users better understand themselves, turning every collaboration into an opportunity for mutual learning and co-evolution.
Subject of Research: An operational model for bidirectional assessment and maturity in human-AI collaboration
Article Title: AI-in-the-human-loop: an operational model for mature human-AI collaboration
Article References: Casadei, R., Delnevo, G., Bassi, B., Ceccarini, C., & Mirri, S. (2026). AI-in-the-human-loop: an operational model for mature human-AI collaboration. Journal of Ambient Intelligence and Humanized Computing. https://doi.org/10.1007/s12652-026-05135-x
Image Credits: AI Generated
DOI: 10.1007/s12652-026-05135-x
Keywords: human-AI collaboration, AI-in-the-human-loop, human-under-test, large language models, metacognitive scaffolding, LLM-as-a-judge, mental models, trust calibration, human-AI co-evolution, AI maturity, hallucination, agentic AI
Cite Scienmag News
Glenn Wilkins. (October 3, 2026). Flipping the Loop: When AI Watches the Human, Not Just the Task. Scienmag. https://scienmag.com/flipping-the-loop-when-ai-watches-the-human-not-just-the-task/
Glenn Wilkins. "Flipping the Loop: When AI Watches the Human, Not Just the Task." Scienmag, 3 October 2026, https://scienmag.com/flipping-the-loop-when-ai-watches-the-human-not-just-the-task/. Accessed 3 October 2026.
Glenn Wilkins. "Flipping the Loop: When AI Watches the Human, Not Just the Task." Scienmag. October 3, 2026. https://scienmag.com/flipping-the-loop-when-ai-watches-the-human-not-just-the-task/

