Every time you ask a smart home to dim the lights, lock the door, and start the coffee maker in one breath, an invisible negotiation takes place. Somewhere in the background, software must find the right services scattered across dozens of devices, chain them together into a working sequence, and deliver the result before your patience runs out. Researchers call this service composition, and it is one of the thorniest unsolved problems in the Internet of Things. Now a team at the University of Kashan in Iran has proposed a new way to solve it, using a form of artificial intelligence in which many small agents learn to cooperate rather than compete. Their framework, described in the journal Cluster Computing, is called HAMRL, short for Hierarchical Attention-based Multi-agent Reinforcement Learning, and in their experiments it outperformed several of the strongest existing methods on the market.
The problem HAMRL tackles is deceptively simple to state but brutally hard in practice. An IoT environment is a shifting landscape of heterogeneous devices: sensors, actuators, gateways, and cloud services, each with different capabilities, response times, and reliability levels. A user request rarely maps onto a single service. Instead, it demands a composition, a carefully ordered chain of services that together fulfill the request while respecting Quality of Service constraints such as latency, availability, and cost. In a static network, engineers could hand-craft these chains. But IoT networks are anything but static. Devices join and leave, wireless links degrade, workloads spike unpredictably, and the composition that worked perfectly yesterday may collapse today. Scalability compounds the difficulty: as the number of devices grows into the hundreds or thousands, the space of possible service combinations explodes combinatorially, overwhelming classical optimization techniques.
Reinforcement learning has emerged as a natural candidate for this challenge. In reinforcement learning, an agent learns by trial and error, taking actions in an environment and receiving rewards or penalties that gradually shape its behavior toward better outcomes. The approach has produced landmark results, from game-playing systems that reached human-level control to adaptive traffic signal networks. Applied to service composition, a learning agent can, in principle, discover which service chains deliver the best quality without needing an explicit model of every device. But single-agent learning hits a wall in large IoT systems, because no single agent can observe or manage the entire network. The field therefore turned to multi-agent reinforcement learning, where each IoT node hosts its own agent and the agents must coordinate. Coordination, however, introduces its own pathology: when every agent is learning simultaneously, each one perceives the others as part of a constantly changing environment, and the learning process can become unstable, a phenomenon researchers describe as the moving target problem.
The Kashan team, Sara Rezaei, Salman Goli-Bidgoli, and Fereshte Dehghani, designed HAMRL around a well-known architectural compromise called centralized training with decentralized execution. The framework combines three ingredients: decentralized actors, a centralized critic, and a hierarchical attention mechanism. Each actor is a small neural network that lives with an individual IoT agent and decides which service to select based on what that agent can locally observe. During training, however, a centralized critic, a second neural network with access to global information, evaluates how good the joint decisions of all agents actually were, providing each actor with a far more stable learning signal than any agent could compute alone. This division of labor, rooted in the actor-critic family of algorithms first formalized in 2000, means the system can learn with global oversight but operate without it, which is essential for real deployments where a global view may not exist at runtime.
The genuinely novel component is the hierarchical attention mechanism. Attention, the computational idea that now underpins modern language models, allows a neural network to weigh the relevance of different pieces of information rather than treating them all equally. HAMRL applies this idea at two levels. At the lower level, each agent fuses its local context, the state of its own device and immediate neighborhood, with global contextual information about the wider network, learning which of the many available signals actually matter for the decision at hand. At the higher level, the framework organizes these attention weights hierarchically, so that agents can reason about relationships across different scales of the network, from a single sensor cluster up to the system as a whole. The practical effect is that an agent choosing a service does not drown in irrelevant data from thousands of distant devices; it learns to focus on the handful of peers and conditions that genuinely influence quality of service.
This fusion of local and global context is what the authors identify as the key to HAMRL’s improved scalability and adaptability. Because each agent’s decision module only attends to the most relevant information, the computational burden does not grow uncontrollably as the network expands. And because the attention weights are learned rather than fixed, the system can re-prioritize its focus when conditions change, for example when a popular service becomes overloaded or a new device with better latency joins the network. Adaptability of this kind is precisely what static composition methods and even earlier learning-based approaches have struggled to deliver under the dynamic conditions that define real IoT deployments.
To find out whether these design choices actually pay off, the researchers ran extensive experiments comparing HAMRL against four state-of-the-art multi-agent reinforcement learning baselines: MAAC, an actor-attention-critic method; COMA, a counterfactual multi-agent policy gradient algorithm; MAA2C, a multi-agent extension of the widely used advantage actor-critic method; and MAPPO, a multi-agent version of proximal policy optimization that has surprised researchers with its effectiveness in cooperative settings. The evaluation focused on three criteria: the quality of service achieved by the composed solutions, the stability of the learning process, and adaptability when environmental conditions shifted. Across all three, HAMRL significantly outperformed the competition, according to the study. The hierarchical attention mechanism appears to give the agents a clearer, less noisy picture of their environment, which translates into steadier learning curves and better final compositions.
The implications reach well beyond benchmark experiments. Service composition sits at the heart of smart infrastructure: intelligent buildings, industrial automation, healthcare monitoring, smart cities, and logistics networks all depend on chaining device-level services into coherent applications. A composition framework that adapts in real time, maintains quality of service under load, and scales to large device populations could make these systems more resilient and less dependent on human reconfiguration. The study also contributes to a broader conversation in machine learning about the role of attention in multi-agent systems, a topic that has attracted growing interest as researchers seek principled ways to tame the complexity of many-agent coordination. By structuring attention hierarchically rather than flatly, HAMRL offers a template that other domains, from drone flocking to traffic control, might borrow.
None of this means the problem is solved. The authors frame HAMRL as a promising solution for next-generation scalable and resilient IoT applications, a carefully chosen phrase that acknowledges the gap between controlled experiments and messy production networks. Real deployments raise questions the paper’s benchmarks cannot fully answer: how the framework behaves under adversarial conditions, how it handles security and privacy constraints in service selection, and how it integrates with existing IoT protocols and standards. The researchers also note that the data supporting their findings are available from the corresponding author upon request, inviting the community to scrutinize and extend the results. Still, the central lesson stands. In ecosystems where no single controller can exist, intelligence has to be distributed, and the hardest part is not making individual agents smart but making them coordinate. HAMRL’s answer, letting agents learn what to pay attention to, locally and globally, at multiple scales, is an elegant step toward IoT networks that quietly assemble themselves, service by service, while the rest of us just ask for coffee.
Subject of Research: Multi-agent reinforcement learning for quality-of-service-aware IoT service composition
Article Title: HAMRL: a hierarchical attention-based multi-agent reinforcement learning approach for IoT-based service composition
Article References: Rezaei, S., Goli-Bidgoli, S., & Dehghani, F. (2026). HAMRL: a hierarchical attention-based multi-agent reinforcement learning approach for IoT-based service composition. Cluster Computing, 29(15), Article 832. https://doi.org/10.1007/s10586-026-06596-7
Image Credits: AI Generated
DOI: 10.1007/s10586-026-06596-7
Keywords: Internet of Things, service composition, multi-agent reinforcement learning, hierarchical attention, quality of service, centralized critic, decentralized actors, actor-critic algorithms, scalability, adaptability, smart infrastructure, Cluster Computing
Cite Scienmag News
Denise Maddox. (October 9, 2026). Smart Devices Learn to Team Up: New AI Framework Tames Chaotic IoT Networks. Scienmag. https://scienmag.com/smart-devices-learn-to-team-up-new-ai-framework-tames-chaotic-iot-networks/
Denise Maddox. "Smart Devices Learn to Team Up: New AI Framework Tames Chaotic IoT Networks." Scienmag, 9 October 2026, https://scienmag.com/smart-devices-learn-to-team-up-new-ai-framework-tames-chaotic-iot-networks/. Accessed 9 October 2026.
Denise Maddox. "Smart Devices Learn to Team Up: New AI Framework Tames Chaotic IoT Networks." Scienmag. October 9, 2026. https://scienmag.com/smart-devices-learn-to-team-up-new-ai-framework-tames-chaotic-iot-networks/

