Every second, billions of connected devices generate tasks that demand immediate answers, from smart factory sensors to augmented reality headsets. Most of these devices are too small, too slow, and too battery-starved to do the heavy computational lifting themselves. The solution that the industry has converged on is multi-access edge computing, or MEC, which places resource-rich servers at the edge of the network, physically close to the devices that need them. But proximity alone does not solve the problem. Someone, or something, must decide in real time whether each incoming task should be executed on the device itself or shipped over a wireless link to the edge server. Making that decision badly wastes energy, inflates delays, or both. A new study published in Cluster Computing by Hala Elhadidy, Heba Saleh, Rawya Rizk, and Walaa Saber of Port Said University and Badr University in Cairo presents a fresh answer: an algorithm called RIASO, short for Real-time Integrated Adaptive Stable Offloading.
The core challenge the researchers set out to tame is one of the hardest in edge computing. Wireless channels fluctuate constantly as signals are blocked, interfered with, and faded. Task arrivals are random and unpredictable. A device’s battery drains as it computes or transmits. If tasks arrive faster than the system can process them, the queues of waiting data at the edge server grow without bound, and the whole system becomes unstable. Traditional optimization methods struggle in this environment because they typically assume knowledge of future conditions that simply does not exist. Machine learning approaches, particularly deep reinforcement learning, can adapt to changing conditions, but they often fail to guarantee the long-term stability of the system’s queues and its average power consumption, which are the metrics that determine whether a network remains usable over hours and days rather than seconds.
RIASO attacks this problem by fusing three techniques that have rarely been combined so tightly. The first is Lyapunov optimization, a mathematical framework that converts long-term stability requirements into per-time-frame decisions. Instead of trying to plan far into an uncertain future, Lyapunov optimization maintains virtual queues that track accumulated data backlog and power consumption, and then makes each frame’s decision in a way that pushes those queues back toward zero. This provides a theoretical guarantee of long-term stability without needing to predict the future. The second ingredient is deep reinforcement learning, in which a neural network learns through trial and error which offloading decisions yield the best outcomes. The third, and arguably the most novel, is multi-output learning, which trains a single neural network to produce many candidate solutions simultaneously rather than just one.
The architecture that ties these together is an actor-critic structure, a well-established reinforcement learning pattern in which one component proposes actions and another evaluates them. In RIASO, the actor is a multi-input multi-output convolutional neural network, dubbed MIMO_CNN, that receives three pieces of information for every user device at each time frame: the current wireless channel gain, the length of the device’s data queue, and the length of its virtual power queue. From this input, the network generates not one but multiple candidate binary offloading decisions, each specifying which devices should compute locally and which should offload to the edge server. The critic then evaluates each candidate analytically, solving the remaining resource allocation problem, which is convex once the binary decisions are fixed, using an efficient primal-dual algorithm. The best action is selected, stored in a replay memory, and used to retrain the actor network.
This multi-output design is where RIASO departs most sharply from conventional actor-critic schemes, which typically use a single-output network and rely on random exploration to discover good actions. By generating many diverse candidate solutions in parallel, the MIMO_CNN naturally balances exploration and exploitation: the system can pick the best of many well-informed options rather than gambling on random perturbations of a single guess. The researchers report that this leads to more robust and faster convergence of the training process, which matters enormously in edge networks with high-dimensional action spaces where the number of possible offloading combinations grows exponentially with the number of devices.
Equally inventive is the way RIASO handles its own training schedule. In standard deep reinforcement learning, the model is retrained every time new experience accumulates, which is computationally expensive and can slow the system down. RIASO instead waits until its replay memory is completely full and then trains on the entire batch at once, setting the batch size equal to the memory size so that every stored sample is used and the memory is emptied. Moreover, training is triggered only if the current loss exceeds the minimum of all previous loss values, preventing wasteful retraining when the model is already performing well. The researchers found that this policy cut the training phase from roughly 10,000 time frames down to just 2,000, a fivefold reduction that directly translates into shorter response times and shorter data queues for users.
The team also carefully tuned the network’s structure. They examined how the number of output branches in the MIMO_CNN affects performance, discovering that with only five output layers the system behaved erratically, with unreasonable queue lengths, while configurations with ten or more layers converged reliably. After 2,000 time frames, the average data queue lengths for ten, twenty, and thirty output layers were nearly identical, so the researchers settled on ten as the sweet spot between convergence quality and computational cost. Similarly, they swept the size of the replay memory and sample batch, finding that a size of 256 offered the best balance: smaller memories risked trapping the model in local optima, while larger ones updated the data too slowly and lengthened training time.
To validate RIASO, the researchers ran simulations on a TensorFlow 2.0 platform modeling a multi-user MEC network in which an access point integrated with an edge server assists a set of user devices whose tasks arrive stochastically across consecutive time frames. Channel gains followed a Rician distribution built on a path-loss model, with devices positioned at varying distances from the server. The benchmarks included LyDROO, a well-known Lyapunov-guided deep reinforcement learning framework; DRL-DO, a distributed offloading framework using multiple convolutional networks; and three naive strategies: executing everything locally, executing everything at the edge, and executing randomly. Across the weighted sum computation rate, average data queue length, average power consumption, average energy queue length, and average response time per channel, RIASO delivered the best computation rate and the shortest data queue of all tested methods.
The comparison with LyDROO was particularly instructive. Under identical conditions and parameters, RIASO performed comparably on most metrics but achieved better results on average response time per channel, average data queue length, and average energy queue length, indicating a more stable system overall. DRL-DO showed better average power consumption than RIASO, but at the cost of instability across the other metrics, which the authors argue is a fatal flaw for real deployments where queue blowups mean dropped tasks and angry users. The naive strategies illustrated the trade-offs vividly: local execution minimized response time but produced a low computation rate and high energy queues, while edge execution minimized power consumption but created very long data queues, and random execution was poor on nearly every measure.
The implications reach well beyond the simulation lab. As 5G and future 6G networks push computing closer to users, the offloading decision layer becomes the nervous system of the entire edge infrastructure, and algorithms like RIASO demonstrate that it is possible to be simultaneously fast, energy-aware, and provably stable without clairvoyant knowledge of the future. The authors point to several directions for extending the work, including multitier offloading across edge, fog, and cloud layers, collaborative offloading among neighboring devices with spare resources, and support for partial offloading in which a task’s data is split between local and remote execution. For now, RIASO stands as a compelling proof that combining Lyapunov optimization’s stability guarantees with the adaptive intelligence of multi-output deep reinforcement learning can produce an offloading algorithm that is greater than the sum of its parts, one that keeps the edge responsive even as the wireless world around it refuses to sit still.
Subject of Research: Real-time adaptive computation offloading in multi-access edge computing using Lyapunov optimization and deep reinforcement learning
Article Title: A novel real-time integrated adaptive stable offloading (RIASO) algorithm in multi-access edge computing
Article References: Elhadidy, H., Saleh, H., Rizk, R., & Saber, W. (2026). A novel real-time integrated adaptive stable offloading (RIASO) algorithm in multi-access edge computing. Cluster Computing, 29(14), Article 796. https://doi.org/10.1007/s10586-026-06444-8
Image Credits: AI Generated
DOI: 10.1007/s10586-026-06444-8
Keywords: edge computing, mobile edge computing, computation offloading, deep reinforcement learning, Lyapunov optimization, multi-output learning, Internet of Things, queue stability, actor-critic, convolutional neural network, energy efficiency, wireless networks
Cite Scienmag News
Marilyn Langley. (October 4, 2026). Smart AI Algorithm Keeps Edge Computing Fast and Stable in Real Time. Scienmag. https://scienmag.com/smart-ai-algorithm-keeps-edge-computing-fast-and-stable-in-real-time/
Marilyn Langley. "Smart AI Algorithm Keeps Edge Computing Fast and Stable in Real Time." Scienmag, 4 October 2026, https://scienmag.com/smart-ai-algorithm-keeps-edge-computing-fast-and-stable-in-real-time/. Accessed 4 October 2026.
Marilyn Langley. "Smart AI Algorithm Keeps Edge Computing Fast and Stable in Real Time." Scienmag. October 4, 2026. https://scienmag.com/smart-ai-algorithm-keeps-edge-computing-fast-and-stable-in-real-time/

