When a test bench fails in the middle of a batch of unmanned aerial vehicle trials, the consequences ripple far beyond the broken machine. Downstream validation tasks stall, technicians sit idle, and the entire testing calendar can slip by days or weeks. A research team in Beijing has now unveiled a hybrid artificial intelligence framework that promises to keep such disruptions from cascading, and the results suggest that a marriage between modern reinforcement learning and a decades-old search technique may be exactly what time-critical industrial scheduling has been waiting for.
The study, published in Applied Intelligence by Zhibin Mao, Yan Gao, Qian Pu, Minghui Wang and Haikuo Shen of Beijing Jiaotong University and the China Academy of Launch Vehicle Technology, tackles a problem that has long frustrated test facility managers: dynamic rescheduling under equipment failure. In their formulation, when a test item breaks down, an estimated repair time becomes available at the very moment of the disruption. The scheduler must then decide, almost instantly, how to reassign remaining tasks across constrained resources so that the overall completion time, known as the makespan, suffers as little as possible.
Mathematically, the researchers reformulated this recovery problem as a static resource-constrained multi-project scheduling problem, or RCMPSP, a notoriously hard combinatorial optimization challenge. In an RCMPSP, multiple projects compete for a limited pool of shared resources, and each task must respect precedence relations, meaning certain activities cannot begin until their predecessors finish. Finding an optimal schedule is computationally intractable for realistic instance sizes, which is why practitioners have historically relied on simple priority rules that assign tasks in a fixed order of importance, accepting suboptimal outcomes in exchange for speed.
The new method, dubbed GRPO-TS, departs from that tradition with a two-stage architecture. In the first stage, a policy network trained with Group Relative Policy Optimization, a reinforcement learning algorithm that has attracted wide attention for its efficiency in large language model training, generates an initial feasible schedule. Unlike conventional proximal policy optimization, GRPO evaluates groups of candidate actions relative to one another, which can stabilize learning and reduce the variance of policy updates. The authors adapted this idea to the scheduling domain, letting the network learn how to sequence competing test tasks under resource contention.
A key technical innovation lies in the architecture of the policy network itself. The researchers equipped it with a mixture-of-experts module, a design in which specialized subnetworks activate selectively depending on the input, allowing different experts to specialize in different scheduling regimes, such as periods of heavy resource competition versus periods dominated by precedence constraints. They also incorporated historical feature fusion, feeding the network information about past states so that it can capture how resource conflicts and task dependencies evolve over the course of a testing campaign. This gives the learned policy a form of temporal awareness that static priority rules fundamentally lack.
Yet a learned policy alone rarely produces a truly polished schedule. That is where the second stage comes in: Tabu Search, a classical metaheuristic introduced in the late 1980s, refines the initial solution through local neighborhood moves, systematically swapping and repositioning tasks while maintaining a tabu list that forbids recently revisited solutions to escape local optima. The division of labor is elegant. The neural policy supplies a high-quality starting point in milliseconds, and the metaheuristic polishes it with targeted local improvements, avoiding the wasteful random exploration that often makes pure metaheuristics slow on large instances.
The experimental evidence is striking. On extended RCMPSP benchmark instances, GRPO-TS achieved the lowest normalized average makespan among all tested baselines, including both classical priority-rule dispatching and modern deep reinforcement learning approaches. It also outperformed classical metaheuristics on medium and large instances, suggesting that the hybrid strategy scales better than either pure learning or pure search. Perhaps most importantly for real-world deployment, the online rescheduling time ranged from just 0.376 to 1.837 seconds, fast enough to re-plan a disrupted testing campaign before technicians have even finished diagnosing the failed equipment.
Robustness to uncertainty was another focus of the evaluation. Estimated repair times are, by definition, estimates, and real maintenance operations routinely deviate from predictions. The team stress-tested their framework by perturbing the estimated repair times by up to plus or minus thirty percent on mixed-scale instances. Even under these perturbations, the makespan increase did not exceed 9.14 percent, indicating that the rescheduled plans remain near-optimal even when the underlying assumptions about repair duration turn out to be substantially wrong. For test facilities where a single day of delay can cost significant sums, that kind of resilience is a meaningful guarantee.
The work sits within a broader and rapidly growing research movement that hybridizes reinforcement learning with evolutionary and local search methods. Recent surveys have documented a surge of such algorithms across domains from satellite scheduling to electric vehicle routing, on the logic that learned heuristics can guide classical optimizers toward promising regions of the search space while the optimizers supply the fine-grained refinement that neural networks struggle to achieve alone. The GRPO-TS framework is a particularly clean instantiation of this philosophy, and its application to UAV batch testing gives it a concrete industrial anchor rather than a purely academic benchmark.
The implications extend beyond drone testing. Resource-constrained multi-project scheduling arises in aircraft maintenance, construction, semiconductor fabrication, cloud computing and any setting where multiple concurrent workloads compete for scarce machines and personnel. The authors note that their data generator, parameter settings and source code are available from the corresponding author upon reasonable request, subject to institutional data-sharing policies, which should facilitate replication and adaptation by other groups. Supported by China’s National Key Research and Development Program, the research signals a future in which the moment a test rig fails, an intelligent scheduler quietly rebuilds the entire plan in under two seconds, and the production line barely notices. For an industry racing to certify fleets of autonomous aircraft, that future may arrive sooner than expected.
Subject of Research: Dynamic rescheduling of resource-constrained UAV batch testing using a hybrid reinforcement learning and Tabu search framework
Article Title: A two-stage framework integrating group relative policy optimization with Tabu search for dynamic rescheduling in UAV testing
Article References: Mao, Z., Gao, Y., Pu, Q., Wang, M., & Shen, H. (2026). A two-stage framework integrating group relative policy optimization with Tabu search for dynamic rescheduling in UAV testing. Applied Intelligence, 56(15), Article 457. https://doi.org/10.1007/s10489-026-07449-x
Image Credits: AI Generated
DOI: 10.1007/s10489-026-07449-x
Keywords: UAV testing, dynamic rescheduling, group relative policy optimization, Tabu search, resource-constrained multi-project scheduling, reinforcement learning, metaheuristics, makespan optimization, mixture-of-experts, equipment failure, combinatorial optimization, artificial intelligence
Cite Scienmag News
Denise Maddox. (September 30, 2026). AI Meets Classic Search: New Two-Stage Algorithm Keeps Drone Testing on Schedule When Equipment Fails. Scienmag. https://scienmag.com/ai-meets-classic-search-new-two-stage-algorithm-keeps-drone-testing-on-schedule-when-equipment-fails/
Denise Maddox. "AI Meets Classic Search: New Two-Stage Algorithm Keeps Drone Testing on Schedule When Equipment Fails." Scienmag, 30 September 2026, https://scienmag.com/ai-meets-classic-search-new-two-stage-algorithm-keeps-drone-testing-on-schedule-when-equipment-fails/. Accessed 30 September 2026.
Denise Maddox. "AI Meets Classic Search: New Two-Stage Algorithm Keeps Drone Testing on Schedule When Equipment Fails." Scienmag. September 30, 2026. https://scienmag.com/ai-meets-classic-search-new-two-stage-algorithm-keeps-drone-testing-on-schedule-when-equipment-fails/

