Evolutionary algorithms have long been the workhorses of optimization, mimicking the blunt logic of natural selection to grind out solutions to problems that defeat conventional mathematics. But a persistent weakness has haunted them for decades: they tend to use the same genetic operator, the same reproductive recipe, throughout an entire run, regardless of whether that recipe is actually helping. Now, a team of researchers in China and Australia has built an algorithm that learns, on the fly, which operator to deploy at every moment of the search — and it does so by borrowing one of the most powerful ideas in modern machine learning, the Deep Q-Network. The result, published in Complex & Intelligent Systems, is a framework called MTEA-RLKS that significantly outperforms six state-of-the-art competitors on two widely used benchmark suites.
The problem the researchers set out to solve sits at the intersection of two challenging branches of evolutionary computation. Multi-objective optimization asks an algorithm to find not one best answer but an entire frontier of trade-off answers — the Pareto front — when goals conflict, as they almost always do in engineering design. Multi-task optimization goes further, demanding that several related problems be solved simultaneously within a single run, with useful knowledge flowing between them so that progress on one task accelerates progress on the others. Combine the two and you get multi-objective multi-task optimization, or MO-MTO, a field that has grown rapidly because so many real-world scenarios — scheduling a factory while tuning its production line, designing a component while optimizing its manufacturing process — naturally present themselves as families of coupled problems.
Most existing MO-MTO algorithms, however, share two structural blind spots. First, they typically rely on a single genetic operator to generate offspring, meaning the crossover and mutation rules applied to parent solutions never change as the population evolves. This is a serious handicap, because the search landscape of an optimization problem changes character dramatically over time: aggressive, exploration-heavy operators are valuable early on, when the population is scattered and the algorithm needs to cover ground, while fine-grained, exploitation-focused operators matter later, when the population has converged near the Pareto front and precision matters more than reach. A fixed operator is a compromise that is rarely optimal at any stage. Second, when these algorithms transfer knowledge between tasks, they confine that transfer to the objective space — the space where solutions are scored — and ignore the decision space, the space of the actual variables being optimized, where much of the most valuable structural information about a problem resides.
MTEA-RLKS attacks both weaknesses at once. Its authors, led by Xuanwei Zhang, Zhiyuan Pan, and Long Chen of Beijing University of Chemical Technology, together with Yushu Du of the University of Melbourne and Junze Zhu of Tongji University, designed a collaborative knowledge transfer mechanism that draws on both spaces simultaneously. In the objective space, the algorithm extracts what the team calls static knowledge: statistical summaries of the global population distribution and of local neighborhoods, capturing where high-quality solutions sit relative to one another. This static knowledge is stable enough to be shared across tasks without much risk of misleading the recipient. In the decision space, by contrast, the algorithm tracks dynamic knowledge — the way the population’s configuration shifts as evolution proceeds — using Gaussian processes, a class of probabilistic models that can represent distributions over functions and quantify uncertainty about how the population will move next.
The division of labor between static and dynamic knowledge is the conceptual heart of the method. Static knowledge, extracted from global distributions and local neighborhoods, provides a coarse map of where promising regions lie, which can be safely exported to a related task even when the two problems differ in detail. Dynamic knowledge, modeled with Gaussian processes, is more like a weather forecast for the evolving population: it tells the algorithm not just where the population is, but how it is trending, allowing the transferred information to be adapted to the recipient task’s current stage of convergence. By synergizing the two, the algorithm guides the generation of offspring that are more likely to be high quality, accelerating the optimization process on every task in the portfolio rather than trading one task’s performance against another’s.
The second pillar of the framework is the part most likely to resonate beyond the evolutionary computation community: a Deep Q-Network, the same architecture that made headlines when it learned to play Atari games from raw pixels, is used to select genetic operators adaptively. In reinforcement learning terms, the state is the current situation of the evolving population, the actions are the available genetic operators, and the reward reflects how much the chosen operator improves the solutions. Crucially, the network is trained online, during the iterations of the evolutionary run itself, so its policy adapts to the evolutionary needs of different solutions as they arise. Early in a run, the network can learn to favor exploratory operators; as the population converges, it can shift toward operators that refine solutions along the Pareto front. No human expert needs to specify a schedule — the algorithm discovers one.
This marriage of deep reinforcement learning with evolutionary search is more than a technical convenience. Adaptive operator selection has been studied for years, usually through bandit-style methods or hand-crafted rules that track simple statistics such as operator success rates. A Deep Q-Network brings something richer: the ability to recognize complex, high-dimensional patterns in the population state and to learn long-horizon strategies rather than myopic one-step choices. In a multi-task setting this matters enormously, because the state of one task’s population can legitimately influence what operator is best for another task, and a deep value function is far better positioned to capture such cross-task dependencies than a running average of success counts.
The empirical case for the method rests on the two standard benchmark suites for this problem class, CEC2017 and CEC2019, which contain collections of multi-objective multi-task problems designed by the evolutionary computation community specifically to stress knowledge transfer and operator choice. Across these test sets, MTEA-RLKS was compared against six state-of-the-art algorithms, and the reported results show that it significantly outperformed all of them. The benchmarks are demanding precisely because they include pairs of tasks with varying degrees of relatedness — some pairs share useful structure, others are nearly independent — so an algorithm that transfers knowledge indiscriminately can be actively harmed by its partner task. The collaborative mechanism’s blend of cautious static transfer and adaptive dynamic transfer appears to navigate that hazard well.
The practical implications stretch across the domains where multi-objective multi-task problems arise naturally. Engineering design routinely involves optimizing several coupled objectives — weight, strength, cost, energy efficiency — across a family of related variants of a product, and an algorithm that can solve the whole family at once, transferring lessons between variants, offers real savings in computation and engineering time. Scheduling, logistics, neural architecture search, and resource allocation all present similar structures. The funding acknowledgments behind the work, including China’s National Key Research and Development Program and the National Natural Science Foundation, signal that the application ambitions are substantial, spanning projects in intelligent manufacturing and information technology.
There are, of course, the usual caveats that accompany any new algorithmic result. The Deep Q-Network adds computational overhead and hyperparameters of its own, and online training during an evolutionary run raises questions about stability that the benchmarks address but real deployments will probe further. The Gaussian process models of dynamic knowledge, while elegant, carry their own scaling considerations as decision dimensions grow. Yet the direction is unmistakable: the boundary between evolutionary computation and deep reinforcement learning, long treated as separate traditions, is becoming a productive seam. MTEA-RLKS demonstrates that a neural network can serve as the decision-making brain of an evolutionary algorithm, choosing its reproductive tools moment by moment while the algorithm itself ferries knowledge between problems in two spaces at once. If the approach generalizes beyond the benchmarks as well as its authors hope, the fixed-operator evolutionary algorithm may soon look like a relic — a one-tool craftsman in a world that has learned to carry a full toolbox and, better still, to know which tool to reach for.
Subject of Research: Adaptive operator selection using deep reinforcement learning for decomposition-based multi-objective multi-task evolutionary optimization
Article Title: Adaptive operator selection via deep reinforcement learning for decomposition-based multi-objective multi-task optimization
Article References: Zhang, X., Pan, Z., Du, Y., Zhu, J., & Chen, L. (2026). Adaptive operator selection via deep reinforcement learning for decomposition-based multi-objective multi-task optimization. Complex & Intelligent Systems. https://doi.org/10.1007/s40747-026-02387-0
Image Credits: AI Generated
DOI: 10.1007/s40747-026-02387-0
Keywords: evolutionary algorithms, multi-objective multi-task optimization, reinforcement learning, Deep Q-Network, adaptive operator selection, knowledge transfer, Gaussian processes, Pareto front, CEC2017 benchmarks, CEC2019 benchmarks, decomposition-based optimization, machine learning
Cite Scienmag News
Gavin Prescott. (October 7, 2026). Deep Reinforcement Learning Teaches Evolutionary Algorithms to Pick the Right Tool. Scienmag. https://scienmag.com/deep-reinforcement-learning-teaches-evolutionary-algorithms-to-pick-the-right-tool/
Gavin Prescott. "Deep Reinforcement Learning Teaches Evolutionary Algorithms to Pick the Right Tool." Scienmag, 7 October 2026, https://scienmag.com/deep-reinforcement-learning-teaches-evolutionary-algorithms-to-pick-the-right-tool/. Accessed 7 October 2026.
Gavin Prescott. "Deep Reinforcement Learning Teaches Evolutionary Algorithms to Pick the Right Tool." Scienmag. October 7, 2026. https://scienmag.com/deep-reinforcement-learning-teaches-evolutionary-algorithms-to-pick-the-right-tool/

