Getting two streams of gas to mix when they are both tearing through a machine at supersonic speeds is one of the deceptively hard problems in aerospace engineering. At such velocities, air and fuel have only microseconds to mingle before the flow exits the combustion chamber, and any incomplete mixing translates directly into lost thrust, wasted propellant, or a weaker laser beam. A team of researchers at the Dalian Institute of Chemical Physics of the Chinese Academy of Sciences has now shown that a deep reinforcement learning system, equipped with a cleverly tuned sampling strategy, can find optimal pulsed-injection settings for supersonic mixing while consuming far fewer of the expensive simulations that normally make such optimization prohibitive. The work, published in the International Journal of Aeronautical and Space Sciences, targets devices as different as scramjet engines, chemical lasers, and supersonic ejectors, all of which live or die by how well their high-speed flows blend.
The central obstacle the researchers set out to attack is not a lack of algorithms but a lack of data. Every candidate design in a supersonic flow problem must be evaluated with computational fluid dynamics simulations that resolve unsteady, compressible, reacting flow, and each of those simulations can take substantial computing time and resources. Reinforcement learning, which typically improves by trial and error across thousands or millions of attempts, is notoriously hungry for exactly this kind of feedback. When each trial costs a full simulation of a shock-laden supersonic shear layer, the standard playbook collapses under its own appetite. The team, led by Meng You, Tingting Liu, Tianzi Bai, Shuqin Jia, and Ying Huai, therefore framed their contribution around sample efficiency: extracting the maximum design insight from the smallest possible number of flow simulations.
Their solution combines two ideas. The first is a deep policy gradient network, a class of reinforcement learning agent that learns a probability distribution over actions and nudges that distribution toward actions that yield better outcomes. In this application, the actions are the parameters of a pulsed fuel injection scheme: the pulsed frequency, the pulsed amplitude, and the mean total pressure of the injected jet. The second idea is an adaptive noise-scaling framework governing how the agent explores. In policy gradient methods, randomness in the action selection is what allows the agent to discover unfamiliar regions of the design space, but too much noise wastes evaluations on poor candidates, while too little traps the agent in mediocre local optima. The adaptive scheme dynamically adjusts the scale of this exploration noise, balancing exploration against exploitation as the optimization proceeds.
A distinctive feature of the architecture is the role of a predictor deep neural network. Rather than requiring a full flow simulation for every candidate the policy network considers, the predictor provides fast approximate feedback on how a given combination of pulse frequency, amplitude, and mean total pressure is likely to perform. This internal surrogate lets the policy network adjust its parameters toward the desired mixing performance without paying the full simulation cost at every step. The predictor effectively acts as a cheap stand-in for the physics, absorbing much of the trial-and-error burden, while the true simulations are reserved for verifying and calibrating the most promising candidates. The result is an optimization loop in which expensive computational fluid dynamics evaluations are spent sparingly and deliberately.
The physical system the team used as its proving ground is a chemical laser, specifically a supersonic flow device in which mixing quality is measured by the small-signal gain coefficient, a quantity that directly determines laser performance. In a chemical oxygen-iodine laser, reactive streams must mix rapidly at supersonic speed for the lasing reaction to proceed efficiently, making the small-signal gain coefficient an unusually sensitive and practically meaningful yardstick for mixing. Pulsed injection itself is a well-established mixing enhancement technique: by modulating the injected jet in time, engineers can excite flow structures that stir the streams far more effectively than a steady jet of the same average strength. The trick has always been choosing the right modulation parameters, a search space large enough that manual tuning or exhaustive sweeps quickly become impractical.
The numerical foundation of the study was established with careful verification. The team constructed three computational grids of increasing refinement, containing 60,000, 120,000, and 540,000 cells respectively, and compared Mach number distributions along the centerline for a representative operating condition with a pulse frequency of 100 kilohertz and a mean total pressure of 250 Torr. The distributions from the two finer grids overlapped almost completely, indicating that 120,000 cells were sufficient to capture the physics. A parallel time-step independence study tested four temporal resolutions, from one-fiftieth to one-four-hundredth of the pulse period, and found that the time-averaged small-signal gain coefficient changed by less than two percent between the two finest steps. On this basis, the researchers adopted the medium grid with the T/200 time step, a pragmatic compromise that kept each simulation affordable without sacrificing accuracy.
With the simulation framework validated, the optimization proceeded on two fronts. The team demonstrated both single-objective and multi-objective optimization of the pulsed injection parameters, the latter addressing the reality that mixing performance often involves trade-offs among competing metrics that cannot all be maximized simultaneously. The adaptive noise-scaling sampling strategy proved decisive: compared with a deep policy gradient network using a conventional sampling strategy, the new framework achieved both superior accuracy and superior stability in locating optimal injection settings. Stability matters as much as accuracy in this context, because an optimizer that lands on a different answer every run is of little use to an engineer who needs a trustworthy design point.
The ultimate test came from verifying the optimized parameters against full simulations of supersonic reactive flows in the chemical laser. The relative errors of the small-signal gain coefficient predicted by the optimization framework, compared with the simulation results, remained below five percent across the cases examined. That level of agreement indicates that the learned policy and its predictor network were not merely exploiting artifacts of the training data but had genuinely captured the relationship between injection parameters and mixing performance. For a quantity as sensitive as small-signal gain in a reacting supersonic flow, sub-five-percent error from a data-efficient learning framework is a substantial result, and it suggests the approach could transfer to other high-speed flow applications where the same trade-off between data cost and optimization quality exists.
The implications reach well beyond chemical lasers. Scramjet engines, which hold fuel and air at supersonic speeds throughout combustion, face perhaps the most demanding mixing problem in propulsion, and pulsed fuel injection has attracted growing attention there, including recent studies of hydrogen-fueled and kerosene-fueled supersonic combustors. Supersonic ejectors and other high-speed flow devices share the same fundamental challenge. In all of these fields, the bottleneck has been the same: the simulations needed to evaluate a design are expensive, so the space of possible designs goes largely unexplored. A framework that demonstrably reduces the data burden while preserving optimization accuracy changes the economics of that exploration, potentially allowing engineers to survey design spaces that were previously out of reach.
The study also sits within a broader movement in fluid dynamics toward machine-learning-driven flow control. Recent years have seen deep reinforcement learning applied to active control of turbulent separation bubbles, jet mixing optimization using Bayesian methods and evolutionary algorithms, and machine learning approaches to scramjet combustor configuration. What distinguishes the Dalian work is its explicit focus on the sample-efficiency problem that has limited practical adoption, and its demonstration on a real, industrially relevant quantity, the gain of a chemical laser, rather than a purely academic benchmark. By coupling a policy gradient agent with adaptive noise scaling and a learned predictor, the researchers offer a template that other groups working on data-limited optimization problems, from propulsion to materials design, may find adaptable. The work was supported by the Chinese Academy of Sciences and the Dalian National Laboratory for Clean Energy, and the authors state they have no competing interests. As hypersonic flight and directed-energy systems mature, tools that squeeze better performance out of every costly simulation are likely to become as valuable as the physics they help to optimize.
Subject of Research: Deep reinforcement learning optimization of pulsed injection parameters for enhancing supersonic mixing in high-speed flow devices such as chemical lasers and scramjets
Article Title: Optimization of Pulsed Injection for Enhancing Supersonic Mixing via Deep Reinforcement Learning with Adaptive Noise Scaling
Article References: You, M., Liu, T., Bai, T., Jia, S., & Huai, Y. (2026). Optimization of Pulsed Injection for Enhancing Supersonic Mixing via Deep Reinforcement Learning with Adaptive Noise Scaling. International Journal of Aeronautical and Space Sciences. https://doi.org/10.1007/s42405-026-01246-z
Image Credits: AI Generated
DOI: 10.1007/s42405-026-01246-z
Keywords: supersonic mixing, deep reinforcement learning, pulsed injection, policy gradient network, adaptive noise scaling, chemical laser, scramjet, computational fluid dynamics, small-signal gain coefficient, sample efficiency, flow control, multi-objective optimization
Cite Scienmag News
Audrey Campbell. (September 30, 2026). AI Learns to Master Supersonic Mixing With Far Fewer Costly Simulations. Scienmag. https://scienmag.com/ai-learns-to-master-supersonic-mixing-with-far-fewer-costly-simulations/
Audrey Campbell. "AI Learns to Master Supersonic Mixing With Far Fewer Costly Simulations." Scienmag, 30 September 2026, https://scienmag.com/ai-learns-to-master-supersonic-mixing-with-far-fewer-costly-simulations/. Accessed 30 September 2026.
Audrey Campbell. "AI Learns to Master Supersonic Mixing With Far Fewer Costly Simulations." Scienmag. September 30, 2026. https://scienmag.com/ai-learns-to-master-supersonic-mixing-with-far-fewer-costly-simulations/

