Tiny robots small enough to swim through the bloodstream have long promised a future in which drugs are delivered to a single diseased cell and surgeons guide machines no wider than a human hair through the most delicate vessels of the brain. The obstacle has rarely been the robots themselves. It has been the intelligence that must steer them. Deep reinforcement learning, the same family of techniques that taught computers to master Go and control fusion plasmas, has emerged as the most promising way to give microrobots autonomous navigation skills. The catch has been time: training a navigation policy could take hours or even days of computation, making it painfully slow to iterate on robot designs, tune parameters, or adapt to new clinical scenarios. A team of researchers in Hong Kong and Harbin now reports in Nature Machine Intelligence a framework that collapses that timeline to minutes, training effective navigation policies in under ten minutes while preserving, and in some respects improving, the quality of the resulting behavior.
The study, led by Yinghan Sun and Lidong Yang of The Hong Kong Polytechnic University together with colleagues including Li Zhang of The Chinese University of Hong Kong and Huijun Gao of the Harbin Institute of Technology, attacks the training bottleneck from two directions at once. The first is raw computational throughput. The researchers built a fully vectorized simulator containing more than 10,000 artificial vascular environments, each a procedurally generated network of vessels in which a virtual microrobot must find its way to a target while avoiding walls, obstacles and flow disturbances. Rather than stepping through environments one at a time, the simulator parallelizes the robot dynamics, the visual feature extraction and the feasibility checks across thousands of environments simultaneously, achieving roughly 190,000 environment transitions per second. That throughput figure is the engine of the entire result: reinforcement learning improves through sheer volume of experience, and generating experience two orders of magnitude faster translates directly into training that finishes before a coffee goes cold.
The technical machinery behind that speed deserves a closer look, because it illustrates a broader trend in robotics research toward GPU-native simulation. In conventional training pipelines, each simulated robot occupies its own process or thread, and the neural network that decides actions must be queried separately for each environment. Vectorized simulation instead treats the entire fleet of environments as a single batched data structure. The equations of motion for thousands of microrobots are advanced in lockstep as tensor operations, the ray-casting routines that let each robot perceive its surroundings are computed for all agents at once, and collision or feasibility checks are evaluated as batched logical operations. The policy network is likewise evaluated in one large forward pass. This eliminates most of the overhead that normally dominates simulation time, and it means the learning algorithm, in this case a proximal policy optimization variant, is fed a continuous torrent of diverse experience rather than a trickle.
Speed alone, however, is a dangerous commodity in reinforcement learning. A fast simulator can simply produce bad policies faster, especially when the reward signal that guides learning is poorly designed. Naive reward functions for navigation tend to produce agents that oscillate, jitter or hug obstacles too closely, because the reward landscape rewards progress toward the goal without sufficiently penalizing erratic control or risky proximity. The second pillar of the new framework is therefore a reward design the authors call task-shaping-regularization, or TSR. It combines three ingredients: a task component that rewards reaching the navigation goal, a shaping component that provides graded feedback as the robot makes progress through the vascular maze, and a regularization component that discourages erratic actions and unsafe clearance margins. The regularizer is what tames the jitter, and the shaping term is what accelerates convergence by giving the learning algorithm a smooth gradient of feedback rather than a sparse reward delivered only on success.
The measured benefits of the TSR framework are specific. Across all evaluated scenarios, the framework reduced action variation by at least 33.7 percent, meaning the trained policies issue smoother, more deliberate control commands rather than rapid oscillations that would be difficult for physical actuation systems to follow. It also increased obstacle clearance by at least 2.1 percent, a modest-sounding margin that matters enormously at micrometer scales, where a few microns of extra clearance can be the difference between a clean transit through a vessel and a collision that strands the robot or damages tissue. The framework also improved final task performance and accelerated convergence, so the policies that emerged from the ten-minute training runs were not merely fast to produce but genuinely competitive with, and in several metrics superior to, policies trained under conventional reward schemes for far longer.
Perhaps the most striking claim in the paper is that the resulting policies support zero-shot deployment, meaning they can be transferred directly from simulation to physical microrobots, and across different microrobot types and navigation scenarios, without any additional training or fine-tuning. The researchers validated this in experiments spanning multiple robot platforms, including magnetically driven helical swimmers and surface-rolling microrobots, in channel environments, in dynamic settings with moving obstacles, and in fluid flow conditions that mimic the perturbations of a living circulatory system. They also demonstrated navigation in three-dimensional scenarios with physical obstacles, including a model of human brain vasculature, one of the most demanding imagined use cases for medical microrobots. The supplementary materials accompanying the paper include videos of these sim-to-real transfers, showing trained policies guiding real devices through environments they had never encountered during training.
Zero-shot transfer of this kind depends on the diversity of the training distribution. Because the simulator contains more than ten thousand distinct vascular environments, drawn from a dataset the team has released publicly, the learned policy is forced to generalize rather than memorize. A policy trained in a handful of environments tends to overfit to their particular geometry, failing catastrophically when the vessel bends differently or an unexpected obstacle appears. A policy trained across ten thousand geometries, with varied flow conditions and obstacle configurations, must instead learn the underlying structure of the navigation problem: how to balance progress against clearance, how to react to visual features of vessel walls, how to maintain control authority in flow. The large-scale dataset used for policy training is available through Zenodo, and the complete codebase, named mr-nav, is available on GitHub with an archived version also deposited on Zenodo, an openness that should allow other groups to reproduce the results and extend the framework to their own robot platforms.
The practical implications extend well beyond a single laboratory’s convenience. In microrobotics, the design loop couples the physical device, the actuation system and the control policy: change the robot’s geometry or magnetization profile and the optimal navigation strategy changes with it. When policy training takes days, researchers are effectively locked out of rapid co-design, because every design iteration demands a fresh, expensive training campaign. When training takes minutes, a researcher can sweep through dozens of candidate designs in a single day, evaluating how each performs under a learned controller, and can re-optimize policies whenever the clinical scenario changes. The authors argue that this substantially shortens the design loop and accelerates the deployment of autonomous microrobots, and the arithmetic supports them: a hundredfold reduction in training time is not an incremental improvement but a change in what kinds of experiments are feasible at all.
The work also lands at a moment of visible momentum for learning-based microrobot control. Recent years have seen reinforcement learning applied to ultrasound-driven microrobots, magnetic helical swimmers, microswarm formation control and three-dimensional positional control, with papers appearing in Nature Machine Intelligence, Science Robotics and IEEE’s robotics transactions. What has distinguished these efforts, and limited them, is the cost of training. The new framework suggests that the field’s computational bottleneck was not intrinsic but architectural, a consequence of simulation pipelines that were never designed for the throughput that modern deep reinforcement learning demands. If vectorized simulation and carefully regularized reward design become standard practice, the barrier to entry for autonomous microrobot navigation drops sharply, potentially bringing dozens of laboratories with strong device fabrication skills but limited machine learning infrastructure into the autonomous navigation arena.
Challenges remain before minute-scale training translates into clinical microrobots. Real vascular environments present imaging noise, physiological motion, complex pulsatile flow and safety constraints that no simulator fully captures, and the gap between the ten thousand training environments and any individual patient’s anatomy will require careful validation. The repeated-trial analyses reported in the paper’s extended data, tracking success rate, completion time, path efficiency and obstacle clearance across fifteen independent trials per trajectory, represent a serious attempt to quantify reliability, but long-horizon behavior inside living organisms remains the ultimate test. Still, the core achievement stands on its own terms: a demonstration that the intelligence for autonomous microrobot navigation can be produced at the pace of experimentation rather than the pace of overnight computation. For a field whose devices are measured in micrometers, that may prove to be the acceleration that matters most.
Subject of Research: Minute-scale deep reinforcement learning training for autonomous microrobot navigation in vascular environments
Article Title: Minute-scale training for microrobot navigation
Article References: Sun, Y., Zhu, A., Ji, X., Li, Y., Zhao, J., Wang, Y., Zhang, L., Gao, H., & Yang, L. (2026). Minute-scale training for microrobot navigation. Nature Machine Intelligence. https://doi.org/10.1038/s42256-026-01305-w
Image Credits: AI Generated
DOI: 10.1038/s42256-026-01305-w
Keywords: microrobots, deep reinforcement learning, autonomous navigation, vectorized simulation, sim-to-real transfer, magnetic microrobots, vascular navigation, reward shaping, Nature Machine Intelligence, targeted drug delivery, policy training, zero-shot deployment
Cite Scienmag News
Denise Maddox. (September 30, 2026). Microrobots Learn to Navigate Blood Vessels in Under Ten Minutes of Training. Scienmag. https://scienmag.com/microrobots-learn-to-navigate-blood-vessels-in-under-ten-minutes-of-training/
Denise Maddox. "Microrobots Learn to Navigate Blood Vessels in Under Ten Minutes of Training." Scienmag, 30 September 2026, https://scienmag.com/microrobots-learn-to-navigate-blood-vessels-in-under-ten-minutes-of-training/. Accessed 30 September 2026.
Denise Maddox. "Microrobots Learn to Navigate Blood Vessels in Under Ten Minutes of Training." Scienmag. September 30, 2026. https://scienmag.com/microrobots-learn-to-navigate-blood-vessels-in-under-ten-minutes-of-training/

