Four-legged robots have become astonishingly agile in recent years. Trained with deep reinforcement learning in physics simulators, they can sprint across rubble, leap onto tables taller than their own shoulders, and squeeze through gaps that would defeat wheeled machines. Yet a stubborn gap has persisted between these flashy locomotion skills and the quieter, equally important discipline of navigation: getting from a distant starting point to a distant goal without falling, colliding, or losing the plot. A new study called Skill-Nav, published in the open-access journal Vicinagearth, proposes a deceptively simple fix for that gap, and it hinges on a single design decision about what kind of instruction a robot’s legs should actually receive.
The core insight from the research team, led by Dewei Wang of the University of Science and Technology of China and the Institute of Artificial Intelligence at China Telecom, with colleagues from the Shanghai Artificial Intelligence Laboratory and Northwestern Polytechnical University, is that the interface between a robot’s planner and its controller matters more than either component alone. Most existing quadrupedal navigation systems pass velocity commands downward: the planner tells the controller to move forward at, say, half a meter per second while turning at a certain rate. The problem, the authors argue, is that velocity-tracking controllers accumulate significant errors, especially over rough ground. A robot asked to hold a precise velocity while clambering over a box or skirting a pit will drift, and those small drifts compound into failed missions over long distances.
Skill-Nav replaces the velocity command with a waypoint: a two-dimensional position, expressed relative to the robot’s own body frame, that the robot should reach. Waypoints are sparse, easy for a planner to generate, and forgiving of imprecision. The low-level locomotion policy, trained entirely with reinforcement learning, is free to choose its own gait and trajectory to hit each waypoint, whether that means climbing, jumping, or carefully threading between obstacles. Meanwhile, the high-level planner does not need to know anything about the fine texture of the terrain. It simply hands down a chain of coordinates, and the legs figure out the rest. This division of labor, the researchers show, lets the system combine an agile learned controller with off-the-shelf planning tools, including classical algorithms like A* and even large language models such as GPT-4.
Training the low-level policy took place in two staged scenarios inside the Isaac Gym GPU physics simulator. In the first, called WP-Fixed, waypoints were pre-placed across terrain units drawn from the robot-parkour literature: boxes to climb, gaps to straddle, obstacles to circumvent. The policy learned basic skills such as mounting platforms and steering around hazards, guided by custom reward functions. One reward encouraged the robot to reach as many waypoints as possible per unit of time; another, a so-called stay reward, used an exponential function of the deviation from default joint positions to teach the robot to stand still at a waypoint until the next command arrived. That staying behavior turns out to be essential for a real navigation system, because a planner may need the robot to pause while it computes the next leg of the route.
The second scenario, WP-Random, was designed to break the rigidity of the first. Terrain units were arranged in a grid, waypoints were selected dynamically within ninety degrees of the robot’s heading and within a distance matched to the terrain-unit size, and obstacles of varying dimensions were scattered across the course. Fine-tuning in this scenario forced the robot to handle irregular, consecutive goals rather than a rehearsed sequence. The team also modified the velocity-direction reward, penalizing any behavior whose heading deviated meaningfully from the direction of the target waypoint, and relaxed regularization terms on vertical motion and body orientation so the robot could jump and climb without being punished for it. An ablation comparison confirmed that both stages were necessary: a policy trained only on fixed waypoints failed to track irregular goals, while one trained only on random waypoints developed a chaotic, excessively jumpy gait unsuitable for deployment.
To make the controller deployable on real hardware, the researchers used a teacher-student distillation scheme. The teacher policy enjoyed privileged information, including detailed terrain scans, that no real robot could observe directly. The student policy learned to reconstruct that information from history: proprioceptive signals captured the terrain properties, while depth images from a camera supplied the obstacle geometry. A clever trick called inflated virtual obstacles was introduced during distillation: the obstacles as perceived by the teacher were enlarged without altering the actual simulation geometry or the depth data, training the student to keep a safer margin from hazards. Depth-image noise was also injected during training to narrow the gap between simulated and real cameras, a standard sim-to-real technique that proved important for transfer.
The evaluation was deliberately adversarial. Eighteen simulated robots were deployed per test task across two benchmarks: a single-traverse task, in which all robots crossed a series of obstacles in the same direction within thirty seconds, and an omni-traverse task, in which robots started at the center of a twenty-one-by-twenty-one-meter terrain with random orientations and had to move more than eight and a half meters outward within sixteen seconds. Against baselines including Rapid Motor Adaptation and Extreme Parkour, the full two-stage Skill-Nav policy came out ahead, particularly in the omni-traverse task with high obstacles, where competing policies either could not traverse the terrain at all or drifted toward obstacles they should have avoided. Heatmaps of position visit frequencies showed the Skill-Nav robots reaching farther positions more often, a direct visual signature of more capable locomotion.
The navigation experiments then demonstrated the payoff of the waypoint interface. In simulation, GPT-4 acted as the high-level planner: prompted with a coarse map of two-meter terrain units, a description of the robot’s capabilities, and definitions of the waypoint format, the language model output a sequence of terrain-unit indices that the low-level controller converted into physical traversal. Notably, the LLM’s waypoints sometimes landed in impractical spots, such as inside a gap or at the edge of a box, and the robot could not always recover gracefully; the authors candidly report partial leg suspension and straddling behavior in those anomalous cases. In the real world, the team deployed a Unitree AlienGo quadruped carrying a Jetson Orin NX onboard computer and an Intel RealSense D435 depth camera. The A* algorithm planned paths over an occupancy map that recorded only wall positions, and the resulting path was segmented into waypoints spaced between half a meter and three meters. The robot’s control policy ran at fifty hertz atop a two-hundred-hertz proportional-derivative joint controller, with a motion capture system providing localization.
The real-world results underline why the waypoint abstraction is robust. The robot successfully reached its target while handling obstacles, and it demonstrated recovery behaviors that no planner had explicitly engineered: when it encountered low obstacles that the depth camera failed to detect, it regained its balance and kept going; when external forces pushed it off its path, it corrected and completed the task; and when a waypoint required a sharp turn, it pivoted quickly to align with the new goal. Because the planner only needed coarse-grained information, the system avoided the expensive, tightly coupled training pipelines of fully learned hierarchical approaches such as Barkour and ANYmal Parkour, which demand fine-grained elevation maps and often struggle to generalize beyond their training distribution.
The broader significance of Skill-Nav lies in what it suggests about the architecture of future autonomous robots. As large language models grow more capable of embodied reasoning, the bottleneck is increasingly the interface between symbolic or semantic planning and physical control. Waypoints are a lingua franca: classical graph-search algorithms speak them, language models can emit them from a plain-language prompt, and learned locomotion policies can consume them. The authors acknowledge limitations, including occasional failures at terrain edges and the lack of a fully end-to-end policy, and they point to future work on edge-collision-free controllers and unified locomotion-navigation learning. But the demonstration that a single waypoint-guided policy, trained in two carefully staged simulated scenarios, can carry a real quadruped across complex terrain while obeying instructions from either a decades-old path-planning algorithm or a frontier language model is a compelling template. It hints at robots that will not merely walk impressively, but actually go somewhere.
Subject of Research: Waypoint-guided reinforcement learning for integrating quadrupedal locomotion skills with hierarchical robot navigation
Article Title: Skill-Nav: enhanced navigation with versatile quadrupedal locomotion via waypoint interface
Article References: Wang, D., Bai, C., Li, C., Shi, J., Ding, Y., Zhang, C., & Zhao, B. (2025). Skill-Nav: enhanced navigation with versatile quadrupedal locomotion via waypoint interface. Vicinagearth, 2(1), Article 7. https://doi.org/10.1007/s44336-025-00015-y
Image Credits: AI Generated
DOI: 10.1007/s44336-025-00015-y
Keywords: quadrupedal robots, reinforcement learning, navigation, waypoints, locomotion policy, large language models, A* path planning, sim-to-real transfer, teacher-student distillation, Isaac Gym, obstacle avoidance, hierarchical control
Cite Scienmag News
Violet Maxwell. (September 30, 2026). Waypoints, Not Velocity Commands, Let Four-Legged Robots Master Long-Distance Navigation. Scienmag. https://scienmag.com/waypoints-not-velocity-commands-let-four-legged-robots-master-long-distance-navigation/
Violet Maxwell. "Waypoints, Not Velocity Commands, Let Four-Legged Robots Master Long-Distance Navigation." Scienmag, 30 September 2026, https://scienmag.com/waypoints-not-velocity-commands-let-four-legged-robots-master-long-distance-navigation/. Accessed 30 September 2026.
Violet Maxwell. "Waypoints, Not Velocity Commands, Let Four-Legged Robots Master Long-Distance Navigation." Scienmag. September 30, 2026. https://scienmag.com/waypoints-not-velocity-commands-let-four-legged-robots-master-long-distance-navigation/

