A quadruped robot was reported to jump autonomously through narrow gates, reaching a maximum speed of up to 2.5 metres per second in the described traversal. The result comes from a preprint describing a two-level learning system: one controller supplies movement skills, while another generates the velocity commands used for gate traversal.
The study asks whether a hierarchical reinforcement-learning pipeline can support a quadrupedal robot’s autonomous jump through a narrow gate and extend to other highly dynamic tasks.
A controller built in two layers
The lower layer was trained to track commanded velocity while imitating motion primitives. The authors used adversarial reinforcement learning for that stage, then froze the resulting locomotion policy. A higher-level policy was trained afterward to generate the command velocity for the task.
The two layers operated at different rates. The high-level policy ran at 10 hertz, or 10 updates per second, while the low-level module ran at 50 hertz. The low-level policy remained frozen during high-level training.
For the gate itself, the perception module used an RGB-D camera, which combines ordinary colour images with depth measurements. It was designed to detect black square frames from those colour and depth signals.
From parallel simulation to a physical robot
Low-level training took place in Isaac Gym with 5,480 parallel agents over 25,000 episodes. Actuator-network parameters, including motor strength, latency and offset, were randomized to improve transfer from simulation to real hardware.
Hardware validation used a 22-kilogram Unitree Aliengo quadruped in a 6.0-metre by 3.0-metre motion-capture venue.
What the reported trials showed
In the reported gate traversal, the foot trajectory avoided the gate edge when rear calf-joint deflection exceeded 140 degrees. The airborne phase was close to 0.44 seconds.
The system was also tested with the gate moved laterally from 1.4 metres to minus 1.4 metres in 0.4-metre steps. The experiment was repeated twice at each position, and all trials in that described sequence were completed without a reported failure.
The authors reported average velocity-tracking errors of 0.0258 metres per second for linear velocity and 0.0857 radians per second for angular velocity. No uncertainty interval or other measure of spread for these averages was reported.
The paper states that the proposed hierarchical pipeline consistently outperformed the compared controllers, including the discrete-skill method HierDisc. That comparison is qualitative: numerical differences from the learning curves, collision results and other plots are not tabulated in the supplied text.
A wider set of manoeuvres
The reported system was extended to additional terrain and dynamic-task scenarios. These included continuous stairs, continuous gaps and a slope, as well as tasks involving a hurdle, a gap and vertically arranged boards.
In the arrangement, the high-level module selects commands while the lower layer supplies continuous locomotion skills. Complete success counts for those extensions were not reported in the supplied analysis.
A demonstration with a narrow evidence base
These results are system-level feasibility evidence under the described conditions. The study used one 22-kilogram Aliengo platform, so it does not establish how the controller would perform on other quadrupedal robots or substantially different hardware.
A central limitation is the way the system represents skills through velocity commands. The authors note that this abstraction can limit the policy’s ability to use task-specific joint-coordination patterns beyond those contained in the imitation data.
The reported perception setup was designed to detect black square frames. The supplied analysis also notes that a complete overall success rate and formal inferential statistical analysis were not reported, while the comparative outcomes were described mainly in qualitative terms.
The findings do not show that quadrupedal robots match biological animals across general locomotion tasks, that the hierarchy is superior for every task, or that it preserves fine-grained joint-level control in all settings.
Open questions include how well the controller transfers to other quadrupedal platforms and whether combining high-level control with joint-level control would improve manoeuvre selection when the motion repertoire or gate geometry changes.
The document identifies itself as arXiv:2608.19977v1, dated 20 August 2026. It reports support from the General Research Fund under Grant 17204222, along with support from the Seed Fund for Collaborative Research and the General Funding Scheme-HKU-TCL Joint Research Center for Artificial Intelligence.
Paper data and sources
Original title: Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space
Authors: Zeren Luo, Jiahui Zhang, Yimin Han et al.
Journal/Repository: Advanced Robotics Research (2025)
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text